Observability Engineer
Location: Liverpool (Hybrid)
Salary: £60,000 – £75,000 per year + excellent benefits package
Job Type: Permanent
Industry: Transportation / Cloud & Infrastructure
Discipline: Cloud, Infrastructure and Services
We are Bridge Staffing Agency, a dedicated recruitment partner committed to connecting exceptional talent with organizations where they can make a lasting impact. We take pride in building collaborative relationships and tailoring our search solutions to your professional goals, your pace, and your people.
We are partnering with a large-scale transportation organisation undergoing significant technology investment to recruit a skilled Observability Engineer.
In this pivotal role, you will take ownership of the monitoring and observability landscape across infrastructure, platforms, networks, and applications. You will work closely with Service Delivery, Infrastructure, and Engineering teams to ensure critical services are measurable, highly visible, and proactively managed through modern telemetry, alerting, and analytics.
Design, implement, and continuously mature enterprise-wide observability and monitoring capabilities.
Manage and optimize telemetry, monitoring, and alerting platforms across hybrid environments.
Develop intuitive, high-impact dashboards and operational reporting that provide actionable insights.
Support incident, problem, and major incident processes through proactive telemetry and root-cause analysis.
Establish and champion observability standards, guidelines, and best practices across technology teams.
Drive continuous service performance improvements using trend analysis and system health data.
Ensure observability and telemetry requirements are embedded into new service transitions from day one.
Proven experience building, operating, and maturing observability solutions within complex, hybrid environments.
Strong hands-on knowledge of telemetry, logs, metrics, and distributed tracing.
Expertise in dashboard creation, alerting strategy optimization, and performance reporting.
Experience with infrastructure, network, and application performance monitoring (APM).
Familiarity with hybrid technology estates incorporating cloud platforms (e.g., Microsoft Azure) and enterprise infrastructure (e.g., VMware, Hyper-V, Veeam, dHCI).
Background in supporting incident management, root cause analysis, and operational resilience initiatives.