SRE Technical Lead- Bristol

FDM Group · Bristol

  • Not stated by the employer
  • Temporary
  • Posted 1 month ago
Apply on FDM Group's site

FDM is a global business and technology consultancy seeking an SRE Technical Lead to support a major global financial services organisation as it establishes its first formal Site Reliability Engineering function. This is initially a 12-month contract with the potential to go permanent and will be a hybrid role based in Bristol. This role offers a unique opportunity to play a key technical leadership role within a newly formed SRE function. Reporting directly to the Head of SRE, you will be responsible for driving the adoption of reliability engineering practices across critical business services, helping to improve service stability, resilience, observability, and operational efficiency.

As a senior individual contributor, you will provide technical leadership rather than people management. You will work closely with platform, infrastructure, engineering, and support teams to implement SRE principles, define reliability standards, reduce operational toil through automation, and establish meaningful service health measurements. You will help accelerate the organisation's transition from reactive production support towards a proactive, engineering-led reliability model, ensuring reliability is designed into services rather than addressed after incidents occur

Responsibilities:

  • Partner with the Head of SRE to implement and embed the organisation's SRE strategy, operating model, and reliability standards.
  • Act as a technical authority for reliability engineering, providing guidance and expertise across application, platform, and infrastructure teams.
  • Define, implement, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Critical User Journeys (CUJs) to establish meaningful service reliability metrics.
  • Drive adoption of SLO-based decision making, supporting teams in balancing reliability, delivery velocity, and operational risk.
  • Identify opportunities to reduce operational toil through automation, including runbook automation, self-healing capabilities, deployment improvements, and recovery processes.
  • Design and implement observability best practices across logging, metrics, tracing, alerting, and dashboarding.
  • Support and improve incident management processes, participating in major incident response and post-incident reviews while driving root cause analysis and preventative actions.
  • Work with engineering teams to improve service resilience, availability, scalability, and recoverability through proactive engineering improvements.
  • Analyse reliability trends and operational data to identify systemic issues and recommend long-term solutions.
  • Contribute to the development of reliability standards, frameworks, and technical roadmaps across both legacy and modern technology environments.
  • Champion engineering excellence and reliability best practices through mentoring, knowledge sharing, and collaboration with technical teams.
  • Support technology transformation initiatives by ensuring operational resilience and reliability requirements are embedded throughout delivery programmes.

Apply for this job

Applications are handled by FDM Group directly. JobLot never asks candidates for payment.