SL2564 - SRE & Service Delivery Lead

Summary

Lead a team of SREs and Production Support Engineers, driving incident response, reliability, and operational improvements across hybrid cloud environments (AWS, Azure, OCI, on-prem, SaaS) for business-critical applications.

Role Overview

We are looking for an experienced SRE & Service Delivery Lead to lead Site Reliability Engineering and production support operations for business-critical applications.

You will be responsible for ensuring system reliability, availability, and performance while leading incident response, service improvements, and a team of SRE and Production Support Engineers.

Key Responsibilities

  • Lead SRE and production support activities across business-critical applications.
  • Manage and mentor a team of SREs and Production Support Engineers.
  • Lead major incident response, troubleshooting, root cause analysis, and service recovery.
  • Monitor system availability, performance, and overall production health against defined SLAs/SLOs.
  • Drive continuous improvements in reliability, operational efficiency, automation, and support processes.
  • Manage performance monitoring, system health checks, and disaster recovery readiness.
  • Develop and maintain operational runbooks, SOPs, and incident documentation.
  • Work closely with Engineering, Infrastructure, Operations, Delivery, and Incident Management teams.
  • Provide timely updates to management and stakeholders during major incidents.
  • Support hybrid environments across AWS, Azure, Oracle Cloud Infrastructure (OCI), on-premises infrastructure, and SaaS platforms.

Requirements

  • Minimum 5 years of experience in SRE, Service Delivery, Production Support, or a similar leadership role.
  • Minimum 5 years of experience working with cloud environments.
  • Strong experience in incident management, production support, system reliability, and service improvement.
  • Hands-on knowledge of observability and monitoring tools/practices.
  • Experience driving support optimization, automation, and operational improvements.
  • Technical understanding of application technologies such as Java, COBOL, and modern front-end technologies.
  • Strong leadership, stakeholder management, communication, and problem-solving skills.
  • Comfortable managing complex, business-critical production environments.
  • Financial Services experience is preferred, particularly within Insurance.
  • Bachelor's degree in Computer Science, IT, or a related discipline is preferred.

See also

SRE jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available