Mainframe SRE
SRE Delivery Manager - Job Description
Job Summary
The SRE Delivery Manager is responsible for leading Service Reliability Engineering teams to ensure high availability, operational excellence, performance, scalability, and reliability of business-critical platforms and services. This role combines technical leadership, service delivery management, stakeholder engagement, incident management, automation strategy, and continuous service improvement to achieve defined SLAs, SLOs, and customer outcomes.
Service Reliability & Operations Leadership
Lead and manage global SRE, Operations, and Production Support teams.
Ensure platform availability, reliability, scalability, and performance.
Drive proactive monitoring and observability strategies.
Establish and track SLOs, SLIs, Error Budgets, and Reliability KPIs.
Improve service resilience through automation and engineering-driven operations.
Incident & Problem Management
Own Major Incident Management (P1/P2) processes.
Conduct post-incident reviews and root cause analysis (RCA).
Drive corrective and preventive actions (CAPA).
Reduce recurring incidents through problem management initiatives.
Lead crisis management and executive communications during critical outages.
Service Delivery Governance
Own operational service delivery commitments and customer satisfaction.
Ensure SLA compliance and continuous improvement.
Conduct regular operational and governance reviews.
Manage service risks, dependencies, and escalations.
Develop operational dashboards and executive reporting.
Automation & Continuous Improvement
Drive AIOps, self-healing, and automation initiatives.
Improve operational efficiency through scripting and workflow automation.
Reduce operational toil and manual interventions.
Define and execute operational excellence programs.
Promote DevOps and SRE best practices across teams.
Cloud & Infrastructure Reliability
Lead cloud operations across AWS, Azure, and GCP environments.
Ensure infrastructure stability, capacity management, and disaster recovery readiness.
Drive platform modernization and reliability engineering initiatives.
Support hybrid cloud and multi-cloud strategies.
Stakeholder & Customer Management
Act as primary escalation point for customers and leadership.
Partner with Engineering, Product Management, Security, Compliance, and Business teams.
Provide executive-level updates on service health and operational performance.
Build trusted relationships with internal and external stakeholders.
Financial & Resource Management
Manage operational budgets and resource planning.
Optimize support costs through automation and process improvements.
Forecast staffing requirements and capacity needs.
Drive productivity and utilization improvements.
Team Leadership
Lead, mentor, and develop SRE engineers and operations teams.
Define career growth plans and competency frameworks.
Foster a culture of ownership, accountability, reliability, and continuous learning.
Build high-performing global teams.
Required Skills & Qualifications
SRE Principles and Practices
DevOps Methodologies
Cloud Platforms (AWS, Azure, GCP)
Kubernetes and Containers
Linux/Unix Administration
Observability Tools: Datadog, Dynatrace, Splunk, Prometheus, Grafana, New Relic
CI/CD Pipelines
Infrastructure as Code (Terraform, Ansible)
Automation & Scripting (Python, PowerShell, Bash)
ITIL Framework and Service Governance
Executive Communication and Stakeholder Management
Key Metrics / KPIs
Service Availability (%)
SLA Achievement (%)
MTTR and MTTD
Incident Volume Reduction
Automation Coverage
Customer Satisfaction (CSAT)
Operational Efficiency Improvements
Error Budget Compliance
Platform Reliability Score
Preferred Qualifications
Bachelor’s degree in Computer Science, Engineering, or related field.
12+ years of IT Operations / Infrastructure experience.
5+ years leading SRE, DevOps, NOC, or Production Support teams.
Experience managing global teams and enterprise customers.
Preferred certifications: AWS/Azure, ITIL, Kubernetes, SRE Foundation.
Executive-Level Profile Summary
Experienced SRE Delivery Manager with expertise in IT Operations, Cloud Infrastructure, Managed Services, Reliability Engineering, and Digital Transformation. Proven track record of leading global teams, driving operational excellence, implementing automation and observability solutions, improving service reliability, reducing operational toil, and delivering exceptional customer outcomes through proactive governance and continuous improvement.
SRE Delivery Manager - Job Description
Job Summary
The SRE Delivery Manager is responsible for leading Service Reliability Engineering teams to ensure high availability, operational excellence, performance, scalability, and reliability of business-critical platforms and services. This role combines technical leadership, service delivery management, stakeholder engagement, incident management, automation strategy, and continuous service improvement to achieve defined SLAs, SLOs, and customer outcomes.
Service Reliability & Operations Leadership
Lead and manage global SRE, Operations, and Production Support teams.
Ensure platform availability, reliability, scalability, and performance.
Drive proactive monitoring and observability strategies.
Establish and track SLOs, SLIs, Error Budgets, and Reliability KPIs.
Improve service resilience through automation and engineering-driven operations.
Incident & Problem Management
Own Major Incident Management (P1/P2) processes.
Conduct post-incident reviews and root cause analysis (RCA).
Drive corrective and preventive actions (CAPA).
Reduce recurring incidents through problem management initiatives.
Lead crisis management and executive communications during critical outages.
Service Delivery Governance
Own operational service delivery commitments and customer satisfaction.
Ensure SLA compliance and continuous improvement.
Conduct regular operational and governance reviews.
Manage service risks, dependencies, and escalations.
Develop operational dashboards and executive reporting.
Automation & Continuous Improvement
Drive AIOps, self-healing, and automation initiatives.
Improve operational efficiency through scripting and workflow automation.
Reduce operational toil and manual interventions.
Define and execute operational excellence programs.
Promote DevOps and SRE best practices across teams.
Cloud & Infrastructure Reliability
Lead cloud operations across AWS, Azure, and GCP environments.
Ensure infrastructure stability, capacity management, and disaster recovery readiness.
Drive platform modernization and reliability engineering initiatives.
Support hybrid cloud and multi-cloud strategies.
Stakeholder & Customer Management
Act as primary escalation point for customers and leadership.
Partner with Engineering, Product Management, Security, Compliance, and Business teams.
Provide executive-level updates on service health and operational performance.
Build trusted relationships with internal and external stakeholders.
Financial & Resource Management
Manage operational budgets and resource planning.
Optimize support costs through automation and process improvements.
Forecast staffing requirements and capacity needs.
Drive productivity and utilization improvements.
Team Leadership
Lead, mentor, and develop SRE engineers and operations teams.
Define career growth plans and competency frameworks.
Foster a culture of ownership, accountability, reliability, and continuous learning.
Build high-performing global teams.
Required Skills & Qualifications
SRE Principles and Practices
DevOps Methodologies
Cloud Platforms (AWS, Azure, GCP)
Kubernetes and Containers
Linux/Unix Administration
Observability Tools: Datadog, Dynatrace, Splunk, Prometheus, Grafana, New Relic
CI/CD Pipelines
Infrastructure as Code (Terraform, Ansible)
Automation & Scripting (Python, PowerShell, Bash)
ITIL Framework and Service Governance
Executive Communication and Stakeholder Management
Key Metrics / KPIs
Service Availability (%)
SLA Achievement (%)
MTTR and MTTD
Incident Volume Reduction
Automation Coverage
Customer Satisfaction (CSAT)
Operational Efficiency Improvements
Error Budget Compliance
Platform Reliability Score
Preferred Qualifications
Bachelor’s degree in Computer Science, Engineering, or related field.
12+ years of IT Operations / Infrastructure experience.
5+ years leading SRE, DevOps, NOC, or Production Support teams.
Experience managing global teams and enterprise customers.
Preferred certifications: AWS/Azure, ITIL, Kubernetes, SRE Foundation.
Executive-Level Profile Summary
Experienced SRE Delivery Manager with expertise in IT Operations, Cloud Infrastructure, Managed Services, Reliability Engineering, and Digital Transformation. Proven track record of leading global teams, driving operational excellence, implementing automation and observability solutions, improving service reliability, reducing operational toil, and delivering exceptional customer outcomes through proactive governance and continuous improvement.
Required Skills & Qualifications
SRE Principles and Practices
DevOps Methodologies
Cloud Platforms (AWS, Azure, GCP)
Kubernetes and Containers
Linux/Unix Administration
Observability Tools: Datadog, Dynatrace, Splunk, Prometheus, Grafana, New Relic
CI/CD Pipelines
Infrastructure as Code (Terraform, Ansible)
Automation & Scripting (Python, PowerShell, Bash)
ITIL Framework and Service Governance
Executive Communication and Stakeholder Management
Preferred Qualifications
Bachelor’s degree in Computer Science, Engineering, or related field.
12+ years of IT Operations / Infrastructure experience.
5+ years leading SRE, DevOps, NOC, or Production Support teams.
Experience managing global teams and enterprise customers.
Preferred certifications: AWS/Azure, ITIL, Kubernetes, SRE Foundation.