SVP, Site Reliability Engineering Lead, SRE & Governance, Group Technology
Summary
Senior leader role at DBS Bank heading a 24/7 Site Reliability Engineering and infrastructure operations function across hybrid cloud and on-prem platforms (OpenShift/Kubernetes, hypervisors, Windows/Linux, databases, mainframe). Day to day involves leading shift-based SRE teams, driving SLA/SLO frameworks, observability, automation, incident governance, and regulatory compliance.
Role Summary
The SVP, Site Reliability Engineering (SRE), will lead and oversee the 24/7 infrastructure operations and reliability engineering function across critical platforms including Hypervisors (VPC, EPC, OPC), OpenShift , Windows, Databases, TWS, and Mainframe environments.
This role is responsible for driving resilience, scalability, automation, and operational excellence across hybrid cloud and on-premises environments, while ensuring alignment with business, risk, and regulatory expectations.
Key Responsibilities
Leadership & Governance
Lead and manage a distributed 24/7 SRE infrastructure team, including shift-based operations and command center functions
Define and execute the SRE strategy aligned to enterprise technology and business priorities
Establish strong governance across incident, problem, change, release, and capacity management
Drive SLA/SLO/SLI frameworks to ensure service reliability and performance targets
Infrastructure & Platform Ownership
Oversee end-to-end reliability of infrastructure platforms:
Cloud & Container: VPC, OpenShift, Kubernetes
Compute & Virtualization: Hypervisors (VMware/others), private cloud platforms
Enterprise Platforms: Windows, Unix/Linux, TWS, Mainframe, Databases
Ensure high availability, resilience, and disaster recovery readiness across all critical systems
Own infrastructure lifecycle including capacity planning, patching, upgrades, and decommissioning
Reliability Engineering & Automation
Champion SRE principles including error budgets, toil reduction, and automation-first mindset
Drive end-to-end observability strategy (monitoring, logging, tracing)
Lead initiatives to reduce MTTR, incident volume, and manual operational effort
Scale automation across deployment, patching, incident resolution, and self-healing capabilities
Operational Excellence
Ensure 24/7 monitoring, incident response, and recovery processes are robust and continuously improved
Lead major incident management and command bridge coordination for critical outages
Conduct RCA, trend analysis, and preventive engineering improvements
Embed ITIL best practices across service management processes
Risk, Compliance & Security
Identify infrastructure risks and drive proactive mitigation strategies
Ensure compliance with regulatory, audit, and internal security requirements
Partner with security teams on hardening, vulnerability management, and access controls
Stakeholder & Cross-Functional Collaboration
Collaborate with application, DevOps, security, architecture, and business teams to improve system reliability
Provide leadership in large-scale transformation programs (cloud adoption, infra modernization, SRE maturity)
Act as a key interface with senior management and external stakeholders
People & Talent Development
Build and develop a high-performing SRE organization across L1/L2/L3 layers
Drive fungibility, cross-skilling, and leadership development within the team
Mentor senior leaders and establish clear career progression frameworks
Requirements
Experience
18+ years of experience in IT infrastructure, SRE, or production operations
Proven leadership in managing large-scale 24/7 infrastructure teams in banking/financial services
Strong experience in hybrid cloud, data center, and enterprise platforms
Technical Expertise
Deep expertise in:
Cloud platforms (private/public cloud architectures)
Container platforms (OpenShift/Kubernetes)
Hypervisors & virtualization technologies
Operating systems (Windows, Linux/Unix)
Databases (MariaDB, Postgres, MSSQL, Redis, DB2)
Enterprise scheduling & legacy systems (TWS, Mainframe)
Strong understanding of DevOps, CI/CD, and infrastructure as code
Leadership & Functional Skills
Strong strategic thinking with ability to translate business goals into technology outcomes
Excellent incident leadership and crisis management skills
Proven track record of driving automation and operational transformation
Strong stakeholder management and executive communication skills
Other Skills
Expertise in ITIL / Service Management frameworks
Strong analytical, problem-solving, and decision-making capabilities
Ability to manage high-pressure situations and multiple priorities
Key Success Metrics (Optional for your slide/JD refinement)
Infrastructure availability (SLA/SLO adherence)
Reduction in MTTR / incident volume
Automation coverage & reduction in manual toil
Capacity utilization and cost optimization
Audit and compliance adherence
