Senior Vice President, Site Reliability Engineer Manager
BNY is seeking an accomplished and strategic Senior Vice President – Site Reliability Engineer (SRE) to lead reliability engineering outcomes for critical technology platforms within Global Payment and Trade. This role is designed for a senior engineering leader who combines deep technical expertise with strong execution discipline, influencing the design, resilience, scalability, and operational maturity of large-scale, business-critical systems.
In this role, you’ll make an impact in the following ways:
- Lead the strategic direction and execution of site reliability engineering practices across critical platforms, with a focus on resilience, scalability, availability, and operational excellence.
- Drive the adoption and maturity of SRE principles, ensuring reliability is engineered into systems from design through production operations.
- Define and champion enterprise-grade observability strategies, including monitoring, alerting, logging, tracing, event correlation, and actionable operational intelligence.
- Establish, refine, and govern SLIs, SLOs, SLAs, and error budgets to create measurable and business-aligned service reliability objectives.
- Lead resilience engineering initiatives, including chaos testing, failure injection, disaster recovery validation, and service hardening, to improve fault tolerance across platforms.
- Oversee the identification and elimination of operational toil through automation, self-healing mechanisms, runbook optimization, and platform engineering practices.
- Provide leadership during major production incidents, guiding incident response, root cause analysis, post-incident reviews, and long-term corrective actions to prevent recurrence.
- Partner with engineering and architecture teams to influence reliability-focused design decisions, ensuring systems are scalable, supportable, and production-ready.
- Drive capacity planning, performance engineering, and production readiness assessments for critical applications and services.
- Evaluate, recommend, and implement modern tools, frameworks, and engineering practices that improve operational visibility, system health, and reliability outcomes.
- Influence and contribute to engineering standards, reliability frameworks, governance practices, and operating models across teams and platforms.
- Act as a senior technical leader and trusted advisor, providing thought leadership, mentorship, and technical direction to engineers and engineering leaders.
- Build strong partnerships with cross-functional stakeholders to align reliability priorities with business objectives, risk management expectations, and client service outcomes.
- Support a culture of continuous improvement, operational accountability, and data-driven decision-making across engineering and support functions.
- Drive reliability transformation initiatives that improve MTTR, service availability, change success rate, alert quality, and platform recovery capabilities.
To be successful in this role, we’re seeking the following:
- Significant experience in Site Reliability Engineering, Reliability Engineering, DevOps, Platform Engineering, or Production Engineering within complex enterprise environments.
- Proven track record of leading large-scale reliability, resilience, and observability initiatives for mission-critical platforms.
- Strong expertise in designing and implementing observability solutions using tools such as Splunk, Prometheus, Grafana, Dynatrace, AppDynamics, or similar platforms.
- Deep hands-on experience in automation, scripting, and infrastructure as code, using technologies such as Python, Shell, Ansible, Terraform, or equivalent.
- Strong experience with chaos engineering, resilience testing, failure scenario design, and service hardening practices.
- Excellent troubleshooting and systems-thinking capability across distributed applications, middleware, infrastructure, cloud, and platform services.
- Experience with cloud platforms such as AWS, Azure, or GCP, including cloud-native reliability practices.
- Strong understanding of Linux/Unix systems, networking, distributed systems architecture, and modern enterprise application landscapes.
- Demonstrated ability to lead technical problem-solving across organizational boundaries and influence outcomes at scale.
- Strong communication, stakeholder engagement, and executive-level presentation skills.
- Experience mentoring engineers and influencing technical direction without necessarily relying on direct line management authority.
- Experience supporting or engineering payments platforms, transaction banking systems, or other high-volume, low-latency, highly available environments.
- Knowledge of banking, financial services, operational risk, and regulatory expectations related to technology resilience and service continuity.
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical discipline, or equivalent practical experience.
Leadership Attributes
- Brings an enterprise mindset, balancing deep technical expertise with strategic business alignment.
- Takes full ownership of outcomes and drives execution through complexity and ambiguity.
- Influences effectively across engineering, operations, architecture, risk, and senior leadership teams.
- Demonstrates sound judgment under pressure, particularly during high-severity incidents and production events.
- Champions continuous improvement, engineering discipline, and operational excellence.
- Encourages innovation while maintaining strong focus on resilience, control, and sustainable engineering practices.