Senior Vice President, Site Reliability Engineer Manager

BNY is seeking an accomplished and strategic Senior Vice President – Site Reliability Engineer (SRE) to lead reliability engineering outcomes for critical technology platforms within Global Payment and Trade. This role is designed for a senior engineering leader who combines deep technical expertise with strong execution discipline, influencing the design, resilience, scalability, and operational maturity of large-scale, business-critical systems.

In this role, you’ll make an impact in the following ways:

  • Lead the strategic direction and execution of site reliability engineering practices across critical platforms, with a focus on resilience, scalability, availability, and operational excellence.
  • Drive the adoption and maturity of SRE principles, ensuring reliability is engineered into systems from design through production operations.
  • Define and champion enterprise-grade observability strategies, including monitoring, alerting, logging, tracing, event correlation, and actionable operational intelligence.
  • Establish, refine, and govern SLIs, SLOs, SLAs, and error budgets to create measurable and business-aligned service reliability objectives.
  • Lead resilience engineering initiatives, including chaos testing, failure injection, disaster recovery validation, and service hardening, to improve fault tolerance across platforms.
  • Oversee the identification and elimination of operational toil through automation, self-healing mechanisms, runbook optimization, and platform engineering practices.
  • Provide leadership during major production incidents, guiding incident response, root cause analysis, post-incident reviews, and long-term corrective actions to prevent recurrence.
  • Partner with engineering and architecture teams to influence reliability-focused design decisions, ensuring systems are scalable, supportable, and production-ready.
  • Drive capacity planning, performance engineering, and production readiness assessments for critical applications and services.
  • Evaluate, recommend, and implement modern tools, frameworks, and engineering practices that improve operational visibility, system health, and reliability outcomes.
  • Influence and contribute to engineering standards, reliability frameworks, governance practices, and operating models across teams and platforms.
  • Act as a senior technical leader and trusted advisor, providing thought leadership, mentorship, and technical direction to engineers and engineering leaders.
  • Build strong partnerships with cross-functional stakeholders to align reliability priorities with business objectives, risk management expectations, and client service outcomes.
  • Support a culture of continuous improvement, operational accountability, and data-driven decision-making across engineering and support functions.
  • Drive reliability transformation initiatives that improve MTTR, service availability, change success rate, alert quality, and platform recovery capabilities.

To be successful in this role, we’re seeking the following:

  • Significant experience in Site Reliability Engineering, Reliability Engineering, DevOps, Platform Engineering, or Production Engineering within complex enterprise environments.
  • Proven track record of leading large-scale reliability, resilience, and observability initiatives for mission-critical platforms.
  • Strong expertise in designing and implementing observability solutions using tools such as Splunk, Prometheus, Grafana, Dynatrace, AppDynamics, or similar platforms.
  • Deep hands-on experience in automation, scripting, and infrastructure as code, using technologies such as Python, Shell, Ansible, Terraform, or equivalent.
  • Strong experience with chaos engineering, resilience testing, failure scenario design, and service hardening practices.
  • Excellent troubleshooting and systems-thinking capability across distributed applications, middleware, infrastructure, cloud, and platform services.
  • Experience with cloud platforms such as AWS, Azure, or GCP, including cloud-native reliability practices.
  • Strong understanding of Linux/Unix systems, networking, distributed systems architecture, and modern enterprise application landscapes.
  • Demonstrated ability to lead technical problem-solving across organizational boundaries and influence outcomes at scale.
  • Strong communication, stakeholder engagement, and executive-level presentation skills.
  • Experience mentoring engineers and influencing technical direction without necessarily relying on direct line management authority.
  • Experience supporting or engineering payments platforms, transaction banking systems, or other high-volume, low-latency, highly available environments.
  • Knowledge of banking, financial services, operational risk, and regulatory expectations related to technology resilience and service continuity.
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical discipline, or equivalent practical experience.

Leadership Attributes

  • Brings an enterprise mindset, balancing deep technical expertise with strategic business alignment.
  • Takes full ownership of outcomes and drives execution through complexity and ambiguity.
  • Influences effectively across engineering, operations, architecture, risk, and senior leadership teams.
  • Demonstrates sound judgment under pressure, particularly during high-severity incidents and production events.
  • Champions continuous improvement, engineering discipline, and operational excellence.
  • Encourages innovation while maintaining strong focus on resilience, control, and sustainable engineering practices.

See also

SRE jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available