Site Reliability Engineering Manager
Summary
Leads and develops a team of site reliability engineers at LayerZero, owning reliability strategy for blockchain node infrastructure — SLOs, capacity planning, incident response, on-call, and postmortems — while driving infrastructure-as-code with Kubernetes and Helm. Requires 6+ years SRE/DevOps experience and 2+ years managing a technical team.
You will lead and develop a team of site reliability engineers, setting technical direction, growth plans, and performance expectations. You will own reliability strategy for blockchain node infrastructure, including service-level objectives, capacity planning, and incident response. You will improve infrastructure-as-code practices, on-call operations, automated detection and triage, and postmortem practices. You will also review designs, investigate critical incidents, and guide architecture decisions.
Responsibilities
- Lead and develop a team of SREs
- Set technical direction, growth plans, and performance expectations
- Own the reliability strategy for blockchain node infrastructure
- Define SLOs, capacity planning, and incident response
- Align reliability investments with business priorities
- Drive infrastructure-as-code practices using Kubernetes and Helm
- Improve on-call structure, incident detection and triage automation, and postmortem culture
- Review designs, investigate complex incidents, and set technical standards
Requirements
- Bachelor's degree in Computer Science, a similar technical field, or equivalent practical experience
- 6+ years of SRE, DevOps, or infrastructure engineering experience
- 2+ years of directly managing or leading a technical team
- Deep familiarity with blockchain node infrastructure, including validator, full, and archive nodes and RPC optimization
- Proficiency in TypeScript or Golang
- Advanced knowledge of Unix/Linux internals and distributed systems or high-availability design
- 3+ years running Kubernetes in production, including Helm chart authoring at scale
- Experience building or scaling an on-call and incident response process
- Excellent communication skills