Staff Site Reliability Engineer - Release Engineering
Summary
Define and scale reliability practices for Plaid’s engineering teams, architecting SLOs, error budgets, and progressive delivery systems while leading incident response and improving release safety.
You will define and scale reliability practices across product engineering. You will architect SLO and error-budget programs, promote progressive delivery and automated safety gates, guide teams toward production readiness, build self-service deployment features, lead critical incident response, and improve release safety for high-velocity development.
Responsibilities
- Lead the expansion of reliability standards across product engineering
- Architect and manage SLO and error-budget frameworks
- Promote progressive delivery and automated safety gates
- Guide product teams toward production readiness
- Collaborate with Platform and Infrastructure teams on self-service platform features
- Direct critical incident response and post-mortem improvements
- Scale safety systems for increased code-change volume
Requirements
- Over 8 years of professional experience in backend systems, SRE, or platform engineering
- Experience designing reliability programs such as service maturity models or SLI frameworks
- Experience building or operating canary rollout systems, metric-gated analysis, or automated rollback infrastructure
- Technical proficiency in software development
- Ability to drive organizational change without formal authority
- Technical judgment in high-stakes production scenarios
- Exposure to Kubernetes, service mesh technologies, Prometheus, or ArgoCD
Benefits
- Equity