Site Reliability Engineer Intermediate to Senior Staff Infrastructure Platforms
You will keep user-facing services and production systems reliable, scalable, and efficient. You will build infrastructure automation, operate and troubleshoot Kubernetes systems, manage infrastructure as code, and safely deliver changes through CI/CD and GitOps. You will participate in on-call rotations, triage alerts, improve runbooks, support incident response and post-incident reviews, and document architecture decisions and repeatable practices.
Responsibilities
- Keep user-facing services and production systems reliable, scalable, and efficient
- Build automation and tooling that reduces toil and replaces manual work with infrastructure-as-code-driven workflows
- Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling
- Write and maintain infrastructure as code and ship changes through CI/CD and GitOps
- Participate in on-call rotations, triage alerts, improve runbooks, and escalate appropriately
- Contribute to observability using metrics, logs, and SLOs
- Participate in incident response and post-incident reviews
- Document runbooks, architecture decisions, and reviews
Requirements
- Experience maintaining reliable production systems using operations and software engineering practices
- Experience building new infrastructure tooling and automation
- Ability to read, debug, and reason about code behavior, performance, and failure modes
- Experience with infrastructure as code and Kubernetes
- Hands-on experience with GCP or AWS
- Familiarity with metrics, logging, alerting, SLOs, and SLIs
- Experience participating in on-call and incident response
- Written communication skills for an asynchronous, distributed environment
- Experience using automation and AI to reduce toil
Benefits
- Health, financial, and well-being benefits
- Flexible Paid Time Off
- Team Member Resource Groups
- Equity compensation and Employee Stock Purchase Plan
- Parental Leave
