Site Reliability Engineer
You will own the reliability, scalability, availability, and performance of production deployment infrastructure. You will build and maintain observability and CI/CD systems, define SLOs and SLIs, scale Kubernetes workloads, participate in on-call rotations, lead incident learning, harden security, and automate operational work.
Responsibilities
- Own and improve production-system reliability, availability, and performance
- Design, implement, and maintain metrics, logging, and distributed tracing
- Build and refine CI/CD pipelines
- Conduct blameless post-mortems and eliminate recurring incident classes
- Define SLOs and SLIs for new features with product engineering
- Scale Kubernetes clusters supporting GPU workloads and AI inference
- Automate recurring operational work
- Participate in a 24/7 on-call rotation
- Harden secrets management, network policies, and runtime security
- Deploy agents for automated runbooks, anomaly detection, and incident triage
Requirements
- 3 to 6 years of SRE, DevOps, or platform engineering experience
- Expertise in Kubernetes and container orchestration at scale
- Proficiency in Go, Rust, or Python
- Experience with AWS, GCP, or bare-metal cloud-native infrastructure
- Experience operating production systems at significant scale
- Incident command experience
Benefits
- Equity stake
- Incident bonuses
- Health benefits
- Flexible PTO