Member of Technical Staff, DevOps
- Own and evolve our CI/CD pipelines: dynamic pipeline generation across a monorepo of Go services, Python model containers, and Helm charts
- Operate and improve our GitOps deployment lifecycle: Helm releases, Kustomizations, and image automation across multiple clusters
- Build and maintain our observability stack: distributed tracing, metrics, dashboards, and alerting across all services and GPU workloads
- Define and track SLOs for core platform services, including session latency, model cold start time, and streaming reliability
- Run incident response: triage production issues, write postmortems, build runbooks, and drive reliability improvements
- Manage infrastructure-as-code across multiple cloud providers and regions: plan/apply workflows, state management, drift detection
- Operate secret management: encrypted secrets, external secret syncing, certificate automation
- Improve deployment safety: canary rollouts, health checks, startup probes, rollback automation
- Manage authentication infrastructure: OIDC federation for CI, workload identity for cloud services, cross-cloud credential management
- Participate in on-call rotation and build the tooling that makes on-call less painful
- You've run production Kubernetes clusters and been on-call for them. You've debugged node scheduling failures, OOM kills, and mysterious pod evictions at 3am
- Strong CI/CD experience: you've built and maintained pipelines for monorepos, not just single-service repos
- GitOps experience: you understand reconciliation loops, drift detection, and why image automation matters
- Infrastructure-as-code fluency with Terraform or similar across multiple environments and cloud accounts
- You know observability beyond just "set up dashboards". You've defined SLOs, built alerting that doesn't page on noise, and used traces to debug cross-service latency issues.
- Comfortable with secret management patterns (KMS, encrypted configs, external secret operators). You've thought about credential rotation and zero-trust.
- Incident response experience: you've triaged production outages, written postmortems that actually led to improvements, and built runbooks that other engineers could follow
- You write code, not just YAML. Proficiency in Go, Python, or Bash for building tooling, automation, and pipeline scripts
- Experience with GPU workloads on Kubernetes: device plugins, GPU-aware scheduling, GPU monitoring
- Multi-cloud operations beyond a single provider
- Real-time or streaming workloads: low-latency systems where p99 matters more than average
- Experience with Helm chart authoring and managing complex value layering across environments
- Familiarity with real-time media or relay infrastructure
- FinOps experience: GPU cost optimization, spot/preemptible instance management
- Engineers who treat infrastructure-as-code as "click around in the console and import later"
- SREs who've only monitored systems but never built the deployment pipelines that ship to them
- Candidates whose CI/CD experience is limited to GitHub Actions for a single-service repo
- People who write alerts that fire every day and then get ignored
We are based in-person in San Francisco. We are also hiring for this role in Europe for on-call coverage and timezone distribution.
- Competitive salary and meaningful early equity
- Visa sponsorship and relocation support
- Generous health, dental, and vision coverage