Senior Site Reliability Engineer DevEx
You will design and build Kubernetes-based infrastructure primitives and automation for CI/CD platforms, build systems, and developer environments. You will develop Kubernetes Operators, scaling automation, GitOps workflows, and secure access capabilities that product teams can adopt directly. You will help build and operate the control plane behind GitHub Actions self-hosted runners, GitHub Apps, Tailscale, Flux, and ephemeral developer environments. Your work will improve the reliability, scalability, and efficiency of how engineers build, test, and deploy software.
Responsibilities
- Design and build infrastructure primitives for CI/CD platforms, build systems, and developer environments
- Build and operate the Kubernetes-based control plane behind the CI/CD platform
- Develop Kubernetes Operators and scaling automation for product-team adoption
- Build GitHub Actions self-hosted runner infrastructure with autoscaling, isolation, and cost and performance tuning
- Build GitHub Apps and GitHub-as-code automation, including permissions and webhooks
- Implement secure network access for CI/CD and remote development environments using Tailscale
- Deploy platform services through GitOps-driven workflows using Flux
- Build ephemeral on-demand developer environments and build systems
Requirements
- 6–9+ years of experience in SRE, platform, or infrastructure engineering
- Experience scaling Kubernetes in high-throughput production environments
- Deep Kubernetes expertise, including internals, scheduler behavior, custom resources, and cluster-scale failure diagnosis
- Experience building platform infrastructure, control planes, or Kubernetes Operators
- Distributed systems and production reliability experience
- Terraform and GitOps ownership
- Experience with GitOps workflows using Flux or ArgoCD
- Experience with CI/CD platforms at scale, including GitHub Actions self-hosted runners, workflows-as-code, GitHub Apps, and build systems
- AWS or cloud infrastructure production experience
- Proficiency in Go or another systems language
- Experience building infrastructure primitives with an automation-first mindset
Benefits
- Long-term incentives
- Comprehensive benefits