Site Reliability Engineer
Summary
Maintain and scale AWS infrastructure, Kubernetes/ECS deployments, and CI/CD pipelines while co-owning production services and participating in on-call rotations.
You will co-own production services, help deliver and operate new and existing services, improve operations through metrics and analysis, develop performance benchmarks, maintain CI/CD tooling, participate in on-call rotations, and manage and scale infrastructure systems.
Responsibilities
- Co-own production services and ensure reliable and scalable operation
- Deliver new features and services and operate existing services
- Identify operational improvements through metric-driven collection and analysis
- Develop and maintain application performance benchmarks
- Improve operational efficiency through code releases and performance monitoring
- Maintain tooling, automation, monitoring, workflow management, and CI/CD
- Participate in the weekly on-call rotation
- Manage and scale infrastructure systems
- Improve CI/CD pipelines and AWS infrastructure
- Implement blue/green and canary deployments
Requirements
- Extensive AWS infrastructure deployment, management, and troubleshooting experience
- Production container deployment lifecycle experience using self-managed Kubernetes, ECS, or EKS
- CI/CD knowledge and custom production deployment tooling experience
- Distributed Linux systems troubleshooting
- Request tracing across applications, systems, and networks
- Automation experience
- Proficiency in at least two programming languages
- Strong written and spoken communication skills
Benefits
- Equity opportunity
- Maternity leave
- Paternity leave
- WeWork Membership
- WFH yearly stipend
- L&D stipend after 6 months