Senior Production Engineer Managed Cloud
You will design and operate reliable managed cloud services at scale. You will define SLIs and SLOs, improve observability and performance, investigate distributed-system reliability issues, optimize AI infrastructure, and contribute to fault-tolerant systems architecture.
Responsibilities
- Design and operate reliable managed cloud services
- Define and improve SLIs and SLOs
- Collaborate with AI, platform, and infrastructure teams
- Optimize training and inference clusters
- Build telemetry and performance-tuning strategies
- Investigate and resolve reliability issues
- Contribute to distributed systems architecture
Requirements
- Strong software engineering background
- Experience building production-grade systems beyond scripting or Bash
- Experience designing and implementing distributed systems
- Experience defining and measuring SLIs and SLOs
- Experience building monitoring and observability systems
- Experience driving performance and reliability improvements
- Experience designing fault-tolerant systems and automated testing strategies
- Proficiency in Python, Go, Java, or C++
- Familiarity with Kubernetes or container orchestration platforms
Benefits
- Industry competitive pay
- Restricted Stock Units
- Health insurance options including HDHP and PPO
- Vision insurance
- Dental insurance
- Employer HSA contributions
- Paid parental leave
- Paid life insurance
- Short-term and long-term disability insurance
- Teladoc
- 401(k) with 100% match up to 4% of salary
- Paid time off
- Paid holidays
- Cell phone reimbursement
- Tuition reimbursement
- Calm app subscription
- MetLife Legal
- Company-paid commuter benefit of $300 per month