Senior Production Engineer - Applied Machine Learning
This position is no longer accepting applications(closed Aug 17, 2026).
Summary
Build and maintain scalable ML infrastructure, including training pipelines, inference serving, and GPU Kubernetes clusters, while ensuring reliability and performance.
- Automate rollback and pre flight checks
- Build online inference serving systems
- Design disaster recovery plans
- Govern compute GPU CPU storage network resources
- Implement CI CD pipelines and canary releases
- Implement observability, monitoring, and alerting
- Maintain Kubernetes GPU cluster operations
- Manage production stability for machine learning training and inference pipelines
- Manage quotas cost attribution and performance tuning
- Orchestrate training scheduling and distributed training workflows
- Perform capacity forecasting and elastic auto scaling
- Run SLO SLA mechanisms and incident post mortems
- Set up on call fault diagnosis and auto healing