Senior Production Engineer - Applied Machine Learning
Summary
Manage production systems for ML training, operate Kubernetes GPU clusters, and build CI/CD pipelines while ensuring reliability, disaster recovery, and observability.
- Automate inspections and pre flight checks
- Build CI/CD pipelines
- Conduct disaster recovery and incident reviews
- Deploy canary releases
- Diagnose faults and perform auto healing
- Enable auto rollback
- Ensure production stability
- Forecast capacity
- Handle scheduling and orchestration
- Implement elastic auto scaling
- Implement reliability engineering for SLO SLA
- Maintain parameter server storage
- Manage distributed training
- Manage production systems for machine learning training
- Manage resource governance for compute storage and network
- Operate Kubernetes GPU clusters
- Perform quota management and cost attribution
- Run on call incident response
- Serve online inference
- Set up observability and alerting
- Tune performance for availability
Perks/Benefits:
- 401k matching
- Dental insurance
- Disability coverage
- Health insurance
- Life insurance
- Paid Holidays
- Paid parental leave
- Paid personal time
- Paid sick days
- Vision insurance
- Wellbeing benefits