freehire launches on Product Hunt on 26 August.

Follow →

Senior Production Engineer - Applied Machine Learning

Summary

Manage production systems for ML training, operate Kubernetes GPU clusters, and build CI/CD pipelines while ensuring reliability, disaster recovery, and observability.

- Automate inspections and pre flight checks - Build CI/CD pipelines - Conduct disaster recovery and incident reviews - Deploy canary releases - Diagnose faults and perform auto healing - Enable auto rollback - Ensure production stability - Forecast capacity - Handle scheduling and orchestration - Implement elastic auto scaling - Implement reliability engineering for SLO SLA - Maintain parameter server storage - Manage distributed training - Manage production systems for machine learning training - Manage resource governance for compute storage and network - Operate Kubernetes GPU clusters - Perform quota management and cost attribution - Run on call incident response - Serve online inference - Set up observability and alerting - Tune performance for availability Perks/Benefits: - 401k matching - Dental insurance - Disability coverage - Health insurance - Life insurance - Paid Holidays - Paid parental leave - Paid personal time - Paid sick days - Vision insurance - Wellbeing benefits

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available