freehire launches on Product Hunt on 26 August.

Follow →

Senior Production Engineer - Applied Machine Learning

This position is no longer accepting applications(closed Aug 17, 2026).

Summary

Build and maintain scalable ML infrastructure, including training pipelines, inference serving, and GPU Kubernetes clusters, while ensuring reliability and performance.

- Automate rollback and pre flight checks - Build online inference serving systems - Design disaster recovery plans - Govern compute GPU CPU storage network resources - Implement CI CD pipelines and canary releases - Implement observability, monitoring, and alerting - Maintain Kubernetes GPU cluster operations - Manage production stability for machine learning training and inference pipelines - Manage quotas cost attribution and performance tuning - Orchestrate training scheduling and distributed training workflows - Perform capacity forecasting and elastic auto scaling - Run SLO SLA mechanisms and incident post mortems - Set up on call fault diagnosis and auto healing

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available