freehire launches on Product Hunt on 26 August.

Follow →

Site Reliability Engineer, Machine Learning Systems - Singapore

Summary

Designs and maintains monitoring, disaster recovery, and resource management tools for large-scale ML systems, ensuring stable training, inference, and offline task execution across global data centers and cloud environments.

- Build monitoring and management tools for ML infrastructure and services - Ensure ML systems run efficiently for training inference and evaluation - Lead disaster recovery and cluster machine governance - Maintain offline tasks services stability across data centers regions and clouds - Manage and plan resource allocation including cost and budget for computing and storage - Provide global on call system and business support Perks/Benefits: - Global on call support

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available