Site Reliability Engineer, Machine Learning Systems - Singapore
Summary
Designs and maintains monitoring, disaster recovery, and resource management tools for large-scale ML systems, ensuring stable training, inference, and offline task execution across global data centers and cloud environments.
- Build monitoring and management tools for ML infrastructure and services
- Ensure ML systems run efficiently for training inference and evaluation
- Lead disaster recovery and cluster machine governance
- Maintain offline tasks services stability across data centers regions and clouds
- Manage and plan resource allocation including cost and budget for computing and storage
- Provide global on call system and business support
Perks/Benefits:
- Global on call support