Machine Learning Platform Engineer
Summary
Build and maintain ML infrastructure, including training pipelines, model serving, and evaluation systems, to ensure scalable, low-latency, and cost-efficient AI deployments.
- Build and operate machine learning infrastructure
- Build platform tooling for experiment evaluation and model release
- Create evaluation and benchmarking infrastructure
- Design training evaluation deployment and inference systems
- Develop data pipelines for training and continuous improvement
- Identify ML stack bottlenecks and improve system performance
- Implement production observability monitoring tracing and alerting
- Improve reliability scalability latency throughput and cost efficiency
- Optimize model serving for high throughput and low latency