Backend Inference Framework Engineer Graduate (AML Inference) - 2027 Start
Summary
Design and optimize backend inference frameworks for large-scale machine learning models. This role focuses on developing distributed service architectures, engine scheduling, and performance tuning for high-concurrency environments to ensure stable and efficient model deployment.
- Design inference service architecture
- Develop inference engine scheduling
- Implement canary release
- Implement model inference services
- Implement monitoring and alerting
- Maintain distributed service architecture solutions
- Optimize distributed high concurrency service performance
- Optimize inference framework core modules
- Resolve performance bottlenecks
- Resolve resource bottlenecks
- Resolve stability issues
- Select and evaluate inference technologies