AI Infrastructure Engineer
You will design AI frameworks for large-scale compute clusters and build distributed training systems for large models. You will optimize GPU utilization, memory use, training throughput, and data-loader efficiency. You will support multimodal large-model and control-model training. You will also build low-latency inference pipelines for real-time robot control using quantization, distillation, and model compilation.
Responsibilities
- Design and build AI frameworks for large-scale compute clusters
- Build and maintain distributed training systems for large-model training
- Optimize GPU utilization, memory consumption, and training throughput
- Optimize data-loader efficiency
- Support training optimization for multimodal large models and control models
- Build low-latency inference pipelines for real-time robot control
- Apply quantization, distillation, and model-compilation techniques to improve inference performance
Requirements
- Experience training on GPU clusters with thousands of GPUs
- Knowledge of distributed training mechanisms such as DDP and FSDP
- Experience with DeepSpeed and Megatron-LM
- Ability to define parallelization strategies based on model structures
- End-to-end performance analysis skills for compute, communication, and data-loading bottlenecks