Senior AI Infrastructure Engineer Model Training
Posted Updated 2
views
You will build high-throughput sensor-data pipelines and distributed GPU training infrastructure. You will optimize parallelism, accelerator utilization, storage, networking, preprocessing, and compute; construct scalable datasets; profile bottlenecks; and scale model architectures to full-cluster training.
Responsibilities
- Design data-loading and streaming systems for multimodal sensor data
- Build and optimize distributed multi-node GPU training infrastructure
- Optimize accelerator utilization through mixed precision, kernel fusion, and memory optimization
- Profile training pipelines and eliminate bottlenecks
- Develop scalable driving-log dataset construction pipelines
- Scale new model architectures to full-cluster training runs
Requirements
- BS, MS, or PhD in computer science or a related field
- 2-3 years of industry experience in ML systems or infrastructure
- Experience with distributed training frameworks and techniques
- Experience building high-performance large-scale training data pipelines
- Understanding of GPU performance, profiling, and interconnects
- Strong Python and PyTorch skills
Benefits
- Equity
- Annual bonuses
- Medical, dental, and vision plans
- Infertility benefits
- Legal services
- Identity and fraud protection
- Hospital indemnity, accident, and critical illness insurance
- Flexible PTO
- 10 paid holidays
- Parental leave
- Dog-friendly office
- Free catered lunch
- Stocked kitchen
- Free EV charging
- Long-term and short-term disability insurance
- Life insurance
- Wellbeing programs
- 401(k)
- Commuter benefits
- FSA
- Dependent Care FSA
- HSA
- Referral and patent bonuses