Member of Technical Staff - Mid-Training Infra
Summary
Design, build, and operate large-scale GPU infrastructure for high-throughput inference and mid-training workloads at an AI lab. Day-to-day work spans distributed systems for synthetic data generation and reinforcement learning, optimizing model execution and GPU utilization (SGLang, Megatron, kernel optimization), and resolving performance bottlenecks across runtimes, networking, and distributed
You will design, build, and operate GPU infrastructure for high-throughput inference and mid-training workloads. You will develop distributed systems for synthetic data generation and reinforcement learning, optimize model execution and GPU utilization, support large-scale evaluations, and resolve performance bottlenecks across runtimes, kernels, networking, and distributed compute.
Responsibilities
- Design build and operate large-scale GPU infrastructure
- Develop systems for synthetic data generation and reinforcement learning pipelines
- Build high-performance inference platforms across thousands of GPUs
- Optimize inference throughput latency and GPU utilization
- Support distributed reinforcement learning and model evaluation workloads
- Improve model execution through kernel optimization model parallelism and GPU runtime improvements
- Diagnose and resolve performance bottlenecks across distributed systems
Requirements
- GPU infrastructure
- Model serving
- Inference
- GPU optimization
- SGLang
- Megatron
- Reinforcement learning
- Distributed system
- Synthetic data
- GPU kernel
- Networking
Benefits
- Stock options
- Medical insurance
- Dental insurance
- Vision insurance
- Life insurance
- Annual wellness allowance
- Daily office lunch and dinner
- 22 weeks of paid parental leave
- Unlimited paid time off in the U.S.
- 30 days of vacation in the U.K.
- Visa sponsorship
- Regular off-sites
- Happy hours
- Team celebrations
