Research Engineer - Distributed Training
You will build and optimize distributed training infrastructure for pre-training and large-scale reinforcement learning workloads. You will improve efficiency across compute, memory, networking, and scheduling, implement low-level optimizations, develop distributed training systems, and collaborate with researchers on frontier-scale model training.
Responsibilities
- Build and optimize distributed training infrastructure
- Improve training efficiency across compute, memory, networking, and scheduling layers
- Design and implement kernel, communication path, and runtime optimizations
- Develop distributed training systems for data, tensor, and pipeline parallel workloads
- Shape the architecture of the RL training stack
- Contribute to open-source libraries and internal infrastructure
- Translate system bottlenecks into concrete improvements
- Track advances in training systems, inference systems, compiler tooling, runtime tooling, and hardware-aware optimization
Requirements
- AI/ML infrastructure engineering experience
- Large-scale model training or inference experience
- PyTorch
- PyTorch Distributed
- DeepSpeed
- FSDP
- Megatron
- vLLM
- Ray
- Training performance optimization
- Data parallelism
- Tensor parallelism
- Pipeline parallelism
- GPU architecture
- Profiling
- Performance debugging
- CUDA
- Triton
- Compiler optimization
- Runtime optimization
- RL training infrastructure
- Multi-node GPU clusters
- High-performance networking
- Open-source contributions
Benefits
- Equity incentives
- Flexible work arrangements
- Remote or in-person work options
- Visa sponsorship
- Relocation assistance
- Quarterly team off-sites
- Hackathons
- Conferences
- Learning opportunities