Member of Technical Staff, Distributed Training Systems
Posted Updated
Build and operate systems that support distributed training across heterogeneous compute, create reliable experiment infrastructure and evaluation harnesses, maintain traceable data and model pipelines, add observability, and turn fragile research prototypes into repeatable runs and trustworthy artifacts.
Responsibilities
- Build and operate distributed training for diffusion-heavy workloads across heterogeneous compute.
- Make experiments reliable with launchers, configurations, checkpointing, logging, metrics, run comparison, and reproducibility.
- Build benchmark and evaluation harnesses for physical AI research, including robotics and world-model experiments.
- Own data and model pipelines so results trace back to a dataset, version, and configuration.
- Add observability for GPU utilization, failure modes, data quality, routing behavior, model quality, and training stability.
- Turn fragile research prototypes into repeatable runs and trustworthy artifacts.
Requirements
- Hands-on experience with distributed training, GPU workloads, experiment infrastructure, or large-scale machine learning systems.
- Ability to debug performance, reliability, and reproducibility problems in complex training and evaluation workflows.
- Taste for simple tools that researchers will adopt.
- Clear communication.
- Strong ownership.
Benefits
- Competitive compensation.
- Meaningful equity.
- Deeply technical culture connecting research and systems work.
- Ownership of foundational infrastructure for frontier AI research.
- Paid travel to top machine learning and systems conferences around the world.