Member of Technical Staff Compute Platform
Summary
Builds and maintains the compute platform for large multi-GPU fleets at an AI company: tooling for automated remediation, topology-aware scheduling, capacity planning, and hardware debugging, plus cluster-wide monitoring, benchmarking, storage replication, and GPU networking. Core stack: Kubernetes, NCCL, GPU/cloud infrastructure.
You will build and maintain tooling for automated remediation, topology-aware scheduling, capacity planning, and hardware debugging. You will improve cluster management for large GPU fleets, implement cluster-wide monitoring and benchmarking, and prepare infrastructure for larger GPU deployments, storage replication, and network performance.
Responsibilities
- Build and maintain tools for automatic remediation, topology-aware scheduling, capacity planning, and hardware debugging
- Design and improve the cluster management stack for large multi-GPU fleets
- Implement cluster-wide monitoring and active performance benchmarking
- Prepare infrastructure for next-generation GPU deployments and larger clusters
- Develop multi-cloud storage, data replication, and GPU network capabilities
- Co-design fault tolerance, node health checks, and remediation strategies
Requirements
- Systems engineering experience focused on cluster-wide behavior and maintenance
- Strong coding ability in systems or GPU infrastructure
- Deep GPU hardware knowledge
- NCCL knowledge
- Kubernetes architecture experience
- Cloud storage expertise across data centers
- Experience handling datasets and checkpointing at scale
Benefits
- Stock options
- Medical, dental, vision, and life insurance
- Annual wellness allowance
- Daily in-office lunch and dinner
- 22 weeks of paid parental leave
- Unlimited paid time off in the U.S.
- 30 days of vacation in the U.K.
- Visa sponsorship support
- Regular off-sites, happy hours, and team celebrations