Staff Software Engineer, HPC
Summary
Build and scale Zoox’s HPC platform using Ray.io, SLURM, and Kubernetes to support autonomous-vehicle AI workloads and developer workflows.
In this role, you will:
-
Design and implement core services and abstractions for distributed compute infrastructure supporting hundreds of thousands of concurrent jobs
-
Work with customer teams and other infrastructure teams to build a multiyear software engineering roadmap for the HPC platform
-
Lead multi-quarter, cross team initiatives that drive org-wide improvements
-
Create production-grade APIs, SDKs, and tools that make it easy for engineers across Zoox to run large-scale distributed workloads
-
Design and improve job scheduling algorithms and auto-scaling policies to maximize reliability and resource availability
-
Design multi-region orchestration strategies that optimize for data locality, reliability, and performance
-
Identify and resolve systemic reliability and performance issues through profiling, analysis, and collaboration with workload owners across multiple teams
-
Evaluate new technologies and paradigms that improve Zoox's computational and storage capabilities
-
Develop capacity planning tools and forecasting models to support Zoox's growing compute needs
-
Mentor junior engineers, guiding them through their career development
Qualifications
-
Experience designing and operating large-scale distributed systems in production
-
Experience with Ray.io, particularly Ray Core and Ray Data (or equivalent technologies)
-
Experience with Kubernetes, particularly for heterogeneous workloads
-
Experience with cloud infrastructure on AWS or similar providers
-
Track record of shipping and operating reliable, highly available scalable infrastructure
-
Demonstrated ability to prioritize development work and build cross-functional consensus around technical tradeoffs
-
Proficiency with Python
Bonus Qualifications
-
Exposure to machine learning workloads (training, inference, data generation)
-
Experience with Kubernetes or SLURM at scale (>10k+ nodes)
-
Experience with SLURM workload manager and advanced scheduling policies
-
Background in algorithmic optimization or operations research
-
Experience building developer tools and platforms used by large engineering organizations