Senior AI Compute Infrastructure Engineer
Summary
Engineer and optimize GPU clusters for AI training and inference, build scheduling/orchestration systems, and improve performance, cost, and reliability of AI infrastructure.
You will operate GPU and accelerator clusters for AI training, inference, evaluation, and experimentation. You will build scheduling and utilization systems, optimize inference performance and cost, develop observability, lead reliability practices, integrate new hardware and runtimes, and support long-term AI infrastructure architecture.
Responsibilities
- Operate GPU and accelerator clusters
- Configure drivers, runtimes, kernels, and device plugins
- Build scheduling, orchestration, placement, and quota systems
- Optimize inference pipelines for latency, throughput, reliability, and cost
- Partner with ML engineers and researchers to remove infrastructure bottlenecks
- Build GPU observability and reporting
- Drive reliability, incident response, alerting, and runbooks
- Evaluate and integrate hardware, accelerators, runtimes, and serving frameworks
- Build tooling for visible and accountable GPU usage
- Contribute to AI infrastructure architecture decisions
Requirements
- 5+ years of infrastructure engineering experience
- GPU compute or ML infrastructure experience
- Distributed systems
- High-performance computing
- Production platform engineering
- GPU cluster operations
- Accelerator-backed infrastructure
- Scheduling and orchestration
- Utilization monitoring
- Cost optimization
- Linux
- Networking
- Storage
- Containers
- Kubernetes
- Distributed runtimes
- Production debugging
- vLLM
- Triton Inference Server
- TensorRT
- Python
- Performance optimization
- Observability
- Incident response