AI Infrastructure Engineer
You will operate and optimize GPU clusters spanning multiple regions, working with Kubernetes, Slurm, and Ray to keep training and inference workloads running smoothly. You'll implement elastic scheduling and unified orchestration, balancing capacity between training and serving jobs. You will manage and tune high-throughput, low-latency inference runtimes, own model hot-swapping and zero-downtime rollout processes, and benchmark performance across a range of model sizes and workloads. You'll also tune distributed communication stacks and build observability tooling so the team can monitor GPU utilization, latency, and anomalies in real time.
Responsibilities
- Operate and optimize GPU clusters using Kubernetes, Slurm, and Ray across multiple regions
- Implement elastic scheduling and unified orchestration for inference and training jobs, including preemption and dynamic capacity arbitration
- Manage and tune vLLM / SGLang runtimes for high-throughput, low-latency serving
- Optimize distributed scheduling for multi-replica, multi-tenant serving
- Own model hot-swapping and zero-downtime rollout paths
- Benchmark and profile performance across workloads and model sizes
- Tune distributed communication stacks including NCCL / RCCL, RDMA over RoCEv2, and InfiniBand
- Build observability with Prometheus, Grafana, and Ray Dashboard to monitor GPU utilization and latency
- Integrate observability with the platform-wide OpenTelemetry + Grafana LGTM+ stack
Requirements
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field; PhD preferred for advanced R&D or innovation-oriented roles
- 3-5+ years in ML Infrastructure, HPC, or Systems Engineering
- Hands-on experience with Kubernetes, Slurm, or Ray
- Familiarity with vLLM, SGLang, or similar inference frameworks
- Strong background in PyTorch / JAX, distributed systems, and communication stacks (NCCL / RCCL, RDMA)
- Proficiency in Python plus one of Go / C++ / Rust
- Experience building observability with Prometheus and Grafana
- Fluent in English; experience working in multinational or cross-cultural environments is a plus
- Experience with major cloud platforms is strongly preferred
Benefits
- Attractive welfare benefits and developmental opportunities such as training and mentoring