Member of Technical Staff - GPU Infrastructure
Summary
Design and deploy large-scale GPU clusters for AI workloads, optimizing networking, filesystems, and performance while providing on-call support and documentation.
You will design, deploy, optimize, and support large-scale GPU infrastructure for customers. You will architect GPU clusters, implement orchestration and high-performance networking, configure parallel filesystems, tune system performance, resolve infrastructure issues across the stack, and provide operational documentation and support.
Responsibilities
- Partner with clients to understand workload requirements
- Design GPU cluster architectures
- Create technical proposals and capacity plans
- Develop deployment strategies for LLM training, inference, and HPC workloads
- Present architectural recommendations
- Deploy SLURM and Kubernetes
- Implement InfiniBand, RoCE, and NVLink networking
- Optimize GPU utilization and memory management
- Configure Lustre, BeeGFS, and GPFS filesystems
- Tune kernel and CUDA configurations
- Resolve customer infrastructure issues
- Implement monitoring, alerting, and automated remediation
- Provide on-call support
- Create runbooks and documentation
Requirements
- 3+ years of hands-on experience with GPU clusters and HPC environments
- SLURM and Kubernetes in production GPU settings
- InfiniBand configuration and troubleshooting
- NVIDIA GPU architecture, CUDA ecosystem, and driver stack
- Ansible and Terraform
- Python, Bash, and systems programming
- Customer-facing technical leadership
- NVIDIA drivers, Fabric Manager, and DCGM
- Docker, Containerd, and Enroot
- Linux kernel tuning
- AI workload network topology
- Power and cooling requirements for high-density GPU deployments
Benefits
- Equity incentives