GPU Compute and Bare Metal DPU Engineer
You will own the complete lifecycle of bare-metal GPU nodes across multiple regions, from provisioning and delivery through operation, break-fix, and decommissioning. You will automate node-delivery pipelines, manage DPU, SmartNIC, and server firmware, improve fleet reliability, lead incident response and root-cause analysis, define GPU infrastructure standards, and participate in a multi-region on-call rotation.
Responsibilities
- Own bare-metal GPU node provisioning, delivery, operation, break-fix, and decommissioning
- Build and operate automated node-delivery pipelines
- Manage DPU, SmartNIC, BMC, BIOS, NIC, and GPU firmware
- Drive fleet reliability and reduce MTTR
- Lead incident response and root-cause analysis
- Improve hardware-health monitoring
- Build runbooks and tooling to reduce manual work
- Partner with Storage, Image, and Network teams on provisioning and handoff
- Define bring-up, rack, capacity, and acceptance standards for GPU SKUs and data-center regions
- Participate in a multi-region on-call rotation
Requirements
- 3+ years in large-scale bare-metal or server-fleet operations, HPC, or cloud infrastructure; 6+ years for Senior level
- Experience operating GPU servers at scale
- GPU driver, CUDA, and firmware management experience
- Strong Linux systems skills
- Experience with PXE, IPMI, Redfish, OS imaging, and automated provisioning
- Familiarity with DPU, SmartNIC, and bare-metal networking
- Ansible, Terraform, Python, or Go experience
- On-call, incident management, and operational runbook experience
Benefits
- Welfare benefits