Escalation L3 Support Engineer
You will own complex escalated customer issues and platform incidents from investigation through resolution. You will troubleshoot GPU compute, networking, storage, drivers, CUDA, control plane, and billing issues; lead root-cause analysis and incident reviews; improve runbooks and monitoring; mentor L1 and L2 support; and participate in an on-call escalation rotation.
Responsibilities
- Own complex escalated customer issues and platform incidents through resolution
- Troubleshoot GPU compute, networking, storage, drivers, CUDA, control plane, and billing issues
- Bridge support with SRE, Compute, and R&D teams
- Drive root-cause analysis and permanent fixes
- Lead or support incident handling and post-incident reviews
- Improve runbooks, monitoring, the knowledge base, and escalation quality
- Mentor L1 and L2 support
- Participate in an on-call escalation rotation
Requirements
- 5+ years in cloud or infrastructure technical support, escalation, or SRE-adjacent roles
- Linux, networking, and cloud infrastructure experience
- GPU, CUDA, or HPC experience strongly preferred
- Experience resolving complex production issues and performing root-cause analysis
- Excellent written English and cross-team communication
- Comfort with on-call work and incident leadership
Benefits
- Welfare benefits