Cloud Operations & Support Engineer (L1 / L2)
You will serve as the front line for a public-cloud GPU platform, handling customer tickets while monitoring platform health and responding to alerts. You will triage and troubleshoot issues across GPU instances, virtual machines, bare metal, networking, CUDA, storage, billing, and quotas, escalate incidents, maintain runbooks, and report operational metrics in a rotating 24/7 schedule.
Responsibilities
- Handle customer support tickets and requests
- Triage, troubleshoot, and resolve customer issues within SLA targets
- Monitor platform health and respond to NOC alerts
- Assess incident severity and initiate incident handling
- Provide customer-facing status updates
- Troubleshoot GPU instances, virtual machines, bare metal, networking, drivers, CUDA, storage, billing, and quota issues
- Escalate issues with complete context
- Maintain runbooks, knowledge bases, and canned responses
- Participate in a 24/7 follow-the-sun shift rotation
- Track and report ticket and alert metrics
Requirements
- 1+ years of experience in cloud or technical support, NOC, or IT operations
- Strong new graduates considered for L1
- Linux, networking, and cloud fundamentals
- Familiarity with GPU, CUDA, containers, virtual machines, bare metal, or storage
- Strong written English and customer communication
- Willingness to work follow-the-sun shifts
Benefits
- Attractive welfare benefits