Technical Service Operations Lead
You will lead responses to major incidents, coordinate investigations and escalations, and keep leadership, customers, and partners informed. You will run blameless post-incident reviews, analyze incident trends, enforce incident-management processes, mentor Operations Engineers, manage operational reporting, and govern service-management workflows.
Responsibilities
- Serve as Incident Commander for major incidents, coordinating cross-functional response teams, investigations, escalations, and SLA resolution.
- Own incident communications for leadership, Customer Success, partners, customers, and the customer-facing status page.
- Facilitate blameless Post-Incident Reviews, identify root causes, assign corrective actions, and track them to closure.
- Analyze incident trends, recurring issues, and production bugs; create Problem tickets and report recommendations.
- Enforce the incident-management framework, including severity, priority, SLA, escalation, and deployment-readiness processes.
- Mentor Operations Engineers on triage, investigation, runbook execution, and documentation quality.
- Produce shift handoff reports and operational reports on incident trends, KPIs, SLA adherence, detection, and repeat incidents.
- Audit service-catalogue completeness and govern JIRA Service Management workflows.
- Cover Operations Engineer duties during absences, breaks, and surge incidents.
- Participate in rotating weekend on-call coverage for major incidents.
Requirements
- 6+ years of experience in incident management, SRE, NOC leadership, or technical operations supporting high-availability, high-transaction production systems.
- Gaming industry experience.
- Incident management experience coordinating multi-team responses, making real-time escalation decisions, and communicating with executive stakeholders.
- Written and verbal English communication skills.
- ITIL knowledge, including incident, problem, and change management lifecycles.
- Observability expertise, including logs, traces, metrics, APM, SLOs, error budgets, burn-rate alerting, and synthetic monitoring.
- Experience with Datadog or equivalent observability tools.
- Experience with PagerDuty or OpsGenie, JIRA or JIRA Service Management, Slack, and Confluence.
- Experience analyzing incident data and producing recommendations.
- Experience with SLA/SLO-driven operations and MTTD, MTTA, and MTTR metrics.
- Experience with or strong interest in AI/ML-assisted operations.
- Comfort with 24x7 shift-based operations and rotating weekend on-call coverage.
Benefits
- Medical coverage
- Dental coverage
- Vision coverage
- Paid time off