Technical Service Operations Lead
You will lead major-incident response, coordinate cross-functional investigations and escalation decisions, and ensure resolution within SLA targets. You will manage incident communications, status-page updates, and blameless post-incident reviews. You will analyze recurring production issues, enforce incident-management processes, mentor Operations Engineers, maintain service-management workflows, and participate in the weekend on-call rotation.
Responsibilities
- Serve as Incident Commander for major incidents, coordinating cross-functional response teams and ensuring resolution within SLA targets.
- Own incident communications and customer-facing status page updates.
- Facilitate blameless post-incident reviews, identify root causes, assign corrective actions, and track them to closure.
- Analyze incident trends, recurring issues, and production bugs; create Problem tickets and report recommendations.
- Enforce the incident management framework, including severity, priority, SLA, escalation, and deployment-readiness processes.
- Mentor Operations Engineers on triage, investigation, runbook execution, and documentation.
- Produce shift handoff reports and operational reporting on incident trends, KPIs, SLA adherence, detection rates, and repeat incidents.
- Audit service catalogue completeness and govern JIRA Service Management workflows.
- Cover Operations Engineer duties during absences, breaks, or surge incidents.
- Participate in the rotating weekend on-call schedule for major incidents.
Requirements
- Previous experience at a gaming company.
- 6+ years of experience in incident management, SRE, NOC leadership, or technical operations supporting high-availability, high-transaction production systems.
- Proven experience coordinating multi-team incident response, making real-time escalation decisions, and communicating with executive stakeholders.
- Excellent written and verbal English communication skills.
- Strong ITIL foundation and practical experience with incident, problem, and change management workflows.
- Observability expertise, including logs, traces, metrics, APM, SLOs, error budgets, burn-rate alerting, and synthetic monitoring.
- Hands-on experience with Datadog, PagerDuty or OpsGenie, JIRA or JIRA Service Management, Slack, and Confluence.
- Ability to analyze incident data and produce actionable recommendations.
- Experience with SLA/SLO-driven operations and MTTD, MTTA, and MTTR metrics.
- Ability to work in 24x7 shift-based operations and participate in rotating weekend on-call coverage.
Benefits
- Unlimited Flexible Time Off
- Gym membership
- Monthly train ticket