Technical Service Operations Lead
You will lead the response to major production incidents, coordinate cross-functional investigations, make escalation decisions, and help ensure resolution within SLA targets. You will communicate incident updates to leadership, customers, and partners; manage status-page updates; and facilitate blameless post-incident reviews. You will analyze recurring issues, enforce incident-management processes, report operational metrics, govern service-management workflows, mentor Operations Engineers, and provide operational coverage during absences and surge incidents.
Responsibilities
- Serve as Incident Commander for major incidents, coordinating response teams, driving investigations, making escalation decisions, and ensuring resolution within SLA targets.
- Own incident communications, including updates to leadership, Customer Success, partners, customers, and the customer-facing status page.
- Facilitate blameless post-incident reviews, identify root causes, assign corrective actions, and track them to closure.
- Analyze incident trends, recurring issues, and production bugs; create Problem tickets; and report recommendations to product and engineering teams.
- Enforce the incident-management framework, including severity, priority, SLA, escalation, and deployment-readiness processes.
- Mentor Operations Engineers on triage, investigation, runbook execution, and documentation quality.
- Produce shift handoff reports and operational reports covering incident trends, KPIs, SLA adherence, detection rates, and repeat incidents.
- Audit service-catalogue completeness and govern JIRA Service Management workflows.
- Cover Operations Engineer duties during absences or surge incidents, including monitoring, triage, ticket creation, and runbook execution.
- Participate in weekend on-call rotation for major incidents.
Requirements
- 6+ years of experience in incident management, SRE, NOC leadership, or technical operations supporting high-availability, high-transaction production systems.
- Incident management experience coordinating multi-team responses, making escalation decisions, and communicating with executive stakeholders.
- Written and verbal English communication skills.
- ITIL knowledge of incident, problem, and change management lifecycles.
- Observability expertise, including logs, traces, metrics, APM, SLOs, error budgets, burn-rate alerting, and synthetic monitoring.
- Experience with Datadog or equivalent observability tools, PagerDuty or OpsGenie, JIRA or JIRA Service Management, Slack, and Confluence.
- Experience analyzing incident data and producing actionable recommendations.
- Experience with SLA/SLO-driven operations and MTTD, MTTA, and MTTR metrics.
- Experience with or interest in AI/ML-assisted operations.
- Ability to work in 24x7 shift-based operations and participate in rotating weekend on-call.