Technical Service Operations Lead TSO Lead
You will lead responses to major incidents, coordinate cross-functional investigations, make escalation decisions, and help ensure resolution within SLA targets. You will manage stakeholder and customer communications, maintain status-page updates, facilitate blameless post-incident reviews, track corrective actions, analyze operational trends, and improve incident-management processes. You will also mentor Operations Engineers, produce operational reports, govern service-management workflows, provide operational coverage when needed, and participate in a rotating weekend on-call schedule.
Responsibilities
- Serve as Incident Commander for major incidents and coordinate cross-functional response teams.
- Drive incident investigations, escalation decisions, and resolution within SLA targets.
- Own incident communications for leadership, Customer Success, partners, and customers.
- Manage customer-facing status page updates.
- Facilitate blameless post-incident reviews and track corrective actions to closure.
- Analyze incident trends, recurring issues, and production bugs.
- Create Problem tickets and report recommendations to product and engineering teams.
- Enforce the incident management framework, including severity, priority, SLA, escalation, and deployment-readiness processes.
- Mentor Operations Engineers on triage, investigation, runbooks, and documentation.
- Produce shift handoff reports and operational KPI reporting.
- Audit service catalogue completeness and govern JIRA Service Management workflows.
- Cover Operations Engineer duties during absences or surge incidents.
- Participate in a rotating weekend on-call schedule for major incidents.
Requirements
- 6+ years of experience in incident management, SRE, NOC leadership, or technical operations supporting high-availability and high-transaction production systems.
- Incident management experience coordinating multi-team responses, making escalation decisions, and communicating with executive stakeholders.
- Excellent written and verbal English communication skills.
- Strong ITIL knowledge across incident, problem, and change management.
- Observability expertise, including logs, traces, metrics, APM, SLOs, error budgets, burn-rate alerting, and synthetic monitoring.
- Hands-on experience with Datadog, PagerDuty or OpsGenie, JIRA or JIRA Service Management, Slack, and Confluence.
- Ability to analyze incident data and provide actionable recommendations.
- Experience with SLA/SLO-driven operations and MTTD, MTTA, and MTTR metrics.
- Experience with or strong interest in AI/ML-assisted operations.
- Ability to work in 24x7 shift-based operations and participate in rotating weekend on-call.
Benefits
- Mac workstation and additional hardware.
- Health insurance covering medical, dental, and optical care for employees and dependants.
- Flexible working hours.
- No dress code.
- New office environment.