Site Reliability Engineering Specialist
|
About BT |
|
BT Group is the UK’s leading communications group and the holding company behind some of the country’s most recognised brands – including BT, EE, Openreach and Plusnet. Our purpose is as simple as it is ambitious: we connect for good. Our customers include consumers, small, medium and large businesses, public sector organisations and other communications providers. BT Group’s role is about setting direction, unlocking value and creating the conditions for our brands and businesses to thrive. Having come through the most capital-intensive phase of our fibre investment, our focus now is on what comes next – simplifying how we operate, using technology and AI to work smarter, and organising ourselves to serve customers better and grow sustainably. Group teams shape strategy, policy, brand, capital allocation and transformation, helping the whole organisation perform at its best. We have a singular culture that unites all our people: we are customer-first challengers, who are committed, clear and connected. These behaviours unite us as one team to deliver for our colleagues, our customers, our stakeholders and the country. Joining BT Group means working at the heart of a business that matters to the UK, with the opportunity to shape decisions, influence outcomes and help set the future course of one of the country’s most important companies. |
About the role
As a Site Reliability Engineer (SRE) within the Network Operations team, BTI International, you will be responsible for ensuring the reliability, resilience and performance of our Global Platforms including Global Fabric. You will collaborate closely with Engineering, Product and ASG teams to embed SRE principles such as automation, observability and proactive incident reduction into day‑to‑day operations. By improving how we monitor, maintain and evolve our services, you will help reduce risk, improve service quality and increase operational efficiency. Through this role, you will support BTI International’s strategy by enabling stable, secure and scalable platforms that support business growth, accelerate delivery of new capabilities, and protect customer experience.
What you will be doing (Role Accountabilities)
-
Provide end-to-end SRE ownership for the Global Fabric service, ensuring platform reliability, performance, resilience, and operational excellence.
-
Own incidents across the customer journey, from detection through to resolution, working with the appropriate ASGs and support teams.
-
Act as the primary operational escalation point for CF-related incidents participating in an on-call rota.
-
Escalate complex technical issues appropriately and coordinate resolution activities
-
Manage incidents through ServiceNow and track defects and improvements through Jira.
-
Perform root cause analysis and implement preventative actions to prevent recurrence along with providing documentation and knowledge transfer sessions.
-
Design, implement, operate, and continuously improve observability and monitoring solutions using Dynatrace.
-
Drive automation initiatives using Ansible, scripting, CI/CD pipelines, and GitOps practices to reduce operational toil and improve service reliability.
-
Define, measure, and report service health metrics, SLIs, SLOs, error budgets, and operational performance dashboards.
-
Fulfil operational service requests in line with agreed SLAs.
-
Support onboarding of new customers and operational readiness activities.
-
Analyse platform, network, and service trends to identify reliability and optimisation opportunities.
What you’ll need to succeed (Skills & Experience)
-
Experience supporting large-scale, high-availability services in an ISP / NaaS / network-centric environment.
-
Experience delivering changes through CI/CD and GitOps processes, including release validation, deployment governance, monitoring verification, and rollback planning.
-
Experience operating customer-facing applications, APIs, and distributed services with a strong focus on reliability, availability, performance, and customer experience.
-
Proven ability to troubleshoot end-to-end customer fulfilment journeys across UI, APIs, middleware, event platforms, and downstream systems.
-
Strong observability skills for CF services, including journey‑based monitoring and synthetic checks.
-
Knowledge of Infrastructure as Code tools like Terraform or Ansible.
-
Knowledge of event-driven architectures, including Kafka concepts such as message delivery, lag monitoring, loss detection, replay, and troubleshooting.
-
Working knowledge of incident/problem management in ServiceNow and delivery tracking in Jira (Scrum / PI planning).
BT Group’s Behaviours
Customer First: Prioritize customer needs in every decision and action.
Challengers: Challenge the status quo and bring innovative ideas to life.Committed: Own outcomes and deliver with integrity.
Clear: Communicate openly and simply, ensuring alignment.
Connected: Collaborate across teams to achieve shared goals.