Site Reliability Engineer
Summary
Site Reliability Engineer building and scaling the observability platform (Prometheus, Grafana, OpenTelemetry, tracing) on GCP and Kubernetes for software-defined networking that supports satellite constellations and ground station networks. Day to day: defining SLOs/SLIs and error budgets, automating with Terraform and GitOps (ArgoCD), and leading monitoring and incident response.
Salary: £130,000 - 150,000 per year
Requirements:- 4+ years of experience in SRE, reliability, or platform engineering, focused on observability for large-scale distributed systems
- Hands-on expertise building, scaling, and operating production observability stacks, and diagnosing complex performance and availability issues
- Strong production experience with GCP and Kubernetes
- Practical Infrastructure as Code and GitOps experience for configuration and deployment management
- Proficiency in Go or Python for automation and tooling
- Proven track record defining, implementing, and governing SLO, SLI, and error budget frameworks for high-availability services
- Design, build, and scale a unified observability platform using metrics, logging, and tracing tools such as Prometheus, Grafana, Loki, OpenTelemetry, and Tempo or Jaeger
- Define and manage end-to-end SLOs, SLIs, and error budgets to support production readiness and reliability
- Partner with engineering teams to embed standards, establish instrumentation best practices, and roll out consistent tooling
- Automate deployment, scaling, and lifecycle management using Terraform and GitOps tools such as ArgoCD
- Collaborate with infrastructure teams to deliver visibility across Kubernetes and multi-cloud environments
- Lead monitoring, alerting, and incident response; foster proactive reliability, blameless post-mortems, and continuous improvement
- ArgoCD
- Cloud
- GCP
- GitOps
- Grafana
- Support
- Kubernetes
- OpenTelemetry
- Prometheus
- Python
- Terraform
- DevOps
More:
We are a pioneering advanced technology organisation and a global leader in software-defined networking platforms for the aerospace sector. Our work supports critical connectivity infrastructure and next-generation space missions. In this remote Site Reliability Engineer role, you will design the visibility layer that powers reliability for satellite constellations, ground station networks, and deep-space communications systems. The salary is £130,000–£150,000. We offer the opportunity to contribute to cutting-edge projects.
last updated 39 week of 2026