Point your AI agent at freehire and let it find you a job.

Get the CLI →

The Business Connection Group

NewBe an early applicant

Site Reliability Engineer

Posted Updated
Discussion

Summary

Site Reliability Engineer building and scaling the observability platform (Prometheus, Grafana, OpenTelemetry, tracing) on GCP and Kubernetes for software-defined networking that supports satellite constellations and ground station networks. Day to day: defining SLOs/SLIs and error budgets, automating with Terraform and GitOps (ArgoCD), and leading monitoring and incident response.

Salary: £130,000 - 150,000 per year

Requirements:
  • 4+ years of experience in SRE, reliability, or platform engineering, focused on observability for large-scale distributed systems
  • Hands-on expertise building, scaling, and operating production observability stacks, and diagnosing complex performance and availability issues
  • Strong production experience with GCP and Kubernetes
  • Practical Infrastructure as Code and GitOps experience for configuration and deployment management
  • Proficiency in Go or Python for automation and tooling
  • Proven track record defining, implementing, and governing SLO, SLI, and error budget frameworks for high-availability services
Responsibilities:
  • Design, build, and scale a unified observability platform using metrics, logging, and tracing tools such as Prometheus, Grafana, Loki, OpenTelemetry, and Tempo or Jaeger
  • Define and manage end-to-end SLOs, SLIs, and error budgets to support production readiness and reliability
  • Partner with engineering teams to embed standards, establish instrumentation best practices, and roll out consistent tooling
  • Automate deployment, scaling, and lifecycle management using Terraform and GitOps tools such as ArgoCD
  • Collaborate with infrastructure teams to deliver visibility across Kubernetes and multi-cloud environments
  • Lead monitoring, alerting, and incident response; foster proactive reliability, blameless post-mortems, and continuous improvement
Technologies:
  • ArgoCD
  • Cloud
  • GCP
  • GitOps
  • Grafana
  • Support
  • Kubernetes
  • OpenTelemetry
  • Prometheus
  • Python
  • Terraform
  • DevOps

More:

We are a pioneering advanced technology organisation and a global leader in software-defined networking platforms for the aerospace sector. Our work supports critical connectivity infrastructure and next-generation space missions. In this remote Site Reliability Engineer role, you will design the visibility layer that powers reliability for satellite constellations, ground station networks, and deep-space communications systems. The salary is £130,000–£150,000. We offer the opportunity to contribute to cutting-edge projects.

last updated 39 week of 2026

Skills

Apply

See also

SRE jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available