Site Reliability Engineer

Summary

The Site Reliability Engineer will manage application infrastructure, CI/CD pipelines, and observability tools to ensure system reliability. The role involves capacity planning, incident response, and automating operational tasks using technologies like Kubernetes, Terraform, and GCP.

You will own application infrastructure and service integrations, build observability with SLOs, SLIs, monitors, alerts, and dashboards, and improve CI/CD automation. You will plan capacity, lead production readiness reviews, support incident response, reduce operational toil, maintain reliability roadmaps, and contribute to architecture and operational standards.

Responsibilities

  • Own Helm charts Terraform configurations Kubernetes deployments and runtime configuration
  • Design and implement SLOs SLIs monitors alerts and dashboards
  • Evolve CI/CD pipelines with deployment and rollback automation
  • Perform capacity planning performance tuning and load testing
  • Run Production Readiness Reviews
  • Investigate incidents and drive post-mortem improvements
  • Maintain runbooks and build operational automation
  • Drive the domain reliability roadmap
  • Participate in planning refinements and architecture reviews
  • Co-author company-wide SLO SLI capacity and operational standards
  • Participate in the SRE duty rotation

Requirements

  • 3+ years of SRE DevOps or platform engineering experience
  • On-call or incident response experience
  • Experience owning SLOs monitoring deployment pipelines and production infrastructure
  • Experience building and shipping backend services
  • Production-quality automation experience in Go PHP Python or Bash
  • Hands-on Kubernetes experience with Helm manifests deployment strategies and debugging
  • Observability experience with Datadog Prometheus Grafana or OpenTelemetry
  • Infrastructure as Code experience with Terraform or Terragrunt
  • GCP experience including IAM networking and managed services
  • Experience maintaining GitLab CI or GitHub Actions pipelines
  • Practical incident response and post-mortem experience
  • Strong collaboration and communication skills
  • Experience with payments fintech e-commerce or gaming systems
  • Kubernetes certifications are a plus
  • Google Cloud Platform certifications are a plus
  • HashiCorp certifications are a plus

Benefits

  • 100% company-paid medical plans
  • 100% company-paid dental plans
  • 100% company-paid vision plans
  • Unlimited flexible time off

See also

SRE jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available