freehire launches on Product Hunt on 26 August.

Follow →

Site Reliability Engineer

You will own the reliability, scalability, availability, and performance of production deployment infrastructure. You will build and maintain observability and CI/CD systems, define SLOs and SLIs, scale Kubernetes workloads, participate in on-call rotations, lead incident learning, harden security, and automate operational work.

Responsibilities

  • Own and improve production-system reliability, availability, and performance
  • Design, implement, and maintain metrics, logging, and distributed tracing
  • Build and refine CI/CD pipelines
  • Conduct blameless post-mortems and eliminate recurring incident classes
  • Define SLOs and SLIs for new features with product engineering
  • Scale Kubernetes clusters supporting GPU workloads and AI inference
  • Automate recurring operational work
  • Participate in a 24/7 on-call rotation
  • Harden secrets management, network policies, and runtime security
  • Deploy agents for automated runbooks, anomaly detection, and incident triage

Requirements

  • 3 to 6 years of SRE, DevOps, or platform engineering experience
  • Expertise in Kubernetes and container orchestration at scale
  • Proficiency in Go, Rust, or Python
  • Experience with AWS, GCP, or bare-metal cloud-native infrastructure
  • Experience operating production systems at significant scale
  • Incident command experience

Benefits

  • Equity stake
  • Incident bonuses
  • Health benefits
  • Flexible PTO

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available