freehire launches on Product Hunt on 26 August.

Follow →

Senior Site Reliability Engineer

Summary

Build and maintain scalable, self-healing infrastructure using Kubernetes, Terraform, and observability tools like Prometheus and Grafana, while automating incident response and reliability processes.

You will build reliable, observable, and self-healing infrastructure at scale. You will automate operational tasks and incident response, improve monitoring and alerting, collaborate on resilient system designs, participate in on-call rotations, lead post-incident reviews, and document operational processes and runbooks.

Responsibilities

  • Improve platform reliability and performance
  • Design, build, and maintain tools that automate operational tasks and incident response
  • Implement and improve monitoring, alerting, and tracing solutions
  • Collaborate on scalable and resilient system designs
  • Participate in on-call rotations
  • Lead post-incident reviews
  • Develop and document operational processes and runbooks
  • Contribute to SLO, SLI, and reliability-metric adoption

Requirements

  • Advanced knowledge of Linux/Unix systems in production environments
  • Experience with Kubernetes and container orchestration
  • Proficiency with Terraform and Ansible
  • Experience with Prometheus, Grafana, Loki, or ELK
  • Familiarity with Bash, Python, Go, or Ruby
  • Working knowledge of Git and CI/CD pipelines
  • Understanding of incident management and root cause analysis
  • Knowledge of cloud-native reliability and security best practices

Benefits

  • Paid Time Off
  • Wellhub
  • Annual bonus based on company and team performance
  • Flexible work hours

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available