freehire launches on Product Hunt on 26 August.

Follow →

Site Reliability Engineer (SRE)

Summary

Maintain and improve the reliability, scalability, and performance of a large-scale cloud SaaS platform using Kubernetes, cloud infrastructure, and automation.

We are looking for a Site Reliability Engineer (SRE) to join a global engineering team building and operating a large-scale cloud SaaS platform. In this role, you'll work closely with experienced SREs and software engineers to improve platform reliability, scalability, and operational excellence.
This position is a great fit for an engineer with experience in SRE, DevOps, or Cloud Operations who wants to deepen expertise in Kubernetes, cloud infrastructure, automation, and production engineering.

Responsibilities
Production Reliability – Help maintain the availability, stability, and performance of our production platform while proactively identifying opportunities to improve system reliability.
Monitoring & Observability – Improve monitoring, dashboards, and alert quality to increase production visibility and reduce alert fatigue.
Automation – Build scripts and automation solutions that reduce manual operational work, improve engineering efficiency, and enhance production reliability.
Production Operations & Incident Response – Participate in troubleshooting production issues, perform root cause analysis, and contribute to long-term reliability improvements.
Cloud Platform Operations – Support and improve our Kubernetes-based cloud platform, CI/CD pipelines, and production infrastructure running on GCP and AWS.
Engineering Collaboration – Work closely with Software Engineers, DevOps, DBAs, and Product teams to improve production reliability and operational excellence across the platform.
On-Call & Operational Excellence – Participate in the team's on-call rotation, supporting reliable production operations. Respond to production incidents alongside experienced SREs while continuously improving operational processes and reducing manual effort through automation.

Requirements
2–3 years of experience in Site Reliability Engineering, DevOps, Cloud Operations, Platform Engineering, or a similar role.
Hands-on experience with Kubernetes and containerized environments.
Familiarity with public cloud platforms (GCP or AWS).
Experience with Linux systems and basic networking concepts.
Experience with scripting or programming (Python, Bash, or similar).
Familiarity with monitoring and observability platforms (Datadog is an advantage).
Strong analytical and troubleshooting skills.
Excellent communication skills and the ability to collaborate with globally distributed engineering teams.
A strong desire to learn, take ownership, and continuously improve systems and processes.
A proactive mindset, strong work ethic, and the hunger to grow as a Site Reliability Engineer.

Nice to Have
Experience with CI/CD pipelines.
Familiarity with Infrastructure as Code tools such as Terraform or Ansible.
Experience with messaging technologies such as Kafka, Pub/Sub, or Redis.
Exposure to Canary, Blue/Green, or Feature Flag deployment strategies.
Understanding of Site Reliability Engineering principles, including SLIs, SLOs, and error budgets.
Experience working in cloud-native or SaaS production environments.