Site Reliability Engineer

Summary

The Site Reliability Engineer will focus on building reliable, scalable systems through infrastructure automation, CI/CD pipeline management, and observability. The role involves collaborating with cross-functional teams to ensure high availability and performance using tools like Kubernetes, Terraform, and various monitoring platforms.

We are looking for a Site Reliability Engineer (SRE) with experience in platform engineering, DevOps, and production operations. The role involves building reliable systems, automating infrastructure, and ensuring observability across mission-critical applications. You will work closely with development, infrastructure, and operations teams to design scalable solutions, improve system resilience, and define operational best practices.

Key Responsibilities

  • System reliability: Ensure high availability, performance, and resilience of production systems.

  • Infrastructure automation: Build and maintain CI/CD pipelines, automate deployments, and manage containerized workloads.

  • Observability & monitoring: Implement logging, metrics, tracing, alerting, and dashboards using modern monitoring tools.

  • Incident response: Integrate systems with enterprise monitoring, SIEM, and incident management workflows; participate in on-call rotations.

  • Operational excellence: Define and implement runbooks, escalation paths, and production support models.

  • Collaboration: Work with cross-functional teams in Agile environments to deliver reliable and secure systems.

  • Continuous improvement: Conduct root cause analysis, implement permanent fixes, and drive automation initiatives to reduce manual effort.

Skills Required

  • 5+ years of experience as an SRE, DevOps Engineer, Platform Engineer, Infrastructure Engineer, or Production Engineer.

  • Solid knowledge of Linux administration, Shell scripting, Git, CI/CD, Docker, and Kubernetes.

  • Hands-on experience with observability platforms (Grafana, Prometheus, ELK, CloudWatch, Azure Monitor, etc.).

  • Experience integrating systems with enterprise monitoring, alerting, SIEM, or incident response workflows.

  • Proven ability to define and implement runbooks, operational procedures, escalation paths, and production support models.

  • Familiarity with cloud platforms (AWS, Azure, GCP) and infrastructure-as-code tools (Terraform, Ansible).

  • Excellent problem-solving skills, ability to troubleshoot complex systems, and experience in Agile/DevOps practices.

  • Excellent communication skills and ability to collaborate with cross-functional stakeholders.


    EA License : 02C3423

    EA Personnel : R22108699

See also

SRE jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available