SRE
Posted Updated 3
views
Summary
The Site Reliability Engineer ensures the availability, performance, and scalability of enterprise platforms through monitoring, automation, and incident management. The role involves working with cloud infrastructure, CI/CD pipelines, and Infrastructure as Code tools to maintain resilient systems.
- Job Summary
- The Site Reliability Engineer (SRE) is responsible for ensuring the availability, performance, scalability, and reliability of enterprise platforms and applications. The role focuses on monitoring, automation, incident management, and continuous improvement, working closely with engineering and DevOps teams to build resilient and highly available systems.
- 2. Key Responsibilities
- Ensure high availability and reliability of production systems and services
Monitor system health using observability tools (logs, metrics, traces)
Define and track SLIs, SLOs, and SLAs to measure system performance
Lead/support incident management, root cause analysis (RCA), and post-incident reviews
Automate operational tasks and implement Infrastructure as Code (IaC) practices
Support and improve CI/CD pipelines for stable and efficient releases
3. Skills & Competencies - Technical Skills
- Cloud Platforms: Azure
Monitoring & Observability: Prometheus, Grafana, Splunk, ELK, Datadog
Containers & Orchestration: Docker, Kubernetes
CI/CD Tools: Jenkins, GitHub Actions, GitLab CI, Azure DevOps
Infrastructure as Code: Terraform, Ansible, CloudFormation
- 2. Key Responsibilities
- Ensure high availability and reliability of production systems and services
Monitor system health using observability tools (logs, metrics, traces)
Define and track SLIs, SLOs, and SLAs to measure system performance
Lead/support incident management, root cause analysis (RCA), and post-incident reviews
Automate operational tasks and implement Infrastructure as Code (IaC) practices
Support and improve CI/CD pipelines for stable and efficient releases
3. Skills & Competencies
- Technical Skills
- Cloud Platforms: AWS / Azure / GCP
Monitoring & Observability: Prometheus, Grafana, Splunk, ELK, Datadog
Containers & Orchestration: Docker, Kubernetes
CI/CD Tools: Jenkins, GitHub Actions, GitLab CI, Azure DevOps
Infrastructure as Code: Terraform, Ansible, CloudFormation