SRE

Open 20d
  • Job Summary
  • The Site Reliability Engineer (SRE) is responsible for ensuring the availability, performance, scalability, and reliability of enterprise platforms and applications. The role focuses on monitoring, automation, incident management, and continuous improvement, working closely with engineering and DevOps teams to build resilient and highly available systems.
  • 2. Key Responsibilities
  • Ensure high availability and reliability of production systems and services
    Monitor system health using observability tools (logs, metrics, traces)
    Define and track SLIs, SLOs, and SLAs to measure system performance
    Lead/support incident management, root cause analysis (RCA), and post-incident reviews
    Automate operational tasks and implement Infrastructure as Code (IaC) practices
    Support and improve CI/CD pipelines for stable and efficient releases
    3. Skills & Competencies
  • Technical Skills
  • Cloud Platforms: Azure
    Monitoring & Observability: Prometheus, Grafana, Splunk, ELK, Datadog
    Containers & Orchestration: Docker, Kubernetes
    CI/CD Tools: Jenkins, GitHub Actions, GitLab CI, Azure DevOps
    Infrastructure as Code: Terraform, Ansible, CloudFormation
  • 2. Key Responsibilities
  • Ensure high availability and reliability of production systems and services
    Monitor system health using observability tools (logs, metrics, traces)
    Define and track SLIs, SLOs, and SLAs to measure system performance
    Lead/support incident management, root cause analysis (RCA), and post-incident reviews
    Automate operational tasks and implement Infrastructure as Code (IaC) practices
    Support and improve CI/CD pipelines for stable and efficient releases
    3. Skills & Competencies
  • Technical Skills
  • Cloud Platforms: AWS / Azure / GCP
    Monitoring & Observability: Prometheus, Grafana, Splunk, ELK, Datadog
    Containers & Orchestration: Docker, Kubernetes
    CI/CD Tools: Jenkins, GitHub Actions, GitLab CI, Azure DevOps
    Infrastructure as Code: Terraform, Ansible, CloudFormation