IT - Site Reliability Engineer

Open 26d posting dated 2 weeks ago

Summary: The main function of the Site Reliability Engineer (SRE) role is to ensure the reliability, availability, and performance of critical systems and services in a 24/7 operational environment. The SRE will leverage expertise in cloud platforms (primarily Microsoft Azure, with additional knowledge of AWS) to maintain seamless operations and respond to incidents promptly.

Responsibilities:

  • Monitor production systems and services using observability tools (logs, metrics, traces, dashboards).

  • Respond to incidents, alerts, and outages in real-time, ensuring rapid resolution and minimal impact.

  • Participate in a rotating on-call schedule, providing support during nights, weekends, and holidays.

  • Design, implement, and maintain observability solutions (e.g., Prometheus, Grafana, ELK and similar tools).

  • Develop and refine dashboards, alerts, and automated health checks for critical infrastructure and applications.

  • Analyze system performance and reliability data to identify trends and prevent future incidents.

  • Collaborate with development, infrastructure, application and security teams to ensure system reliability and scalability.

  • Automate operational tasks and incident response processes using scripting and configuration management tools.

  • Document procedures, runbooks, and incident reports for knowledge sharing and continuous improvement.

  • Conduct post-incident reviews and root cause analysis to drive improvements in reliability and response.

  • Propose and implement enhancements to monitoring, alerting, and operational processes.

Key Requirements:

  • Bachelor's degree in Information Technology, Computer Science, Business Administration, or related field.

  • 2-5 years of experience in cloud engineering and operations engineering.

  • Experience with Azure services; AWS and GCP knowledge is a plus.

  • Hands-on experience with Infrastructure-as-Code (IaC) tools such as Terraform.

  • Strong scripting skills in Python, Bash, or PowerShell.

  • Familiarity with GitLab CI/CD tools.

  • Proficiency in monitoring and logging tools (e.g., native cloud tools, OpenMetrics, OpenTelemetry).

Nice to Have:

  • Master's degree or relevant certifications.

  • Experience in developing health checks for critical infrastructure.

Other Details:

  • Location: Pune

  • Team Structure: 24/7 shift rotation