IT - Site Reliability Engineer
Summary: The main function of the Site Reliability Engineer (SRE) role is to ensure the reliability, availability, and performance of critical systems and services in a 24/7 operational environment. The SRE will leverage expertise in cloud platforms (primarily Microsoft Azure, with additional knowledge of AWS) to maintain seamless operations and respond to incidents promptly.
Responsibilities:
Monitor production systems and services using observability tools (logs, metrics, traces, dashboards).
Respond to incidents, alerts, and outages in real-time, ensuring rapid resolution and minimal impact.
Participate in a rotating on-call schedule, providing support during nights, weekends, and holidays.
Design, implement, and maintain observability solutions (e.g., Prometheus, Grafana, ELK and similar tools).
Develop and refine dashboards, alerts, and automated health checks for critical infrastructure and applications.
Analyze system performance and reliability data to identify trends and prevent future incidents.
Collaborate with development, infrastructure, application and security teams to ensure system reliability and scalability.
Automate operational tasks and incident response processes using scripting and configuration management tools.
Document procedures, runbooks, and incident reports for knowledge sharing and continuous improvement.
Conduct post-incident reviews and root cause analysis to drive improvements in reliability and response.
Propose and implement enhancements to monitoring, alerting, and operational processes.
Key Requirements:
Bachelor's degree in Information Technology, Computer Science, Business Administration, or related field.
2-5 years of experience in cloud engineering and operations engineering.
Experience with Azure services; AWS and GCP knowledge is a plus.
Hands-on experience with Infrastructure-as-Code (IaC) tools such as Terraform.
Strong scripting skills in Python, Bash, or PowerShell.
Familiarity with GitLab CI/CD tools.
Proficiency in monitoring and logging tools (e.g., native cloud tools, OpenMetrics, OpenTelemetry).
Nice to Have:
Master's degree or relevant certifications.
Experience in developing health checks for critical infrastructure.
Other Details:
Location: Pune
Team Structure: 24/7 shift rotation