Point your AI agent at freehire and let it find you a job.

Get the CLI →

Icertis

New

Lead Software Engineer, Cloud Site Reliability (SRE)

Posted 2 views
Discussion

Summary

Lead 24x7 cloud SRE/NOC operations for Icertis (Pune, India; remote-flagged): act as major incident manager for P1/P2 events, run Azure infrastructure and AKS clusters, and drive observability and automation. Core stack: Datadog, Azure Monitor, Terraform, Helm, ServiceNow, Power BI, and PowerShell/Python/Bash scripting.

Required Skills:

  • 7–12 years of experience in CloudOps / SRE / NOC environments (24x7 operations)

  • Strong expertise in Azure Infrastructure (VMs, Networking, Storage)

  • Hands-on experience with Azure Kubernetes Service (AKS), Kubernetes, Docker

  • Strong experience with monitoring and observability tools (Datadog, Azure Monitor)

  • Proven experience in Incident Management / Major Incident Handling, Monthly reporting

  • Experience with Infrastructure as Code (Terraform, ARM templates, Helm)

  • Scripting skills in PowerShell, Python, or Bash

  • Experience with ServiceNow (Incident, Problem, Change modules and dashboards)

  • Good understanding of distributed systems and cloud-native architecture

  • Excellent communication, leadership, and problem-solving skills


Role Responsibilities:

  • Lead 24x7 NOC operations with mandatory rotational shifts ensuring system availability and SLA adherence

  • Act as Major Incident Manager (P1/P2 incidents), driving triage, war room coordination, and stakeholder communication

  • Implement and enhance observability practices across logs, metrics, and traces

  • Work with tools like Datadog and Azure Monitor for monitoring and alerting

  • Drive proactive monitoring, alert tuning, anomaly detection, and AIOps initiatives

  • Manage Azure infrastructure and AKS clusters, including troubleshooting, scaling, and performance tuning

  • Build automation and self-healing workflows using Terraform, ARM, Helm, Power Automate, and scripting

  • Collaborate with engineering teams to improve reliability, deployment pipelines, and cloud-native architecture

  • Develop dashboards and reports using Power BI and ServiceNow

  • Handle Monthly Business reviews and leadership reporting

  • Mentor team members and drive process standardization and operational excellence

Preferred Certification:

  • Experience in multi-cloud environments (Azure/AWS)

  • Exposure to AIOps / predictive monitoring / self-healing systems

  • Azure / Datadog / Kubernetes certifications

Skills

What Lead SRE jobs ask for — and how much of it you have →
Apply

See also

SRE jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available