Principal Site Reliability Engineer (Python, Java and Automation Specialization)
Job Responsibilities
- Own and improve the reliability, scalability, performance, and operational efficiency of critical OCI Compute services.
- Lead investigation and resolution of complex production incidents; drive mitigation, recovery, RCA, and follow-up improvements.
- Improve service KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures.
- Build automation and tooling to reduce recurring operational toil and improve production safety.
- Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements.
- Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response.
- Support major upgrades, migrations, patching, capacity planning, performance tuning, security vulnerability management and production rollouts.
- Troubleshoot complex distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes.
- Contribute to incident-management practices, operational readiness, and service ownership improvements.
- Mentor team members and act as a technical resource for partner teams.
- Participate in a 12x7 on-call rotation and lead response during customer-impacting incidents.
Mandatory Skills
- Total Experience of 7+ Years
- 5+ years of experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role.
- Experience operating and improving highly available distributed systems in production.
- Strong programming or scripting skills in Python, Java, Go, or similar languages.
- Strong hands-on experience with Linux, cloud infrastructure, networking, compute, and storage.
- Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing.
- Experience defining or improving service SLIs, SLOs, KPIs, and operational procedures.
- Strong incident-management, troubleshooting, RCA, and problem-solving skills.
- Experience with deployment pipelines, release validation, automation, and change-management practices.
- Understanding of distributed systems, service dependencies, capacity planning, and performance tuning.
- Ability to work independently on complex technical issues and collaborate across engineering teams.
- Strong written and verbal communication skills.
Preferred Skills
- Experience with OCI and cloud infrastructure services.
- Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation.
- Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD.
- Experience with service migrations, fleet maintenance, upgrades, patching, or large-scale rollouts.
- Experience with architecture reviews, operational-readiness reviews, and post-incident improvements.
- Experience mentoring engineers or leading technical initiatives across teams.
- Familiarity with security, compliance, and access-control practices in production environments.
Self-Test Questions
- Does the candidate have 5+ years of SRE, Production Engineering, Cloud Operations, or Systems Engineering experience?
- Has the candidate owned or significantly improved a critical production service or infrastructure component?
- Can they lead a complex incident end-to-end: investigation, mitigation, recovery, RCA, and follow-up actions?
- Do they have strong hands-on experience with Linux, cloud infrastructure, distributed systems, networking, compute, or storage?
- Are they proficient in Python, Java, Go, or a similar language for automation, tooling, and troubleshooting?
- Have they built or improved automation, CI/CD pipelines, deployment validation, operational tooling, or remediation workflows?
- Do they have experience with monitoring, alerting, dashboards, logs, metrics, tracing, and SLO/KPI-driven reliability improvement?
- Can they work independently on ambiguous technical problems, collaborate across engineering teams, mentor peers, and participate in a 12x7 on-call rotation?
Role Details
|
Field
|
Requirement
|
|
Role
|
Site Reliability Engineer 4 (IC4)
|
|
Experience
|
5+ Years
|
|
Location
|
Bangalore Only
|
|
Work Mode
|
Hybrid - 3 days from Office
|
|
Primary Skills
|
Site Reliability Engineering, Production Engineering, Linux, OCI/Cloud Infrastructure, Distributed Systems, Python/Java/Go, Automation, CI/CD, Monitoring and Observability, Incident Management, RCA, SLOs/SLIs/KPIs, Performance Tuning, Capacity Planning, AIOps, Deployment Validation, On-Call Operations, Security Vulnerability Management, Security, Vulnerability
35% Ops, 65% - Automation and Coding
|
Career Level - IC4