freehire launches on Product Hunt on 26 August.

Follow →

Principal Site Reliability Engineer (Python, Java and Automation Specialization)

Summary

Principal Site Reliability Engineer at Oracle responsible for improving reliability, scalability, and operational efficiency of critical OCI Compute services. The role involves incident resolution, automation development, observability, and participating in on-call rotations, with strong emphasis on Python, Java, and automation skills.

Job Responsibilities

  • Own and improve the reliability, scalability, performance, and operational efficiency of critical OCI Compute services.
  • Lead investigation and resolution of complex production incidents; drive mitigation, recovery, RCA, and follow-up improvements.
  • Improve service KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures.
  • Build automation and tooling to reduce recurring operational toil and improve production safety.
  • Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements.
  • Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response.
  • Support major upgrades, migrations, patching, capacity planning, performance tuning, security vulnerability management and production rollouts.
  • Troubleshoot complex distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes.
  • Contribute to incident-management practices, operational readiness, and service ownership improvements.
  • Mentor team members and act as a technical resource for partner teams.
  • Participate in a 12x7 on-call rotation and lead response during customer-impacting incidents.

Mandatory Skills

  • Total Experience of 7+ Years
  • 5+ years of experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role.
  • Experience operating and improving highly available distributed systems in production.
  • Strong programming or scripting skills in Python, Java, Go, or similar languages.
  • Strong hands-on experience with Linux, cloud infrastructure, networking, compute, and storage.
  • Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing.
  • Experience defining or improving service SLIs, SLOs, KPIs, and operational procedures.
  • Strong incident-management, troubleshooting, RCA, and problem-solving skills.
  • Experience with deployment pipelines, release validation, automation, and change-management practices.
  • Understanding of distributed systems, service dependencies, capacity planning, and performance tuning.
  • Ability to work independently on complex technical issues and collaborate across engineering teams.
  • Strong written and verbal communication skills.

Preferred Skills

  • Experience with OCI and cloud infrastructure services.
  • Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation.
  • Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD.
  • Experience with service migrations, fleet maintenance, upgrades, patching, or large-scale rollouts.
  • Experience with architecture reviews, operational-readiness reviews, and post-incident improvements.
  • Experience mentoring engineers or leading technical initiatives across teams.
  • Familiarity with security, compliance, and access-control practices in production environments.

Self-Test Questions

  1. Does the candidate have 5+ years of SRE, Production Engineering, Cloud Operations, or Systems Engineering experience?
  2. Has the candidate owned or significantly improved a critical production service or infrastructure component?
  3. Can they lead a complex incident end-to-end: investigation, mitigation, recovery, RCA, and follow-up actions?
  4. Do they have strong hands-on experience with Linux, cloud infrastructure, distributed systems, networking, compute, or storage?
  5. Are they proficient in Python, Java, Go, or a similar language for automation, tooling, and troubleshooting?
  6. Have they built or improved automation, CI/CD pipelines, deployment validation, operational tooling, or remediation workflows?
  7. Do they have experience with monitoring, alerting, dashboards, logs, metrics, tracing, and SLO/KPI-driven reliability improvement?
  8. Can they work independently on ambiguous technical problems, collaborate across engineering teams, mentor peers, and participate in a 12x7 on-call rotation?

Role Details

Field
Requirement
Role
Site Reliability Engineer 4 (IC4)
Experience
5+ Years
Location
Bangalore Only
Work Mode
Hybrid - 3 days from Office
Primary Skills
Site Reliability Engineering, Production Engineering, Linux, OCI/Cloud Infrastructure, Distributed Systems, Python/Java/Go, Automation, CI/CD, Monitoring and Observability, Incident Management, RCA, SLOs/SLIs/KPIs, Performance Tuning, Capacity Planning, AIOps, Deployment Validation, On-Call Operations, Security Vulnerability Management, Security, Vulnerability

35% Ops, 65% - Automation and Coding


Career Level - IC4

See also