freehire launches on Product Hunt on 26 August.

Follow →

Principal Site Reliability Engineer (Python, Java and Automation Specialization)

Job Responsibilities

  • Own and improve the reliability, scalability, performance, and operational efficiency of critical OCI Compute services.
  • Lead investigation and resolution of complex production incidents; drive mitigation, recovery, RCA, and follow-up improvements.
  • Improve service KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures.
  • Build automation and tooling to reduce recurring operational toil and improve production safety.
  • Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements.
  • Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response.
  • Support major upgrades, migrations, patching, capacity planning, performance tuning, security vulnerability management and production rollouts.
  • Troubleshoot complex distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes.
  • Contribute to incident-management practices, operational readiness, and service ownership improvements.
  • Mentor team members and act as a technical resource for partner teams.
  • Participate in a 12x7 on-call rotation and lead response during customer-impacting incidents.

Mandatory Skills

  • Total Experience of 7+ Years
  • 5+ years of experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role.
  • Experience operating and improving highly available distributed systems in production.
  • Strong programming or scripting skills in Python, Java, Go, or similar languages.
  • Strong hands-on experience with Linux, cloud infrastructure, networking, compute, and storage.
  • Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing.
  • Experience defining or improving service SLIs, SLOs, KPIs, and operational procedures.
  • Strong incident-management, troubleshooting, RCA, and problem-solving skills.
  • Experience with deployment pipelines, release validation, automation, and change-management practices.
  • Understanding of distributed systems, service dependencies, capacity planning, and performance tuning.
  • Ability to work independently on complex technical issues and collaborate across engineering teams.
  • Strong written and verbal communication skills.

Preferred Skills

  • Experience with OCI and cloud infrastructure services.
  • Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation.
  • Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD.
  • Experience with service migrations, fleet maintenance, upgrades, patching, or large-scale rollouts.
  • Experience with architecture reviews, operational-readiness reviews, and post-incident improvements.
  • Experience mentoring engineers or leading technical initiatives across teams.
  • Familiarity with security, compliance, and access-control practices in production environments.

Self-Test Questions

  1. Does the candidate have 5+ years of SRE, Production Engineering, Cloud Operations, or Systems Engineering experience?
  2. Has the candidate owned or significantly improved a critical production service or infrastructure component?
  3. Can they lead a complex incident end-to-end: investigation, mitigation, recovery, RCA, and follow-up actions?
  4. Do they have strong hands-on experience with Linux, cloud infrastructure, distributed systems, networking, compute, or storage?
  5. Are they proficient in Python, Java, Go, or a similar language for automation, tooling, and troubleshooting?
  6. Have they built or improved automation, CI/CD pipelines, deployment validation, operational tooling, or remediation workflows?
  7. Do they have experience with monitoring, alerting, dashboards, logs, metrics, tracing, and SLO/KPI-driven reliability improvement?
  8. Can they work independently on ambiguous technical problems, collaborate across engineering teams, mentor peers, and participate in a 12x7 on-call rotation?

Role Details

Field
Requirement
Role
Site Reliability Engineer 4 (IC4)
Experience
5+ Years
Location
Bangalore Only
Work Mode
Hybrid - 3 days from Office
Primary Skills
Site Reliability Engineering, Production Engineering, Linux, OCI/Cloud Infrastructure, Distributed Systems, Python/Java/Go, Automation, CI/CD, Monitoring and Observability, Incident Management, RCA, SLOs/SLIs/KPIs, Performance Tuning, Capacity Planning, AIOps, Deployment Validation, On-Call Operations, Security Vulnerability Management, Security, Vulnerability

35% Ops, 65% - Automation and Coding


Career Level - IC4

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available