freehire launches on Product Hunt on 26 August.

Follow →

Senior Site Reliability Engineer

Summary

Maintain and improve Oracle Cloud Infrastructure’s compute services by automating operations, reducing incidents, and ensuring high availability and performance.

Job Responsibilities

  • Improve the reliability, scalability, performance, and operational efficiency of assigned OCI Compute services and components.
  • Investigate and resolve complex production incidents; contribute to mitigation, recovery, RCA, and follow-up actions.
  • Own and improve service-level KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures for assigned systems.
  • Build automation and tooling to reduce operational toil and improve production safety.
  • Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements.
  • Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response.
  • Support upgrades, migrations, patching, capacity planning, performance tuning, security vulnerability management and production rollouts.
  • Troubleshoot distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes.
  • Contribute to incident-management practices, operational readiness, and service ownership improvements.
  • Share technical knowledge and support team members through documentation, reviews, and collaboration.
  • Participate in a 12x7 on-call rotation and support response to customer-impacting incidents.

Mandatory Skills

  • 4–8 years of experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role.
  • Experience operating and improving highly available production systems.
  • Strong programming or scripting skills in Python, Java, Go, or similar languages.
  • Hands-on experience with Linux, cloud infrastructure, networking, compute, and storage.
  • Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing.
  • Experience owning or improving service SLIs, SLOs, KPIs, and operational procedures.
  • Strong incident troubleshooting, RCA, debugging, and problem-solving skills.
  • Experience with deployment pipelines, release validation, automation, and change-management practices.
  • Understanding of distributed systems, service dependencies, capacity planning, and performance tuning.
  • Ability to work independently on technical problems and collaborate effectively with engineering teams.
  • Strong written and verbal communication skills.

Preferred Skills

  • Experience with OCI and cloud infrastructure services.
  • Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation.
  • Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD.
  • Experience with service migrations, fleet maintenance, upgrades, patching, or production rollouts.
  • Experience with architecture reviews, operational-readiness reviews, and post-incident improvements.
  • Experience contributing to technical initiatives, knowledge sharing, code reviews, or operational improvements within the team.
  • Familiarity with security, compliance, and access-control practices in production environments.

Self-Test Questions for TA

  1. Does the candidate have 4–8 years of relevant SRE, Production Engineering, Cloud Operations, or Systems Engineering experience?
  2. Has the candidate independently operated or improved a production service, system, or infrastructure component?
  3. Can they investigate production incidents and contribute to mitigation, recovery, RCA, and follow-up actions?
  4. Do they have hands-on experience with Linux, cloud infrastructure, distributed systems, networking, compute, or storage?
  5. Are they proficient in Python, Java, Go, or a similar language for automation, tooling, and troubleshooting?
  6. Have they built or improved automation, deployment validation, CI/CD pipelines, or operational tooling?
  7. Do they have experience with monitoring, alerting, logs, metrics, tracing, and service health indicators such as SLOs or KPIs?
  8. Can they work independently on assigned technical problems, collaborate with partner teams, and participate in a 12x7 on-call rotation?

Career Level - IC3

See also