Senior Site Reliability Engineer
Summary
Maintain and improve Oracle Cloud Infrastructure’s compute services by automating operations, reducing incidents, and ensuring high availability and performance.
Job Responsibilities
- Improve the reliability, scalability, performance, and operational efficiency of assigned OCI Compute services and components.
- Investigate and resolve complex production incidents; contribute to mitigation, recovery, RCA, and follow-up actions.
- Own and improve service-level KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures for assigned systems.
- Build automation and tooling to reduce operational toil and improve production safety.
- Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements.
- Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response.
- Support upgrades, migrations, patching, capacity planning, performance tuning, security vulnerability management and production rollouts.
- Troubleshoot distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes.
- Contribute to incident-management practices, operational readiness, and service ownership improvements.
- Share technical knowledge and support team members through documentation, reviews, and collaboration.
- Participate in a 12x7 on-call rotation and support response to customer-impacting incidents.
Mandatory Skills
- 4–8 years of experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role.
- Experience operating and improving highly available production systems.
- Strong programming or scripting skills in Python, Java, Go, or similar languages.
- Hands-on experience with Linux, cloud infrastructure, networking, compute, and storage.
- Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing.
- Experience owning or improving service SLIs, SLOs, KPIs, and operational procedures.
- Strong incident troubleshooting, RCA, debugging, and problem-solving skills.
- Experience with deployment pipelines, release validation, automation, and change-management practices.
- Understanding of distributed systems, service dependencies, capacity planning, and performance tuning.
- Ability to work independently on technical problems and collaborate effectively with engineering teams.
- Strong written and verbal communication skills.
Preferred Skills
- Experience with OCI and cloud infrastructure services.
- Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation.
- Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD.
- Experience with service migrations, fleet maintenance, upgrades, patching, or production rollouts.
- Experience with architecture reviews, operational-readiness reviews, and post-incident improvements.
- Experience contributing to technical initiatives, knowledge sharing, code reviews, or operational improvements within the team.
- Familiarity with security, compliance, and access-control practices in production environments.
Self-Test Questions for TA
- Does the candidate have 4–8 years of relevant SRE, Production Engineering, Cloud Operations, or Systems Engineering experience?
- Has the candidate independently operated or improved a production service, system, or infrastructure component?
- Can they investigate production incidents and contribute to mitigation, recovery, RCA, and follow-up actions?
- Do they have hands-on experience with Linux, cloud infrastructure, distributed systems, networking, compute, or storage?
- Are they proficient in Python, Java, Go, or a similar language for automation, tooling, and troubleshooting?
- Have they built or improved automation, deployment validation, CI/CD pipelines, or operational tooling?
- Do they have experience with monitoring, alerting, logs, metrics, tracing, and service health indicators such as SLOs or KPIs?
- Can they work independently on assigned technical problems, collaborate with partner teams, and participate in a 12x7 on-call rotation?
Career Level - IC3