Site Reliability Engineer - Database
Who We Are
Oracle is looking for an experienced Site Reliability Engineer to join our growing Cloud DevOps team supporting the Oracle Database Autonomous Recovery Service.
This is an opportunity to work on a mission-critical Oracle Cloud service that helps customers protect, recover, and secure their data at scale. You will help operate, automate, and improve the reliability of cloud infrastructure and database recovery services used by customers who depend on Oracle for resilient enterprise data protection.
You will work across cloud infrastructure, database technologies, automation, and operational reliability to solve complex problems, reduce manual effort, and prevent recurring issues. This role is ideal for someone with strong production operations experience, a passion for automation, and expertise in backup, restore, and recovery.
What You’ll Do
As a Site Reliability Engineer, you will help ensure Oracle’s Recovery Cloud Services operate reliably, securely, and efficiently. You will work with engineering and operations teams to improve service availability, automate complex workflows, and support the delivery of a high-quality cloud service.
- Support the operation and reliability of Oracle Database Autonomous Recovery Service.
- Build and improve automation for cloud infrastructure, service operations, and recovery workflows.
- Troubleshoot complex production issues across infrastructure, database, and cloud service components.
- Identify recurring operational problems and develop automation to prevent them.
- Contribute to the release, maintenance, and continuous improvement of cloud services.
- Work across teams to ensure service components integrate and operate seamlessly.
- Support production systems and database environments with a focus on reliability, scalability, and service quality.
- Participate in an out-of-hours on-call rota to support mission-critical cloud services.
What You’ll Bring
You bring strong problem-solving skills, sound judgment, and a continuous improvement mindset to help improve the reliability, scalability, and operational efficiency of production systems.
Minimum Qualifications
- Experience managing, supporting, or operating production cloud services, large-scale distributed systems, database environments, or mission-critical infrastructure.
- Experience releasing, maintaining, troubleshooting, or improving cloud services or production systems.
- Hands-on scripting or programming experience using one or more of the following: Python, Perl, Unix Shell, SQL, or PL/SQL.
- Experience improving operational processes, reducing recurring issues, and supporting service reliability.
- Experience in two or more of the following areas:
- Oracle Database
- Linux
- RMAN
Preferred Qualifications
- Experience in site reliability engineering, cloud operations, database systems, or production infrastructure support.
- Experience with Terraform or other Infrastructure as Code tools.
- Experience with Oracle Exadata.
- Experience with Zero Data Loss Recovery Appliance.
What We Offer
The opportunity to help operate and improve Oracle Recovery Cloud Services, supporting mission-critical database recovery for customers around the world.
The chance to work with Oracle engineering, operations, and database experts to deliver reliable, secure, and scalable cloud services.
A collaborative, high-performing environment where service reliability, operational excellence, practical problem-solving, and continuous improvement drive customer success.
Career growth through hands-on experience with production cloud services, Oracle Database technologies, site reliability engineering practices, automation, and large-scale infrastructure operations.
Competitive compensation, comprehensive benefits, flexible work arrangements, and the opportunity to build your career with one of the world’s leading cloud technology companies.
#LI-MP1
Career Level - IC3