Senior Core Infrastructure Engineer
Summary
Build and maintain automation and observability tooling for Oracle Cloud’s GPU fleets, ensuring high availability and performance across regions.
We are seeking a highly skilled and motivated Developer to join the AI2 Ops team supporting GPU infrastructure in Oracle Cloud Infrastructure (OCI). In this role, you will design, develop, deploy, and maintain automation and operational tooling for GPU fleets across multiple regions, ensuring high availability, scalability, performance, and efficient capacity utilization. You will collaborate closely with engineering, product, and operations teams to build robust automation, observability, and reliability solutions while continuously improving GPU operations.
- Design, build, and maintain software, automation, and operational tooling for OCI GPU infrastructure across multiple geographic regions.
- Automate GPU infrastructure provisioning, configuration, validation, and deployment using Python, Bash, Terraform, and related tooling.
- Collaborate with software engineers, hardware teams, and operations partners to build scalable, reliable, and highly available GPU platform services.
- Build and improve monitoring, alerting, and diagnostics for GPU fleet health, performance, capacity, and utilization using tools such as Grafana.
- Participate in incident response and root-cause analysis to remove blockers affecting GPU capacity, availability, and regional deployments.
- Continuously improve AI2 Ops processes, GPU fleet automation, and OCI region build readiness.
- Participate in on-call rotations and provide support for critical infrastructure issues.
- Document operational procedures, automation workflows, troubleshooting guides, and runbooks.
- Collaborate with cross-functional teams to design and roll out new GPU capacity, cloud region builds, and expansions.
Required Qualifications:
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 3+ years of software development or infrastructure automation experience with strong proficiency in Python and Bash.
- Strong Linux systems experience, including troubleshooting, scripting, process management, networking fundamentals, and system administration.
- Hands-on experience with infrastructure-as-code and automation tools, especially Terraform, and with RESTful APIs.
- Strong problem-solving and troubleshooting skills.
- Excellent communication and teamwork skills.
Preferred Skills:
- Experience participating in or leading on-call operations and incident response.
- Familiarity with DevOps practices and continuous integration/continuous deployment (CI/CD).
- Experience operating or automating GPU, compute, or other large-scale cloud infrastructure.
- Exposure to containerization and orchestration technologies such as Docker and Kubernetes.
- Experience with observability tooling, including metrics, logging, dashboards, and alerting.
- Experience with CI/CD and Agile methodologies, especially Scrum.
Career Level - IC3