Senior Site Reliability Engineer (AI Hardware & Infrastructure)
Summary
The Senior Site Reliability Engineer will ensure reliability, performance, and scalability of the company's global AI compute infrastructure—including bare‑metal servers, high‑density GPU racks, and advanced networking—by writing Python automation, designing BGP routing, and collaborating with data‑center technicians.
We are looking for an elite Site Reliability Engineer to join the core team responsible for our global AI compute infrastructure. Your mission will be to ensure that the physical and virtualized backbone of our AI platform—the bare-metal servers, high-density GPU racks, and the advanced network that connects them—is exceptionally reliable, performant, and scalable.
This is a role for a hands-on engineer who is as comfortable writing Python automation and Infrastructure-as-Code as they are designing BGP routing strategies and collaborating with data center technicians.
Who We're Looking For (Your Profile):
- You have a deep background in Site Reliability or Production Engineering, built on a solid Computer Science foundation and proven experience managing large-scale, mission-critical infrastructure.
- You are an exceptional Python programmer. You don't just write scripts; you build scalable, robust operational tools and automation frameworks from the ground up.
- You are a Networking expert. You have a strong, practical understanding of advanced network topologies, high-bandwidth routing and switching, BGP, and the complexities of dual-stack IPv4/IPv6 environments.
- You live and breathe Observability. You have hands-on, expert-level experience with modern monitoring stacks like Prometheus, Grafana, OpenTelemetry, and Loki.
- You are an operational leader. You have extensive experience designing service rollout strategies, defining meaningful alerting thresholds, creating clear technical runbooks, and leading incident response "war rooms."
- You are a natural owner. You thrive on solving ambiguous, complex technical problems and have a proven ability to take a challenge from a vague idea to a production-grade, fully automated solution.
- You are a strong collaborator, able to partner effectively with external data center vendors and coordinate with on-site field technicians to ensure maximum uptime.