Senior Site Reliability Engineer (AI/ML Platform)
Summary
The Senior Site Reliability Engineer will own reliability, observability, and automation for a large‑scale, GPU‑powered AI/ML platform, managing Kubernetes clusters, monitoring stacks, and infrastructure‑as‑code.
We are looking for a seasoned Site Reliability Engineer to join the team responsible for the backbone of our global AI/ML services. This isn't your typical SRE role. You won't just be maintaining systems; you'll be the guardian of a massive, distributed AI compute platform that processes workloads at an incredible scale. You will ensure that our AI models and GPU-powered infrastructure are not just fast, but fundamentally reliable, observable, and built to last.
If you are passionate about building and operating large-scale systems and are excited by the unique challenges of the AI/ML world, this is the role for you.
Who We're Looking For (Your Profile):
- You are a true Site Reliability, Platform, or Infrastructure Engineer at heart, with a proven track record of managing complex, large-scale distributed systems.
- Kubernetes is your natural habitat. You have deep, practical experience managing large-scale containerized environments and understand the complexities of orchestration under heavy load.
- You speak the language of observability fluently, with hands-on experience using tools like Prometheus, Grafana, and distributed tracing systems to make systems transparent and understandable.
- You are a strong programmer. You write clean, scalable automation scripts and infrastructure-as-code using Python or Go and tools like Terraform.
- You have a genuine curiosity or, ideally, direct experience with the unique challenges of AI/ML infrastructure, such as model serving pipelines, inference engines, or managing GPU-accelerated workloads.
- You are a problem-solver who takes full ownership of issues from start to finish. When you see a problem, you don't just fix it—you figure out how to prevent it from ever happening again.
- You excel at collaboration and enjoy mentoring other engineers, helping them adopt SRE principles and build more reliable software from the ground up.