Point your AI agent at freehire and let it find you a job.

Get the CLI →

Link Group

NewBe an early applicant

Senior Site Reliability Engineer (AI/ML Platform)

Discussion

Summary

The Senior Site Reliability Engineer will own reliability, observability, and automation for a large‑scale, GPU‑powered AI/ML platform, managing Kubernetes clusters, monitoring stacks, and infrastructure‑as‑code.

We are looking for a seasoned Site Reliability Engineer to join the team responsible for the backbone of our global AI/ML services. This isn't your typical SRE role. You won't just be maintaining systems; you'll be the guardian of a massive, distributed AI compute platform that processes workloads at an incredible scale. You will ensure that our AI models and GPU-powered infrastructure are not just fast, but fundamentally reliable, observable, and built to last.

If you are passionate about building and operating large-scale systems and are excited by the unique challenges of the AI/ML world, this is the role for you.

Who We're Looking For (Your Profile):

  • You are a true Site Reliability, Platform, or Infrastructure Engineer at heart, with a proven track record of managing complex, large-scale distributed systems.
  • Kubernetes is your natural habitat. You have deep, practical experience managing large-scale containerized environments and understand the complexities of orchestration under heavy load.
  • You speak the language of observability fluently, with hands-on experience using tools like Prometheus, Grafana, and distributed tracing systems to make systems transparent and understandable.
  • You are a strong programmer. You write clean, scalable automation scripts and infrastructure-as-code using Python or Go and tools like Terraform.
  • You have a genuine curiosity or, ideally, direct experience with the unique challenges of AI/ML infrastructure, such as model serving pipelines, inference engines, or managing GPU-accelerated workloads.
  • You are a problem-solver who takes full ownership of issues from start to finish. When you see a problem, you don't just fix it—you figure out how to prevent it from ever happening again.
  • You excel at collaboration and enjoy mentoring other engineers, helping them adopt SRE principles and build more reliable software from the ground up.

See also

ML / AI jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available