Lead Site Reliability Engineer - Infrastructure & DevOps
Summary
Lead a team to design, deploy, and maintain highly available cloud infrastructure for a generative AI platform using Kubernetes and DevOps best practices.
Job Title: Lead Site Reliability Engineer (SRE) Overview / Summary We are seeking a Lead Site Reliability Engineer to help drive the reliability, scalability, and operational excellence of a rapidly growing Generative AI platform. This role provides technical leadership while designing and supporting highly available cloud infrastructure powering modern AI and data-driven applications. The ideal candidate combines deep expertise in Site Reliability Engineering, cloud infrastructure, Kubernetes,…