The Senior Staff Site Reliability Engineer will lead the architecture and development of NVIDIA's enterprise AI runtime platform, focusing on Kubernetes-based systems for deploying and scaling AI applications and databases. The role involves building control-plane services, optimizing GPU scheduling, and establishing standards for high-performance AI inference infrastructure.
Sign in to see your match