Own a multi-GPU inference platform end to end: the operating system, storage, networking, observability, and model-serving stack. The role deploys and benchmarks open-weight models, manages routing, quotas, and GPU monitoring, and helps engineers integrate internal models, using Linux with Ansible, Terraform, Docker, and Kubernetes.
Sign in to see your match