GPU Infrastructure Site Reliability Engineer(On-site, L2- Face to Face)
Summary
Maintains and optimizes GPU and AI infrastructure to ensure high availability and reliability in a production environment.
Job Title: Site Reliability Engineer (SRE) Location: Sunnyvale, CA (On-site) Duration: Long-Term Contract Job Description We are seeking a highly motivated Site Reliability Engineer (SRE) to support mission-critical AI and GPU infrastructure in a high-performance production environment. The ideal candidate will have experience supporting GPU platforms, embedded infrastructure, compute, networking, and storage systems while ensuring high availability, reliability, and operational excellence. As …