AI Infrastructure Manager
AI Infrastructure Operations & Maintenance
- Manage and maintain AI computing infrastructure, including VMware virtualization, Linux & Windows Server environments, and x86-based systems
- Perform system health checks, capacity planning, and performance tuning for AI workloads
- Troubleshoot complex issues across virtualization, storage, AD, and security systems
- Lead incident response and problem resolution for critical infrastructure failures
Security & Access Management
- Implement and maintain security controls including AD authentication, bastion host access, and network security policies
- Conduct regular security audits, vulnerability assessments, and patch management
- Ensure compliance with IT security standards and regulations
Infrastructure Management
- Administer Nvidia AI SU and HPC systems
- Administer VMware vSphere/ESXi environments, Linux & Windows Server systems
- Manage enterprise storage solutions (SAN/NAS) for AI data requirements
- Oversee x86 server hardware maintenance and lifecycle management
- Collaborate with network and cloud teams for integrated infrastructure solutions
- Supervise infrastructure operations team and coordinate cross‑functional projects
- Develop and maintain system documentation, operational procedures, and incident reports
- Provide technical guidance and mentorship to junior engineers
Qualifications & Experience
- Bachelor's degree or higher in Computer Science, Information Technology, or related field
- 8+ years in IT infrastructure operations, preferably in AI/cloud/data center environments
- Strong expertise in Nvidia SU, VMware, Windows Server, x86 systems, storage, AD, and cybersecurity
- Experience with AI/ML infrastructure support is highly desirable
- Certifications (Preferred): Nvidia, VMware (VCP/VCAP), Microsoft (MCSE), ITIL, CISSP, or other relevant certifications
- Fluent in English and Mandarin (both written and spoken); Cantonese proficiency is a strong plus
- Knowledge of containerization /cluster technologies (Docker/Kubernetes)
- Experience with private network and hybrid cloud environments (AWS/Azure)