AI Cluster Engineer
You'll study, understand, and optimize large-scale AI computing clusters, keeping hardware and software environments running smoothly. You will provide advanced engineering support across the entire AI compute stack, covering systems (operating system configuration, job scheduling, containerization), networking (high-speed interconnects, fabric management, and optimization for AI model training), and storage (deploying and maintaining high-performance parallel file systems for large datasets and AI workloads). You'll collaborate with cross-functional teams on system upgrades or expansions, document architecture, configurations, and operational processes, and participate in long-term planning for AI infrastructure growth. You will ensure high system availability and performance tuning, and implement best practices for security, operations, and maintenance.
Responsibilities
- Study, understand, and optimize large-scale AI computing clusters
- Maintain and monitor AI cluster hardware and software environments
- Provide advanced engineering support across the entire AI compute stack
- Configure operating systems, job scheduling, and containerization
- Manage high-speed interconnects and fabric optimization for AI model training
- Deploy and maintain high-performance parallel file systems for large datasets and AI workloads
- Collaborate with cross-functional teams on system upgrades or expansions
- Document architecture, configurations, and operational processes
- Participate in long-term planning for AI infrastructure growth
- Ensure high system availability and performance tuning
- Implement best practices for security, operations, and maintenance
Requirements
- 5 years of experience in large-scale server or data center operations, HPC/distributed cluster systems, or GPU server environments
- Network, storage, or systems engineering experience
- Bachelor's Degree in Computer Science, Information Technology or related field
- Strong understanding of server architecture, operating systems, networking, and storage
- Experience with large-scale compute systems, HPC, or distributed clusters preferred
- Ability to analyze and solve complex system and network issues
- Familiarity with AI/ML infrastructure or GPU-based environments is an advantage
- Excellent communication, documentation, and teamwork skills
- Background in data center operations or enterprise IT environments preferred
- Hands-on experience with cluster management tools or monitoring systems preferred
- Ability to work independently and proactively research new technologies
Benefits
- Attractive welfare benefits and developmental opportunities such as training and mentoring