Senior AI Cloud Network Operations Engineer
You monitor the full network estate around the clock, respond to outages and device failures, manage network tickets, investigate AI network performance issues, produce health reports and post-incident reviews, support customer and R&D change requests, and execute controlled network changes.
Responsibilities
- Monitor switches, routers, optical transport, device health, bandwidth, and traffic load
- Monitor AI-specific network signals
- Respond to network outages, link failures, and device-down events
- Investigate jitter, NCCL throughput degradation, and AI network performance issues
- Manage network request tickets from creation through resolution
- Produce network health reports and post-incident reviews
- Maintain the network operations knowledge base
- Handle network change requests for internal R&D teams and customers
- Execute network changes under change management procedures
Requirements
- 5+ years of network operations experience
- Degree in computer science, telecommunications, electronics, or a similar field
- Knowledge of BGP, OSPF, VXLAN, EVPN, and ECMP
- Experience with enterprise-grade switches and routers
- Working knowledge of InfiniBand or RoCEv2
- Understanding of PFC and ECN congestion control
- Familiarity with Zabbix, Prometheus, Grafana, Cacti, and MTR
- Fluency in Chinese and English
- Experience with large-scale GPU clusters is preferred
- Familiarity with NCCL and MPI is preferred
- Proficiency with NVIDIA UFM is preferred
- CCIE, JNCIE, NCP-AIN, or advanced networking certifications are preferred
- Optical transmission or global backbone experience is preferred
- Python or Go network automation experience is preferred
Benefits
- Attractive welfare benefits