freehire launches on Product Hunt on 26 August.

Follow →

Senior AI Data Center Network Engineer

Open 27d

You will architect high-availability network solutions for AI Cloud Data Centers, covering DCN, DCI, WAN, and backbone networks. You will monitor, troubleshoot, and tune performance for large-scale GPU clusters running on InfiniBand or RoCEv2 fabrics, and lead deep-dive investigations into complex issues affecting AI training and inference performance such as RDMA packet loss, latency, link flapping, and NCCL communication timeouts. You will manage NVIDIA Quantum and Spectrum series switches, ConnectX NICs, NetQ, and the UFM platform to keep the network fabric healthy and stable. You will lead network architecture changes, capacity expansions, cutovers, and firmware upgrades with zero incidents, and respond rapidly to critical incidents with immediate mitigation and thorough root cause analysis. You will build and maintain network monitoring and observability platforms, and develop automation tools to improve operational efficiency and standardize workflows.

Responsibilities

  • Architect high-availability network solutions for AI Cloud Data Centers covering DCN, DCI, WAN, and backbone networks
  • Orchestrate daily monitoring, troubleshooting, and performance tuning for large-scale GPU clusters based on InfiniBand or RoCEv2 fabrics
  • Lead investigations into complex network issues affecting AI training and inference performance, such as RDMA packet loss, latency, link flapping, and NCCL communication timeouts
  • Manage NVIDIA Quantum series switches, NVIDIA Spectrum series switches, ConnectX NICs, NetQ, and the UFM platform to ensure fabric health and stability
  • Lead network architecture changes, capacity expansions, cutovers, and firmware upgrades ensuring smooth execution with zero incidents
  • Provide rapid response to critical network incidents, implementing immediate mitigation measures and conducting thorough Root Cause Analysis
  • Build and maintain network monitoring and observability platforms utilizing Zabbix, Prometheus, Grafana, or Telegraf
  • Develop network automation tools or platforms using Python, Go, Ansible, or Terraform to improve operational efficiency and standardize workflows

Requirements

  • Bachelor's degree or above in Computer Science, Telecommunications, or a related field
  • 10+ years of experience in large-scale network operations or architecture
  • Proficient in the TCP/IP protocol stack; extensive mastery of routing protocols (BGP, OSPF, ISIS) and Data Center technologies (EVPN-VXLAN)
  • Familiarity with InfiniBand architecture and Subnet Manager principles, with hands-on experience operating NVIDIA (Mellanox) switches; or proficiency in Ethernet-based RoCEv2 technology with deep understanding of PFC and ECN mechanisms, adaptive routing, congestion control like Spectrum-X CC, ZTRRTT CC, DCQCN
  • Proficient in Linux system operations
  • Hands-on experience in the construction or operations of large-scale GPU clusters (1,000+ GPUs), specifically utilizing NVIDIA H100 or GB200 platforms
  • Familiarity with the principles of AI distributed training communication libraries (e.g., NCCL, MPI) and ability to diagnose/infer network issues by analyzing application-level training logs
  • Proficiency with advanced features of NVIDIA UFM such as SHARP and network telemetry
  • Professional certifications such as CCIE, JNCIE, NVIDIA-Certified Professional AI Networking (NCP-AIN), or Advanced Networking Specialty certifications from major public clouds
  • Experience with Optical Transmission equipment (DWDM/DCI) or managing global backbone networks
  • Mastery of at least one scripting language (Python or Go) with experience in developing network automation platforms or writing operational scripts

Benefits

  • Attractive welfare benefits
  • Training and mentoring opportunities

See also