Cluster Network Engineering Lead
Summary
Lead the design and operation of high-performance AI cluster networks, focusing on InfiniBand/RoCE fabrics, RDMA performance, and fabric reliability for large-scale training and inference workloads.
You will lead the design and operation of AI cluster network fabrics, owning InfiniBand and RoCE architectures, RDMA performance, and fabric reliability for large-scale training and inference workloads. You will define fabric evolution, congestion control, topology, observability, vendor strategy, migration plans, and engineering standards while mentoring senior engineers and resolving architecture escalations.
Responsibilities
- Define long-range architecture for AI fabric evolution across RoCE, InfiniBand, Clos, spine-leaf, mesh, and optical fabric designs
- Lead migration from 100G to 400G, 800G, and 1.6T
- Own congestion control strategy, queue management, path diversity, and routing policy
- Drive topology decisions that improve NCCL all-reduce performance, completion times, latency stability, and fault containment
- Establish observability standards for fabric telemetry, queue behavior, packet loss, jitter, retries, and job impact
- Set cable plant strategy across fiber topology, optics qualification, and DAC/AOC standards
- Lead vendor engagement and define upgrade and migration playbooks
- Mentor principal and staff engineers and serve as the final escalation point for fabric architecture decisions
Requirements
- Demonstrated experience in hyperscale networking, HPC fabrics, RDMA systems, or distributed systems networking at large scale
- Deep expertise in ECMP, adaptive routing, queueing theory, network telemetry pipelines, InfiniBand, and optical systems
- Proven experience designing and operating AI or HPC interconnects
- Strong knowledge of network behavior under distributed training frameworks and collective communication libraries such as NCCL and UCX
- Experience leading technology transitions across multiple hardware generations and mixed-vendor environments
- Ability to work credibly with hardware vendors, datacenter teams, software platform leaders, and executive stakeholders