Point your AI agent at freehire and let it find you a job.

Get the CLI →

bitdeer

NewBe an early applicant

Senior AI Data Center Network Engineer

Posted Updated
Discussion

Summary

Senior network engineer on Bitdeer's US AI Cloud team, architecting, deploying, and operating high-availability networks for large-scale GPU clusters. Core work spans InfiniBand, RoCEv2, VXLAN EVPN, and SDN fabrics, with automation via Python/Ansible/Terraform and monitoring via NVIDIA UFM, NetQ, Zabbix, and Prometheus.

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit

Position Overview
We are seeking an experienced and highly skilled Senior AI Data Center Network Engineer to join the Bitdeer AI Cloud team in the US. In this role, you will architect, deploy, and operate high-performance, high-availability network solutions for our large-scale GPU clusters. You will be instrumental in ensuring the stability, low latency, and optimal performance of our AI infrastructure, focusing heavily on advanced networking technologies such as InfiniBand, RoCEv2, and SDN architectures. You will collaborate with cross-functional teams to build the backbone of our AI Cloud, enabling cutting-edge AI training and inference workloads for our global customers.

Key Responsibilities

  • Network Architecture & Design: Architect scalable, high-availability network solutions for AI Cloud Data Centers, including Spine-Leaf topology, DCN, DCI, and backbone networks using VXLAN EVPN and SDN technologies.
  • High-Performance AI Networking: Design, deploy, and optimize large-scale GPU clusters utilizing InfiniBand and RoCEv2 (Lossless Ethernet) fabrics. Configure and manage NVIDIA Spectrum/Quantum series switches.
  • Cluster Operations & Troubleshooting: Lead deep-dive investigations into complex network issues affecting AI workloads, such as RDMA packet loss, congestion control (PFC/ECN), latency, and NCCL communication timeouts.
  • Network Automation & DevOps: Develop and maintain network automation tools and platforms using Python, Ansible, and Terraform to implement Infrastructure as Code (IaC) and streamline configuration management.
  • Fabric Management & Monitoring: Utilize tools like NVIDIA UFM (Unified Fabric Manager), NetQ, Zabbix, and Prometheus to ensure real-time monitoring, network telemetry, and overall fabric health.
  • Cloud & Container Networking: Support network integration for Kubernetes/Docker container environments and hybrid cloud deployments, ensuring seamless connectivity and security.
  • Incident Response & Change Management: Lead critical network changes, capacity expansions, and firmware upgrades. Provide rapid response to incidents, implement mitigations, and conduct Root Cause Analysis (RCA).

Requirements

  • Bachelor's degree or above in Computer Science, Network Engineering, Telecommunications, or a related field.
  • Minimum 8-10 years of experience in large-scale data center network operations, architecture, and engineering.
  • Deep expertise in the TCP/IP protocol stack and core routing protocols (BGP, OSPF, ISIS), as well as Data Center technologies (VXLAN EVPN, Spine-Leaf).
  • Extensive hands-on experience with HPC/AI networking: InfiniBand architecture, Subnet Manager, or Ethernet-based RoCEv2, including PFC, ECN, and congestion control mechanisms.
  • Proficiency in configuring and managing data center switches and routers from major vendors (e.g., Cisco Nexus, NVIDIA/Mellanox, Juniper).
  • Strong skills in network automation and scripting (Python, Ansible, Terraform) and Linux system administration.
  • Experience with Kubernetes container networking (CNI) and cloud-native architectures.
  • Professional fluency in English and Chinese is required to effectively collaborate with global teams.

Equal Opportunity Employer
Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Skills

See also

Network Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available