Network Engineer (Data Centre / GPU Infrastructure)
Summary
Design and operate high-performance networking for an AI data centre, deploying VXLAN/EVPN/BGP and automating with Ansible, Python, and monitoring tools.
We are looking for a Network Engineer to join our team. You will design, deploy, and operate high-performance networking for an AI data centre, supporting customers running large-scale AI and HPC workloads on GPU clusters.
You will work with advanced routing protocols, Linux-based network operating systems, and automation tools in a fast-growing AI infrastructure environment.
Responsibilities
Design, develop, and operate high-performance networking solutions for GPUaaS platforms using VXLAN, EVPN, and BGP
Provide operational support for GPU infrastructure services, ensuring SLA compliance and minimising downtime
Monitor, troubleshoot, and optimise network performance using tools like Zabbix, Netbox, Ansible, and Python scripts
Implement network security measures and support industry security certifications
Coordinate with cross-functional teams (internal and external) to deliver network solutions on time
Prepare technical reports — outage reports, SLA reports, and network optimisation plans — for management and customers
Participate in on-call rotation / work outside standard hours when required (nights, weekends, public holidays)
Requirements
Must-Have
Degree in Computer Science, Information Technology, Network Engineering, or a related field — OR equivalent networking certifications plus Linux certification
Hands‑on experience deploying and troubleshooting VXLAN, EVPN, and BGP
Strong understanding of TCP/IP, VLANs, and subnetting
Working knowledge of network operating systems: Cumulus Linux, Arista, Ubuntu, Proxmox, or Redhat
Experience with network monitoring and automation: Zabbix, Netbox, Ansible, Python scripting, CI/CD
Strong verbal and written communication skills in English
Customer‑service oriented and comfortable working with multiple stakeholders
Good-to-Have
Experience with RDMA, InfiniBand, or RoCE protocols
Understanding of how AI and HPC workloads interact with high-performance networks
System-level experience with GPU‑accelerated networking
Knowledge of cloud architectures (IaaS, PaaS) and NVIDIA GPU architecture
Experience with Linux, hypervisors, storage (NFS, Object), and infrastructure-as-code