freehire launches on Product Hunt on 26 August.

Follow →

Principal Network Engineer

Summary

Design, automate, and operate large-scale InfiniBand and Ethernet network fabrics for a GPU cloud, using Python, Ansible, and GitOps in a high-performance data center environment.

Principal Network Engineer, AI Infrastructure & High-Performance Networking

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud strengthens technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

At Nscale, our Engineering team plays a critical role in deploying and operating the infrastructure and software platforms that power our customers.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.

About the Role

The Network Engineering Team is responsible for the design, validation, and ongoing operation of all networking services that underpin both the internal management platform and the customer-facing cloud infrastructure — including high-performance Ethernet fabrics, InfiniBand, WAN connectivity, and data center networking. The team also acts as a 3rd/4th line escalation point for the support organization.

As a Principal Network Engineer, you will be a senior technical authority for Nscale’s AI-optimized network fabrics. You will set technical direction across low-latency, high-bandwidth InfiniBand and Ethernet networks supporting large-scale training and inference workloads; own critical technical domains end to end; and raise the bar for architecture, automation, operational rigor, and engineering standards across the organization.

You will combine deep hands-on engineering with broad architectural influence. You’ll define reference architectures, drive consistency across sites, lead complex technical decisions and escalations, and mentor engineers while partnering closely with deployment, data center operations, platform engineering, and vendors.

What You'll Be Doing

  • Define, design, validate, and evolve large-scale InfiniBand/RoCE and Ethernet fabric architectures at rack, row, and data center scale, with tight integration to bare-metal provisioning and cluster management systems.
  • Own technical direction for high-performance Ethernet fabrics, including BGP, EVPN, VXLAN, LACP, and QoS, and establish reference architectures and standards implemented consistently across sites.
  • Design and engineer perimeter and security infrastructure — firewalls, NAT, VPN, and security policy architecture — across WAN and data center edge environments.
  • Lead network automation strategy in a GitOps model, building and guiding Python/Ansible tooling for provisioning, configuration validation, and compliance, with version-controlled configuration and CI/CD-driven change across multi-vendor environments.
  • Drive operational excellence by leading complex escalations and root-cause analysis for performance and stability issues, and systematically reducing reactive toil through runbooks, automation, and measurable SLOs.
  • Set the direction for network observability, telemetry, monitoring, and alerting to provide clear visibility into fabric health, performance, and traffic patterns.
  • Ensure the accuracy and reliability of source-of-truth network inventory and configuration data, with changes flowing through structured engineering and change-management practices.
  • Partner with deployment, data center operations, platform engineering, systems, storage, and vendors on new site delivery and platform evolution.
  • Act as a technical mentor and force multiplier across the team through architecture reviews, design reviews, incident leadership, documentation, and knowledge sharing.
  • Identify systemic risks and architectural gaps across sites and drive durable solutions that improve scalability, reliability, and operational simplicity.

About You (Skills / Qualifications)

  • 10+ years of network engineering experience, with significant depth in HPC, AI, hyperscale, or large-scale data center environments.
  • Extensive hands-on experience with RDMA-aware networking for AI/HPC workloads, including InfiniBand and/or RoCE, subnet managers such as OpenSM/UFM, and fabric orchestration.
  • Expert-level knowledge of modern data center routing and control planes, including BGP, EVPN-VXLAN, and Clos/spine-leaf architectures, with production experience on platforms such as Cumulus, Nokia, or Arista EOS.
  • Strong network automation expertise using Python and Ansible, Git-based workflows, and modern infrastructure-as-code and pipeline tooling such as Terraform, GitLab CI, or GitHub Actions; you treat the network as code rather than managing devices by hand.
  • Deep design and engineering experience with firewall platforms such as Juniper SRX and/or Palo Alto, including security policy architecture, high-availability design, and multi-tenant segmentation.
  • Experience designing network telemetry and observability for high-throughput, performance-sensitive environments.
  • Proven ability to lead complex technical decisions and incidents across networking, systems, storage, and HPC/AI workload teams, with the judgment to balance performance, reliability, operability, and delivery velocity.
  • Demonstrated experience defining architecture, standards, and technical strategy beyond a single project or site, and influencing engineering teams without relying on formal authority.
  • Strong communication and mentoring skills, with the ability to make complex technical trade-offs clear to engineering leaders, operators, and cross-functional partners.
  • Hands-on, adaptable, and comfortable operating with high ownership in a fast-paced environment building next-generation infrastructure for ML scale-out.

Equal Opportunities Statement

At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enriches our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.

If there’s anything we can do to accommodate your specific situation, please let us know.

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range
$270,000$330,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

What this application asks

greenhouse

First Name, Last Name, Email, Phone, Resume/CV, Cover Letter, Location

  • LinkedIn Profile optional
  • Website optional
  • Do you have the legal right to work in the country in which this job is located, without requiring visa sponsorship? choose one
  • What is your current or most recent employer? optional
  • What is your preferred name?

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available