Tech jobs
Job listings
AI Cluster Architect
Designs and optimizes large-scale GPU clusters for AI workloads, balancing power budgets, networking fabrics (InfiniBand/RoCE), and hardware constraints to maximize GPU density while meeting performance needs.
Principal Software Engineering Manager
Overview The HPC/AI (High-Performance Computing and Artificial Intelligence) organization is on a mission to build the next generation of distributed AI supercomputers - systems that deliver unprecedented computational…
Senior Technical Program Manager, Cluster Operations & Quota Management
Owns the end-to-end system that turns GPU allocation decisions into usable capacity for Microsoft AI’s compute fleet, coordinating moves, provisioning, and validation across teams.
Platform Hardware Engineer – Data Center Accelerator Platfor
Design and validate next-gen accelerator hardware platforms for AI/ML and HPC, spanning board design, manufacturing, bring-up, and system integration.
Customer Support Lead
We're building the company which will de-risk the largest infrastructure build-out in history. When people finance GPU clusters, the datacenters housing them, and the infrastructure powering them, they need…
RF Antenna and Radome Design Engineer
Senior Principal RF Antenna and Radome Design Engineer Location: US-AZ-TUCSON-M02 ~ 1151 E Hermans Rd ~ BLDG M02 Time Type: Full time Job Description Date Posted: 2026-06-15 Country: United States of America Location:…
Staff Software Engineer - Managed Kubernetes
Build and lead Lambda’s Managed Kubernetes platform for AI workloads, integrating NVIDIA’s GPU ecosystem and designing GPU-aware orchestration, storage, and networking for bare-metal clusters.
Identity & Access Management (IAM) Engineer with Security Clearance
Implement access controls and permission mechanisms for securing high-performance computing systems with an active TS/SCI clearance.
Vulnerability Management Lead with Security Clearance
Lead a large-scale vulnerability management program for 30,000+ enterprise assets, including servers, HPC systems, cloud workloads, and network infrastructure.
RF Range Architect and Integrator
Leads the design and implementation of RF antenna, radome, and radar cross-section ranges/chambers, collaborating with cross-functional teams and government stakeholders to validate compliance and modernize testing infrastructure.
Senior Rack Integration & Hardware Engineer
Leads hardware integration for AI/data center racks, designing and validating next-gen rack solutions, resolving cross-functional technical challenges, and ensuring deployment readiness for AMD’s AI infrastructure.
Member of Technical Staff - Thermal Simulation Engineer
Build and validate thermal-physics solvers and training data for an AI-powered hardware-design assistant, then run customer cases and present results to engineering teams.
Domain Architect - AI Compute
Salary: $150 – $170 per hour About the Role The Domain Architect - AI Compute acts as the primary technical authority for the physical and logical lifecycle of high-performance GPU compute fleets across diverse client…
IT Project Manager - PMP required
Oversee hybrid IT projects for a pediatric research group, coordinating hybrid teams and infrastructure deployments with a focus on data center and HPC environments.
Software Engineer, GPU Cluster Infrastructure
Build and operate a unified GPU compute platform for AI training and inference, hiding cloud complexity behind Kubernetes operators and schedulers that manage multi-cloud NVIDIA clusters.
Technical Program Manager
Technical Program Manager at an AI-infra startup coordinating capacity planning, SLA design, and cross-team execution for GPU clusters used in training and inference workloads.
GPU Cluster Architect
Designs, validates, and deploys large-scale GPU clusters for an AI-infra startup, negotiating with CSPs and vendors to ensure optimal power, cooling, and high-speed networking for AI workloads.
Platform Hardware Engineer, Data Center Accelerator Platforms
Design and validate next-gen accelerator boards for AI/ML/HPC using Openchip silicon, spanning schematic reviews, PCB layout, bring-up, validation, and system integration.