freehire launches on Product Hunt on 26 August.

Follow →

Member of Technical Staff - Platform

Summary

The Platform team member manages the infrastructure layer between bare metal and inference engines, focusing on Kubernetes orchestration, cluster scaling, networking, and observability for AI compute clusters. You will work across heterogeneous GPU hardware to ensure reliable, scalable deployment of AI models in company-owned data centers.

About ai&

ai& is a new global AI technology company dedicated to meeting the world's growing demand for AI. Our vision is twofold: to serve as a premier AI lab specializing in localization, and to act as a global infrastructure and compute provider. We are building a unified, optimized global platform that integrates next-generation data centers and infrastructure, heterogeneous compute serving, and advanced model services. We believe that the most effective way to build and scale AI is to own the stack from top to bottom.

At ai&, we empower small teams with the autonomy needed to tackle significant challenges. Our approach is to deconstruct large problems into manageable components and solve complex issues collaboratively. We seek highly motivated, mission-driven individuals who demonstrate strong personal agency. We value curiosity as the foundation of talent, and we are looking for people eager to develop alongside our evolving technology and expanding business.

We are actively hiring worldwide, with presence in Tokyo, SF, Austin, and Toronto. We are more than happy to meet exceptional talent where they are.

Role overview
The Platform team turns raw heterogeneous compute into a serving platform. ai& owns its data centers and runs AMD, NVIDIA, and Tenstorrent silicon side by side. Your job is everything between the bare metal and the inference engines: cluster orchestration, node lifecycle, scaling, networking, observability, and reliability.
This is not cloud consumption. When capacity is short, you add nodes we own. When a fabric misbehaves, you debug it down to the switch. The platform must let a small team operate hundreds of nodes across multiple sites without heroics, and it must scale by an order of magnitude over the next two years as new sites come online.
You will work directly with the inference team, which owns the engines and serving gateway, and the data center team, which owns power, cooling, and physical deployment. You own the layer that makes their work composable.
Responsibilities

  • Compute orchestration Run Kubernetes across GPU clusters in ai&-owned data centers. Own node lifecycle from bring-up and burn-in through drain and repair, across multiple accelerator vendors.

  • Scaling Build the capacity and scheduling machinery that places inference workloads across heterogeneous silicon and multiple sites, and that lets us bring a new site from empty racks to serving traffic on a predictable timeline.

  • Reliability Define and hold SLOs for the platform. Build the observability stack (metrics, logs, tracing, alerting) and the failure isolation that keeps one bad node or one bad rollout from becoming an incident.

  • Networking and data Operate high-bandwidth fabrics for multi-node inference. Solve model weight distribution: getting hundreds of gigabytes onto the right nodes fast, every time a model ships.

  • Deployment machinery Own CI/CD and GitOps for the fleet. Infrastructure as code, reproducible node images, safe rollouts.

You may be a fit if you have the following skills

  • Production Kubernetes at scale You have operated large multi-cluster Kubernetes environments, ideally with GPU scheduling, device plugins, and topology-aware placement.

  • Systems depth Strong Linux fundamentals. You can reason about NUMA, PCIe, NICs, and storage, and you debug from symptoms to root cause without guessing.

  • Networking fundamentals You understand L2/L3, and ideally RDMA fabrics (InfiniBand or RoCE) in production.

  • Infrastructure as code Terraform or equivalent, GitOps workflows, and the discipline to keep the fleet reproducible.

  • Ownership under load You have carried a pager for systems that matter and you build so the pager stays quiet.

  • Relevant tooling Go or Python, Prometheus-family observability, and comfort automating anything you do twice.

  • Great team spirit A mission-driven approach to engineering, valuing clear communication, hands-on execution, and collective success over individual silos.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available