Silicon Infrastructure Engineer
You will build and operate the infrastructure layer powering heterogeneous GPU clusters and datacenters. You will design provisioning, scheduling, isolation, monitoring, networking, storage, and lifecycle-management systems; develop infrastructure software in Rust and Python; operate Kubernetes and Slurm environments; secure multi-tenant compute; collect telemetry and infrastructure state; and debug failures across Linux, networking, storage, virtualization, GPUs, and distributed systems.
Responsibilities
- Build and operate infrastructure across heterogeneous GPU clusters and datacenters
- Design systems for provisioning, scheduling, isolating, monitoring, and managing compute
- Build infrastructure software in Rust and Python for node management, orchestration, telemetry, networking, and cluster operations
- Operate and extend Kubernetes and Slurm environments
- Design secure multi-tenant compute environments
- Debug failures across Linux, networking, storage, schedulers, virtualization, GPUs, and distributed systems
- Build systems for collecting and querying infrastructure state, telemetry, inventory, health, and utilization data
- Automate infrastructure deployment and lifecycle management across clusters
- Design systems that degrade predictably, recover automatically, and expose failure information
Requirements
- Strong Linux systems knowledge
- Strong Rust knowledge and experience building production systems software
- Working knowledge of SQL and experience with data-intensive backend systems
- Experience operating containerized and virtualized workloads in production
- Deep familiarity with Kubernetes and/or Slurm
- Understanding of virtualization fundamentals including hypervisors, KVM/QEMU-style architectures, virtio, device passthrough, and workload isolation
- Strong understanding of distributed systems
- Ability to debug across application, kernel, networking, scheduler, hypervisor, and physical server layers
- Ability to operate independently and own systems from design through production
- BS, MS, or equivalent experience in computer science, computer engineering, electrical engineering, or a related technical field
Benefits
- Meaningful equity
- Health coverage
- Free meals