Senior Staff Software Engineer Kubernetes Infrastructure
Summary
Senior staff engineer who designs, provisions, and operates customer compute environments at fal — building dedicated Kubernetes and Slurm clusters on bare-metal NVIDIA GPU infrastructure, covering Linux/OS provisioning, networking, storage, observability, and reusable ops tooling in Python/Go.
You will design, provision, operate, upgrade, recover, and decommission customer compute environments. You will build Kubernetes and Slurm clusters, Linux provisioning workflows, GPU infrastructure, networking, storage, observability, and reusable operational tooling.
Responsibilities
- Design and deliver the full lifecycle of customer compute environments
- Automate infrastructure delivery and operations with AI
- Provision dedicated Kubernetes and Slurm clusters
- Build Linux images and automated OS-provisioning workflows
- Operate NVIDIA GPU infrastructure
- Design Kubernetes and data-center networking
- Configure distributed and shared storage
- Build monitoring, alerting, diagnostics, and automated recovery
- Develop reusable tooling, standards, documentation, and runbooks
- Translate workload requirements into infrastructure designs
Requirements
- Linux
- Kubernetes
- Bare metal
- etcd
- containerd
- CNI
- CSI
- KVM
- QEMU
- libvirt
- VFIO
- NVIDIA GPU
- TCP/IP
- VLAN
- routing
- tcpdump
- Wireshark
- Ansible
- Slurm
- Python
- Go
Benefits
- Equity
- Visa sponsorship
- Relocation assistance to San Francisco
- Health insurance
- Dental insurance
- Vision insurance
- Regular team events and offsites
