Member of Technical Staff - Compute Platform
Summary
Builds and maintains the platform that schedules, monitors, and runs AI workloads, using Python, Rust, Kubernetes, and cloud infrastructure.
You will build the platform software and infrastructure used to manage and monitor AI workloads. You will develop web interfaces, Python APIs and backend services, real-time debugging tools, distributed training infrastructure in Rust, automation pipelines, cloud resources, container orchestration, and scheduling systems for heterogeneous hardware.
Responsibilities
- Build web interfaces for AI workload management and monitoring
- Develop REST APIs and backend services in Python
- Create real-time monitoring and debugging tools
- Implement resource management and job control features
- Design distributed training infrastructure in Rust
- Build networking and coordination components
- Create Ansible infrastructure automation pipelines
- Manage cloud resources and container orchestration
- Implement scheduling systems for CPU, GPU, and TPU hardware
- Integrate backend features into existing infrastructure
Requirements
- Strong Python backend development with FastAPI and async
- Modern frontend development with TypeScript, React/Next.js, and Tailwind
- Developer tools and dashboard development
- RESTful API design and implementation
- Rust systems programming
- Ansible and Terraform automation
- Kubernetes
- GCP or other cloud platforms
- Prometheus and Grafana
- GPU computing or ML infrastructure experience