Member of Technical Staff - Training Platform
Summary
Build and operate a hosted training platform for AI models, including Kubernetes orchestration, GPU scheduling, observability, and a Next.js/React UI.
You will build the hosted training platform and Kubernetes-based infrastructure for managed GPU training. You will develop orchestration, scheduling, autoscaling, control-plane agents, observability, backend APIs, monitoring tools, and product interfaces, while integrating new training capabilities.
Responsibilities
- Design and operate Kubernetes-based training and inference orchestration
- Build and maintain Helm charts for reproducible training stacks
- Develop Python control-plane agents
- Implement scheduling and autoscaling for heterogeneous GPU hardware
- Operate GitOps workflows
- Build model caches, checkpoint pipelines, and shared storage
- Operate observability systems
- Build job submission and live run monitoring surfaces
- Develop FastAPI backend services and REST APIs
- Build real-time monitoring and debugging tools
- Ship product interfaces in Next.js, React, and TypeScript
- Interface with trainers, inference servers, and environment servers
- Productize new training capabilities
Requirements
- Knowledge of open model families and fine-tuning techniques
- LoRA
- QLoRA
- RLHF
- RLAIF
- vLLM
- SGLang
- TensorRT-LLM
- GPU hardware
- Distributed training
- NCCL
- Kubernetes
- Helm
- CRDs
- KEDA
- GPU operator
- Terraform
- Ansible
- Prometheus
- Grafana
- Loki
- OpenTelemetry
- DCGM
- Linux
- Python
- FastAPI
- SQLAlchemy
- TypeScript
- React
- Next.js
- Tailwind
- REST
- tRPC
Benefits
- Significant equity
- Flexible work arrangement
- Full visa sponsorship
- Relocation support
- Professional development budget
- Regular team off-sites
- Conference attendance