Member of Technical Staff - ML Operations
Summary
Builds and operates the ML infrastructure behind Veeda AI's Physical AI world models: experiment tracking tools, inference fleet orchestration, model evaluation in CI, large-scale data pipelines, and visualization platforms. Day-to-day is platform/backend engineering with Python, Go/Rust/Java, CI/CD, and distributed systems for training and serving.
ABOUT US
Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.
RESPONSIBILITIES
Experiment Lifecycle Tracking and Tooling: Design, build and deploy tools that own how a run is defined, launched, resumed, and killed. Build and operate experiment databases for code version tracking, data version tracking, reproducibility, and checkpoint ancestry.
Inference Fleet Orchestration: Design, build, and operate the serving control plane that accepts large volumes of concurrent client requests and assigns them across inference clusters. Develop cache-aware admission, routing, batching, and scheduling policies that improve cache locality, balance workload, protect tail latency, and keep the fleet highly utilized and reliable. Partner with ML Performance on model runtime, kernel, and per-worker throughput optimization.
Model Evaluation in CI: Design, build, and operate automatic model checkpoint evaluation systems on seeded rollout and policy-success suites, run per-change and nightly.
Data Pipeline Operations: Design, build, and operate high-performance, fault-tolerant, distributed backend services and event-driven systems for our large scale data processing pipeline.
Visualization Platforms: Design, build, and deploy experiment observability and dataset visualization platform(s) that provides interactive data visualization, progress tracking, search, and comparison.
End-to-End Ownership: Lead projects through the complete software lifecycle, including technical specs, implementation, CI/CD, on-call support, and production observability.
REQUIREMENTS
Bachelor's degree or equivalent hands-on experience in Computer Science, Engineering, or a related technical field.
Proficiency in at least one scripting (Python or Bash) and one compiled (Java, Rust, or Go) languages.
Experience in shipping production-quality developer tools.
Proficient in CICD automations (pipelines, runners, deployment)
Experience in building reproducible pipelines end to end, and can say precisely which parts of a training run are bit-reproducible, which are not, and why.
One of the following:
Proven understanding of event-driven architecture, concurrency models, fault tolerance, and data consistency patterns.
Experience building or operating large-scale inference control planes or distributed serving infrastructure, with hands-on work in traffic management, admission control, request scheduling, routing, batching, or cache-aware load balancing; able to reason clearly about cache locality, queueing, tail latency, availability, and fleet utilization.
Experience working with multi-node workloads and building around slurm based scheduling systems.
Experience working on in-production model evaluation frameworks, in particular regression testing of large multimodal models.
NICE TO HAVE
You have run experiment tracking at scale, logging video, 3D, and trajectory artifacts rather than only scalars.
You have built evaluation harnesses for generative or embodied models, where quality is a distribution rather than a pass/fail.
Full stack development experience with web-based front-end.
You have orchestrated ML workflows with Argo Workflows, Flyte, or Ray, and know where each one breaks.
You have built GPU-hour attribution that maps cluster spend back to specific experiments and teams.
You have contributed to open-source ML tooling, or published on evaluation or reproducibility methodology.
Skills
As published by ashby · 2 questions
Name, Email, Resume
- Why do you want to work at Veeda AI?
- Describe your most significant accomplishment