ML Infrastructure Engineer
About the Role
This is a hands-on ML Infrastructure Engineer role at an early-stage enterprise AI startup, where you'll own the end-to-end inference and model-serving infrastructure that keeps production AI agents running reliably and at scale. You'll sit at the intersection of ML and platform engineering, directly shaping the systems that power real-world, high-stakes deployments in regulated industries like insurance, banking, and healthcare.
What You'll Do
Own inference and model-serving infrastructure end to end, from design through production deployment.
Build and scale systems that enable AI agents to run reliably and efficiently under increasing concurrency.
Collaborate closely with ML and infrastructure teams to ensure seamless integration and performance optimization.
Identify infrastructure bottlenecks and drive cross-functional solutions across engineering teams.
What We're Looking For
5+ years of experience building and operating ML inference systems, model-serving platforms, or ML infrastructure in production.
Hands-on experience designing and scaling inference-serving infrastructure using frameworks such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems.
Strong track record optimizing production ML systems for latency, throughput, and reliability at scale.
Experience with containerization and orchestration (Docker, Kubernetes) for deploying and scaling ML workloads.
Experience building distributed systems that handle concurrent requests and manage resource allocation under load.
Proficiency with observability and debugging tooling for production systems (e.g., Prometheus, Grafana, ELK, distributed tracing).
Cloud platform experience on AWS, GCP, or Azure for deploying and managing ML systems.
Proficiency in at least one systems or backend language — Python, Go, Rust, C++, or Java.
Nice to have: experience with knowledge graphs, semantic search, or graph databases (e.g., Neo4j, Amazon Neptune); real-time or low-latency inference systems; agentic or multi-step reasoning pipelines; enterprise data infrastructure or integration platforms.
Location
On-site in San Mateo, CA. No visa sponsorship is available for this role.