Kubernetes ML Inference Engineer: Model Serving
Summary
Builds and optimizes Kubernetes-based infrastructure to deploy and serve AI models in real-time for healthcare applications.
Abridge is seeking an ML Infrastructure Engineer, Model Inference in San Francisco to build and optimize the core inference infrastructure powering our AI‑driven healthcare solutions. You will collaborate across Infrastructure and Research to deploy, optimize, and orchestrate AI models at scale.
Ideal candidates have 2+ years of production ML infrastructure experience, strong Kubernetes know‑how, and a track record of engineering scalable APIs and distributed systems for real-time workloads.
a16z