Point your AI agent at freehire and let it find you a job.

Get the CLI →

clera

NewBe an early applicant

ML Infrastructure Engineer

Posted Updated
Discussion

Summary

Hands-on infrastructure engineer at an early-stage enterprise AI company, owning inference and model-serving infrastructure end to end. Day to day: designing, deploying, and scaling systems so AI agents run fast and reliably under growing concurrent load, using Kubernetes, cloud platforms, ML serving frameworks, and observability tooling.

About the Role

This is a hands-on infrastructure engineering role at an early-stage enterprise AI company building a context and data governance layer for AI agents deployed in highly regulated industries. You will own the inference and model-serving infrastructure end to end, making production AI agents fast, reliable, and scalable as concurrency grows.

What You'll Do

  • Design, build, and own inference and model-serving infrastructure from initial architecture through production deployment.

  • Scale systems that enable AI agents to run reliably and efficiently under increasing concurrent load.

  • Identify and resolve infrastructure bottlenecks in collaboration with ML and platform engineering teams.

  • Drive performance optimization across latency, throughput, and reliability for production workloads.

What We're Looking For

  • 5+ years building and operating ML inference systems, model-serving platforms, or ML infrastructure in production environments.

  • Hands-on experience designing and scaling inference-serving systems using frameworks such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom solutions.

  • Strong distributed systems fundamentals, including experience managing concurrent requests and resource allocation under load.

  • Proficiency with containerization and orchestration technologies, particularly Docker and Kubernetes, for ML workloads.

  • Experience with cloud infrastructure platforms (AWS, GCP, or Azure) for deploying and managing ML systems.

  • Solid monitoring and observability skills using tools such as Prometheus, Grafana, ELK, or distributed tracing solutions.

  • Proficiency in at least one systems or backend language: Python, Go, Rust, C++, or Java.

  • Familiarity with knowledge graphs, semantic search, or graph databases is a plus.

  • Background in agentic or autonomous AI systems, real-time inference, or enterprise data infrastructure is a plus.

Location

On-site in San Mateo, California, United States. Visa sponsorship is not available for this role.

Skills

See also

DevOps jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available