Point your AI agent at freehire and let it find you a job.

Get the CLI →

clera

NewBe an early applicant

ML Infrastructure Engineer

Posted Updated
Discussion

Summary

Clera, an early-stage enterprise AI company building a context and data governance layer for AI agents in regulated industries, seeks a hands-on ML Infrastructure Engineer to own inference and model-serving infrastructure end to end. Day-to-day work centers on scaling and optimizing serving systems (Triton/KServe/TorchServe, Docker/Kubernetes, cloud) for latency, throughput, and reliability in pro

About the Role

This is a hands-on infrastructure engineering role at an early-stage enterprise AI company building a context and data governance layer for AI agents in highly regulated industries. You will own the inference and model-serving infrastructure end to end, ensuring AI agents run reliably, accurately, and at scale in production environments where performance is non-negotiable.

What You'll Do

  • Design, build, and operate inference and model-serving infrastructure from development through production deployment.

  • Scale systems to support AI agents running reliably under increasing concurrency and production load.

  • Identify and resolve infrastructure bottlenecks in close collaboration with ML and platform engineering teams.

  • Optimize systems for latency, throughput, and reliability at scale.

What We're Looking For

  • 5 or more years building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments.

  • Hands-on experience designing and scaling inference serving infrastructure using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems.

  • Strong systems engineering fundamentals with expertise in distributed systems, containerization, and orchestration (Docker, Kubernetes).

  • Demonstrated ability to optimize production ML systems for latency, throughput, and reliability under high concurrency.

  • Experience with cloud infrastructure platforms such as AWS, GCP, or Azure for deploying and managing ML workloads.

  • Proficiency with monitoring, observability, and debugging tools such as Prometheus, Grafana, ELK, or distributed tracing frameworks.

  • Proficiency in at least one systems programming or backend language: Python, Go, Rust, C++, or Java.

  • Experience with knowledge graphs, semantic search, or graph databases (e.g., Neo4j, Amazon Neptune) is a plus.

  • Familiarity with agentic AI systems, autonomous agents, or multi-step reasoning pipelines is a plus.

  • Experience with enterprise data infrastructure, data pipelines, or data integration platforms is a plus.

Location

This role is on-site in San Mateo, California. Visa sponsorship is not available.

Skills

See also

DevOps jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available