Senior Site Reliability / DevOps Engineer– AI Products
Summary
Senior SRE Engineer at Navan in Tel Aviv builds and maintains reliable, scalable platforms for AI-powered travel and expense systems, focusing on observability, automation, and AI provider integrations.
At Navan, "It's all about the user. All of them." We're passionate about providing a seamless one-stop experience for business travelers, no matter how they travel, where they stay, or where they're going. Navan is building cutting-edge solutions at the intersection of travel, expense, payments, and AI. As a leader in the AI for Travel domain, we are using intelligent, practical AI experiences to make business travel simpler, faster, and more reliable for travelers, travel managers, finance teams, and support teams.
We are seeking a Senior Site Reliability / DevOps Engineer to ensure the scalability, performance, and reliability of our user-facing generative AI features.
In this role, you will bridge the gap between traditional infrastructure and cutting-edge machine learning.
You will build and maintain the high-throughput, low-latency systems required to serve AI models directly to millions of users.
This position is based out of our new Tel Aviv office.
What You'll Do:
- Infrastructure Ownership: Design, build, and scale the infrastructure hosting our user-facing AI applications and inference engines.
- Performance Optimization: Optimize system latency, specifically targeting Time-to-First-Token (TTFT) and total round-trip time for user requests.
- GPU & Resource Orchestration: Manage and scale GPU clusters within Kubernetes to maximize utilization and minimize operational costs.
- Resiliency & Fallbacks: Build robust fallback systems, circuit breakers, and rate-limiting infrastructure to handle upstream LLM API failures and traffic spikes.
- Monitoring & Observability: Implement deep observability for AI workloads, tracking custom metrics like token usage, model drift, and GPU memory saturation.
What We're Looking For:
- SRE Fundamentals: 4+ years of experience in SRE, DevOps, or Production Engineering roles supporting high-traffic, user-facing applications.
- LLMOps / AI Infrastructure Expertise: Experience working with AI workloads (such as serving models using vLLM, Server TGI or working with cloud providers like Bedrock, OpenAI, etc.).
- Container Orchestration: Strong expertise in Kubernetes (EKS, GKE, or AKS) and infrastructure-as-code (Terraform).
- AI/ML Ecosystem: Hands-on experience with inference servers (e.g., vLLM, TGI) and vector databases (e.g., Pinecone, Milvus, Qdrant).
- Programming: Proficiency in Python and Go for automation, tooling, and backend optimization.
- Cloud Architecture: Deep experience managing cloud compute resources, specifically specialized GPU instances
- Models AI and Code:
Ability to build automation processes that not only update code versions, but also support testing and safe deployment of new models (Shadow Deployments, Canary releases for models) Product thinking and user orientation (User-Facing) - Advanced Observability:
Mastery of tools like OpenTelemetry, Prometheus, Datadog or Grafana, with the ability to trace agent-based systems and complex model calls. - Cost & Capacity Optimization:
Ability to manage the high costs of GPU/Inference in a productive architecture without compromising availability or performance. - Empathy for the end-user experience:
Understanding that every millisecond of latency or error in the stream directly impacts customer retention. - Preferred Qualifications
Experience building semantic caching layers to reduce LLM API costs.
Active contributor to open-source LLMOps or MLOps projects. (edited)
Navan uses AI-assisted Automated Employment Decision Tool (Metaview) to assist with evaluating resumes against job qualifications for this role. All final decisions are made by human recruiters and hiring managers.
Human oversight: Metaview does not automatically reject candidates or make final hiring decisions. Our recruiters and hiring managers review all outputs and make the final hiring decision regarding every application.
-
Your rights: If you prefer to have your application reviewed without AI assistance, you may request a human evaluation by entering your email here. Your decision to do so will not affect how your candidacy is evaluated.
Please refer to our Candidate Privacy Notice for more information about our processing of personal data, and your rights.
As published by greenhouse
First Name, Last Name, Email, Phone, Resume/CV, Cover Letter, Location
- Preferred First Name optional
- Website optional
- How did you hear about this job?
- Are you currently eligible to work in the country outlined for this position, and authorized to work for Navan on an ongoing indefinite basis? choose one
- Will you now or in the future require sponsorship by Navan to attain or maintain your employment eligibility? choose one
- LinkedIn Profile
- This role operates on a hybrid work model, requiring you to work from one of our global offices 4 days per week (including 2 Fridays each month). Are you comfortable with this policy, and are you currently located at, or willing to relocate to, one of our office locations? choose one
- Have you ever been employed by, applied to, or are you currently employed by Navan or any of its affiliated or group companies (Reed & Mackay)? choose one
