Senior AI Engineer - Observability
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior AI Engineer - Observability based in United States.
This is a hands-on senior engineering opportunity focused on making AI-powered products reliable, measurable, and production-ready.
You’ll build and improve LLM, RAG, semantic search, and agentic experiences while helping define how AI quality is measured.
A major focus will be designing evaluation pipelines, golden datasets, regression testing, and meaningful quality gates.
You’ll also own observability across AI systems, tracking quality, latency, cost, reliability, safety, and real-world user feedback.
Working closely with engineering, product, QA, security, and business teams, you’ll turn production signals into actionable improvements.
You’ll help establish reusable standards, tooling, and engineering practices that enable teams to ship AI faster and with greater confidence.
This role is ideal for an experienced software engineer who combines strong technical depth with product judgment and a rigorous quality mindset.
Accountabilities:
- Partner with product engineering teams to design, build, and enhance AI-powered features using LLMs, RAG, semantic search, and agentic workflows.
- Contribute directly to production codebases using Python, C#/.NET, or related technologies, while improving prompts, retrieval, context assembly, ranking, grounding, and response-generation patterns.
- Evaluate architecture and technology trade-offs across AI quality, latency, cost, privacy, maintainability, and operational reliability.
- Design and implement AI evaluation pipelines, scoring methodologies, regression frameworks, and quality gates that establish whether features are ready for production.
- Build and maintain versioned golden datasets covering real-world use cases, edge cases, failure modes, and customer-critical workflows.
- Apply LLM-as-judge, heuristic, human-feedback, and task-specific evaluation approaches, while assessing whether evaluation methods accurately measure meaningful outcomes.
- Instrument LLM interactions, RAG pipelines, tool calls, and agent workflows using observability platforms and establish dashboards, alerts, and production monitoring.
- Track and analyze latency, token usage, cost, retrieval quality, groundedness, safety signals, failure modes, and user feedback to identify regressions and improvement opportunities.
- Develop reusable libraries, SDKs, templates, documentation, and reference implementations for AI instrumentation, evaluation, tracing, cost reporting, and incident response.
- Coach product teams on AI quality, observability, evaluation design, and reliable release practices, helping establish consistent standards across product areas.
- Support responsible AI delivery by ensuring appropriate handling of sensitive data and alignment with security, privacy, SOC 2, ISO 27001, and data-residency requirements.
- Contribute to model and provider evaluations, migrations, fallback strategies, rollout plans, and ongoing optimization of AI-powered systems.
- 5+ years of professional software engineering experience building and supporting production systems.
- Hands-on experience developing or operating LLM-powered features, RAG systems, AI workflows, or comparable AI applications.
- Strong programming capabilities in Python, C#/.NET, or both, with the ability to contribute effectively to production engineering environments.
- Practical understanding of prompts, embeddings, vector search, retrieval quality, orchestration patterns, grounding, and modern LLM application architecture.
- Demonstrated experience designing evaluations, quality metrics, regression frameworks, or benchmarking approaches, with strong judgment around whether an evaluation measures what truly matters.
- Strong observability fundamentals covering tracing, logging, metrics, alerting, production debugging, and analysis of system behavior.
- Experience with CI/CD, Git-based workflows, cloud environments, and production release practices; experience with Azure DevOps is advantageous.
- Familiarity with AI-assisted development tools such as Claude Code, PlayerZero, or comparable technologies.
- Experience with AI observability and LLMOps platforms such as Arize, Langfuse, LangSmith, W&B, Humanloop, or Helicone is preferred.
- Knowledge of OpenTelemetry, Azure Monitor, Application Insights, or similar observability technologies is beneficial.
- Experience with vector search technologies such as Azure AI Search, Pinecone, Qdrant, Weaviate, or pgvector is advantageous.
- Familiarity with frameworks such as Semantic Kernel, LangChain, LlamaIndex, AutoGen, or related AI orchestration technologies is a plus.
- Experience with LLM-as-judge evaluation, RAG evaluation, semantic similarity, hallucination detection, groundedness scoring, human-feedback workflows, or dedicated evaluation frameworks is preferred.
- Experience with A/B testing, online experimentation, product analytics, regulated environments, PII handling, or data-residency requirements is beneficial.
- Strong communication, collaboration, decision-making, and stakeholder management skills, with the ability to translate ambiguous AI behavior into concrete engineering actions.
- A strong product mindset, curiosity about real-world AI behavior, adaptability, accountability, and a genuine commitment to quality and continuous improvement.
- Fully remote position within the United States.
- Company-provided equipment, including laptop and required software.
- Comprehensive medical and prescription drug plan options, along with dental and vision coverage.
- Employer contribution to a Health Savings Account (HSA) for eligible employees enrolled in a high-deductible healthcare plan.
- Medical and dependent-care Flexible Spending Accounts.
- Basic life insurance valued at $50,000 or one times annual salary, whichever is greater.
- Short-term disability, long-term disability, and Accidental Death & Dismemberment coverage at no employee cost.
- 401(k) retirement savings plan with automatic enrollment after the initial eligibility period and a company match of up to 4%.
- Paid time off and holidays.
- Employee support programs and benefits designed to promote well-being, flexibility, and work-life balance.
- A collaborative, growth-oriented environment where employees can influence engineering standards and contribute to meaningful AI product development.
- Opportunities to shape reusable AI engineering practices, observability standards, and quality frameworks across multiple product teams.
Requirements
Benefits
As published by lever
Resume/CV, Full name, Email, Phone, Current location, Current company