freehire launches on Product Hunt on 26 August.

Follow →

Staff Applied AI Engineer, Product & Agent Performance

Summary

Build and refine AI agents for healthcare workflows, focusing on reliability, safety, and production readiness through prompting, retrieval, evaluation, and escalation strategies.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Applied AI Engineer, Product & Agent Performance based in the United States.

This is a staff-level individual contributor role focused on making AI agents reliable, measurable, and production-ready in complex healthcare workflows.
You will shape agent behavior across prompting, retrieval, context, memory, tool use, evaluation, and human escalation.
Your work will directly influence how AI systems perform across real-world, high-impact use cases serving healthcare organizations.
You will establish rigorous evaluation practices that identify failure modes, quantify risk, and guide model and product release decisions.
The role combines deep technical execution with product judgment, requiring you to translate production evidence into practical improvements.
You will work closely with Product and Engineering to build AI systems that are accurate, steerable, transparent, cost-efficient, and trustworthy.
This is an opportunity to define a scalable product-layer AI performance discipline in an environment where safety and responsible innovation matter.

Accountabilities

  • Design, implement, and continuously improve agent behavior across live, long-horizon, multi-turn, and multi-agent workflows.
  • Architect retrieval and context strategies that deliver the right source data to models in the right structure while keeping agents grounded in reliable information.
  • Design memory and state-management approaches for multi-turn and multi-agent experiences, determining what information should be retained, summarized, or discarded.
  • Develop prompt and context templates using few-shot examples, structured formats, reasoning scaffolding, and other techniques to create consistent agent behavior.
  • Improve agent performance through experimentation with prompting, tool-use strategies, retrieval, and context construction rather than relying on assumptions.
  • Build production-representative evaluation suites and regression checks to measure accuracy, reliability, regressions, failure modes, edge cases, latency, and cost.
  • Create evaluation rubrics, quality heuristics, and performance thresholds that account for the severity and business or safety impact of failures, not simply their frequency.
  • Design and validate escalation mechanisms that route uncertain or high-risk cases to human review while maintaining safe and consistent behavior.
  • Establish cost-aware approaches to AI performance, balancing accuracy, reliability, latency, context efficiency, and tool-call usage.
  • Baseline existing behavior, conduct comparative evaluations, and assess model or system changes to make evidence-based go/no-go recommendations before customer release.
  • Maintain product-level AI documentation, including model cards, intended-use guidance, limitations, known failure modes, and performance information.
  • Partner closely with Product and Engineering to ensure agentic systems are not only capable but also steerable, trustworthy, transparent, and scalable.
  • Translate production failures and performance evidence into clear diagnoses, experiments, fixes, and actionable recommendations for cross-functional teams.
  • Requirements

    • 8+ years of production software engineering experience, including at least 3 years of hands-on ownership of ML, LLM, or agentic systems in production.
    • Professional experience working with AI systems in healthcare, finance, or another regulated environment where reliability, safety, and transparency are important.
    • Demonstrated ability to diagnose agent failures and determine whether improvements should come from instructions, retrieval, context, memory, tool use, or other system components.
    • Strong understanding of how to evaluate AI failures based on severity, risk, and cost rather than frequency alone.
    • Hands-on experience designing and implementing RAG architectures and production-grounded evaluation frameworks.
    • Experience developing fallback mechanisms, human-in-the-loop workflows, escalation logic, or comparable safety mechanisms for automated systems.
    • Practical familiarity with AWS AI/ML services, including Bedrock and SageMaker, sufficient to build, test, and evaluate AI systems in production environments.
    • Strong software engineering foundations and the ability to move comfortably between technical implementation, experimentation, evaluation, and product-level decision-making.
    • Evidence-driven judgment and confidence to challenge launch decisions when performance or safety standards are not met.
    • A strong builder mentality, with the ability to move quickly from identifying a production problem to designing an experiment, validating a solution, and implementing a fix.
    • Equivalent practical experience demonstrating staff-level technical depth may be considered in place of specific educational credentials.
    • Experience applying AI to healthcare data or workflows, particularly where calibrated uncertainty and transparency affect clinicians, care teams, or patients, is highly valued.
    • Experience with long-horizon, multi-turn, or multi-agent systems and product-level AI documentation such as model cards is a plus.
    • Benefits

      • Competitive compensation: $175,000–$200,000 per year.
      • Meaningful technical ownership: Define how AI agent performance, safety, reliability, and production readiness are measured.
      • High-impact scope: Influence prompts, retrieval, context, memory, evaluations, escalation patterns, and AI product decisions at scale.
      • Mission-driven work: Contribute to technology designed to improve healthcare delivery and patient outcomes.
      • Remote-friendly culture: Flexible working environment designed to support distributed collaboration.
      • Professional development: Employee-driven programs and initiatives supporting personal and career growth.
      • Collaborative community: Work alongside a diverse, talented, energized, and purpose-driven team.
      • Cross-functional exposure: Partner closely with Product and Engineering while translating production evidence into improvements across AI-powered workflows.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available