Software Quality Engineer – Conversational AI / LLM Testing

Summary

This role involves designing and executing quality strategies for conversational AI and LLM-based applications within the airline and travel industry. The engineer will use Python to automate testing, evaluate non-deterministic AI outputs, and ensure the reliability of booking and passenger service workflows.

Software Quality Engineer – Conversational AI / LLM Testing

Experience\: 4-9 Years
Location \: Noida / Kolkata
Domain \: Airlines / Travel Technology

Job Summary

We are seeking a highly skilled Software Quality Engineer with hands-on experience in testing Conversational AI, LLM-based applications, chatbots, virtual assistants, and Agentic AI systems. The ideal candidate will have a strong automation and quality engineering background, deep understanding of how modern conversational systems operate, and solid experience within the Airline or Travel domain, particularly around Booking, Post-Booking, and Check-In workflows.

This role involves validating AI-driven customer conversation journeys, ensuring response quality, evaluating non-deterministic outputs, and establishing robust testing strategies for intelligent systems.

Key Responsibilities

  • Own and drive the quality strategy for conversational AI products and customer-facing virtual assistants.
  • Design, execute, and automate test scenarios for LLM-powered applications, chatbots, voice assistants, RAG-based systems, and agentic workflows.
  • Validate conversational flows across booking, post-booking servicing, and check-in experiences within airline and travel ecosystems.
  • Develop and maintain automated test frameworks using Python.
  • Create test datasets, golden datasets, and evaluation benchmarks to measure AI system performance.
  • Design evaluation mechanisms for non-deterministic AI outputs using\:
    • Rubric-based scoring
    • LLM-as-a-Judge approaches
    • Human review workflows
    • Statistical quality assessment methods
  • Validate prompt behavior, context retention, retrieval accuracy, tool/function calling, and response consistency.
  • Identify conversational AI failure modes such as hallucinations, context loss, prompt injections, retrieval issues, and reasoning errors.
  • Collaborate with Product Managers, Data Scientists, AI Engineers, and Developers to improve system quality and reliability.
  • Participate in release validation, regression testing, performance testing, and production quality monitoring.

Mandatory Skills & Qualifications

Quality Engineering

  • 4+ years of experience in Software Quality Engineering, QA Automation, or SDET roles.
  • Proven experience owning the quality of a product or platform from testing strategy to production validation.
  • Strong understanding of software testing methodologies, automation frameworks, and quality metrics.

Python Automation

  • Strong scripting and automation expertise in Python.
  • Experience developing automated test suites, validation scripts, and data-driven testing frameworks.

Conversational AI / LLM Testing

Hands-on experience testing one or more of the following\:

  • Large Language Model (LLM) applications
  • Conversational AI platforms
  • Chatbots and Virtual Assistants
  • Voice Assistants
  • Retrieval Augmented Generation (RAG) solutions
  • Agentic AI applications

AI/LLM Evaluation Expertise

Practical understanding of\:

  • Prompt engineering principles
  • Context windows and memory handling
  • Function/Tool Calling
  • Retrieval pipelines
  • LLM evaluation methodologies
  • AI quality metrics and benchmarking

Experience with\:

  • Golden datasets
  • Human evaluation workflows
  • Statistical validation approaches
  • Non-deterministic output testing

Domain Knowledge

Strong experience in Airline or Travel domain, including\:

  • Flight Booking journeys
  • Reservation Management
  • Post-Booking Services
  • Check-In processes
  • Passenger servicing workflows

Preferred Skills

  • Experience with OpenAI, Azure OpenAI, Anthropic, Gemini, or similar AI platforms.
  • Understanding of NLP concepts and conversational design.
  • Exposure to API testing and automation tools.
  • Experience with CI/CD pipelines and test automation integration.
  • Knowledge of cloud platforms (Azure, AWS, GCP).
  • Familiarity with observability and monitoring tools for AI applications.

See also

AI Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available