QA Engineer (C1) - Agentic AI

Key Responsibilities

  • 1. Agentic AI Testing, Evaluation & Automation
  • Define and execute testing strategies for LLM-based, multi-agent, RAG, and Agentic AI systems.
  • Validate autonomous agent behavior, reasoning, memory, tool usage, API/database integrations, and end-to-end workflows.
  • Evaluate AI outputs for accuracy, relevance, groundedness, consistency, completeness, toxicity, bias, hallucination risk, and guardrail compliance.
  • Define AI quality KPIs such as hallucination rate, groundedness score, agent success rate, task completion rate, response relevancy, latency, cost efficiency, and user satisfaction.
  • Build automated evaluation pipelines, quality scoring mechanisms, dashboards, and CI/CD-integrated quality gates.
  • Develop reusable test harnesses, simulators, and benchmarking frameworks to compare models, prompts, and agent configurations.
  • 2. Team Leadership & Capability Building
  • Build and lead a team of Agentic AI Quality Engineers.
  • Define team structure, testing standards, best practices, and governance models.
  • Mentor QA engineers in AI testing methodologies, evaluation techniques, and automation frameworks.
  • Drive innovation and adoption of emerging AI testing tools and technologies.
  • Collaborate with Product, Engineering, Data Science, and AI Research teams to improve overall AI quality.
  • 3. Reporting & Stakeholder Management
  • Provide quality assessments and recommendations to leadership and stakeholders.
  • Present testing outcomes, risk assessments, KPI trends, and model evaluation reports.
  • Drive quality governance for Agentic AI initiatives across the organization.
  • Ensure traceability of testing activities, evaluation criteria, and quality benchmarks.

Required Skills & Experience

Technical Skills

  • 7–12 years of experience in Software Testing, Quality Engineering, or Test Automation.
  • Minimum 3+ years of hands-on experience in GenAI, LLM Testing, Agentic AI Testing, or AI Quality Engineering.
  • Strong understanding of LLMs, AI agents, RAG, prompt validation, tool calling, agent memory, MCP, and multi-agent orchestration.
  • Experience defining AI quality metrics, evaluation methodologies, benchmarking frameworks, and model comparison approaches.
  • Hands-on automation experience with Python, Playwright, Pytest, API automation, test framework development, and CI/CD quality gates.
  • Experience with AI evaluation frameworks such as DeepEval, Ragas, LangSmith, OpenAI Evals, or equivalent tools.
  • Exposure to cloud platforms such as Azure, AWS, or GCP.

Soft Skills

  • Strong communication and stakeholder management skills.
  • Analytical mindset with strong problem-solving ability.
  • Self-driven, outcome-oriented, and capable of leading multiple initiatives in a fast-evolving AI ecosystem.
  • Experience testing enterprise Agentic AI platforms and autonomous AI systems.
  • Hands-on exposure to frameworks or tools such as LangGraph, CrewAI, AutoGen, Semantic Kernel, Microsoft Copilot Studio, TruLens, or Promptfoo.
  • Exposure to AI observability, monitoring, model governance, responsible AI, and AI safety practices.
  • Experience building AI quality dashboards and KPI reporting systems.
  • ISTQB, AI Testing, GenAI, Azure AI, AWS AI, or equivalent certifications