Senior Automation QA (JavaScript) Engineer
Summary
Senior automation QA engineer (India-based) building the quality foundation for an enterprise AI Agent Development Platform: designing Playwright/TypeScript test suites, testing Temporal workflows and non-deterministic LLM behavior via statistical and semantic assertions, and wiring tests into CI/CD release gates.
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Automation QA (JS) Engineer based in India.
As a Senior Automation QA Engineer, you will help establish the quality foundation for an enterprise-grade AI Agent Development Platform. You will design and own automation across frontend, backend, long-running workflows, multi-tenant infrastructure, security boundaries, and performance requirements. Because the platform relies on non-deterministic AI behavior, you will develop sophisticated testing approaches based on statistical assertions, semantic evaluation, and adversarial scenarios rather than simple expected-output checks. Your work will directly influence release readiness, client milestone sign-offs, and production reliability. You will collaborate closely with AI engineers, platform engineers, and evaluation specialists to embed quality and testability from the earliest design stages. This is a high-impact opportunity to shape reusable quality engineering practices for rapidly evolving AI-powered systems.
Accountabilities:
You will own the automation strategy across the platform, building reliable, reproducible, and scalable test infrastructure that provides defensible evidence of product quality and production readiness.
- Design, build, and maintain automated test suites covering frontend and backend functionality.
- Develop contract tests that protect the agent-developer experience as SDKs and platform capabilities evolve.
- Automate testing of long-running Temporal workflows, including worker failures, provider outages, retry storms, and other injected failures.
- Verify durable execution guarantees, including zero lost workflow runs, idempotency, recovery behavior, and correct compensation or saga execution.
- Automate multi-tenant isolation testing, including cross-tenant data access, configuration leakage, and cost-attribution correctness.
- Validate agent execution sandboxing and egress controls through negative testing of unauthorized tool calls and network destinations.
- Design testing strategies for LLM-driven and other non-deterministic behavior using statistical assertions, repeated-run consistency, semantic similarity scoring, and confidence thresholds.
- Identify and quarantine flaky tests while distinguishing expected model variance from genuine regressions.
- Build and maintain adversarial test corpora covering prompt injection, tool-call hijacking, data-exfiltration scenarios, and other AI security risks.
- Test Human-in-the-Loop escalation mechanisms to ensure confidence thresholds trigger correctly and gated or irreversible actions cannot proceed without authorization.
- Partner with evaluation teams on golden datasets, synthetic-data pipelines, and CI integrations that make AI quality measurable and repeatable.
- Automate verification of performance and reliability requirements, including p95 invocation overhead, concurrency targets, queue backpressure, and LLM provider failover.
- Use Langfuse, OpenTelemetry, and tracing data to validate trace completeness, token and cost accounting, and anomalous-run alerting.
- Integrate automated tests into CI/CD pipelines as hard release gates with per-metric regression detection.
- Produce auditable and reproducible test-evidence packages that support client milestone sign-offs.
- Collaborate with AI and platform engineers during system design to establish acceptance criteria, testability, and observability requirements.
- Maintain test environments, mocked LLM and provider layers, and synthetic data generators to keep testing efficient, deterministic where possible, and cost-effective.
- Mentor mid-level QA engineers and develop reusable AI testing frameworks, patterns, and best practices.
- Contribute to the broader quality engineering practice through knowledge sharing and technical enablement.
- 8+ years of experience in test automation, API testing, and building reusable test frameworks adopted by other engineers.
- Strong proficiency in Playwright and TypeScript is mandatory.
- Extensive experience developing Playwright/TypeScript-based automation frameworks and integrating them into secure CI/CD environments.
- Hands-on experience testing LLM-powered or other non-deterministic systems, including statistical assertions, semantic scoring, and model-variance management.
- Deep understanding of Large Language Model behavior, prompt sensitivity, and common RAG failure modes.
- Strong experience testing asynchronous and event-driven systems, including failure injection, idempotency verification, and eventual-consistency assertions.
- Experience with Temporal or a similar workflow orchestration engine is highly desirable.
- Strong CI/CD expertise, particularly with GitHub Actions or equivalent platforms, and experience establishing automated release gates.
- Solid SQL and data-validation skills, including the ability to build synthetic test-data pipelines and assess dataset quality.
- Strong analytical skills and experience interpreting evaluation scores to provide actionable feedback to AI engineers and support prompt tuning.
- Security-focused and adversarial mindset, with familiarity with prompt-injection techniques or strong application security testing experience.
- Ability to determine appropriate testing depth under delivery and milestone pressure and confidently communicate quality evidence to technical and non-technical stakeholders.
- Strong understanding of automation strategy, software quality principles, debugging, and testability.
- Ability to work collaboratively with AI engineers, platform engineers, product teams, and evaluation specialists.
- Strong communication, mentoring, and knowledge-sharing skills.
- Experience working in fast-paced, cross-functional engineering environments.
- Experience with performance and load-testing tools such as k6 or Locust.
- Experience testing multi-tenant SaaS isolation and security boundaries.
- Familiarity with AI evaluation and observability platforms such as Langfuse, LangSmith, Promptfoo, DeepEval, or Arize Phoenix.
- Experience producing compliance evidence for frameworks such as SOC 2 or GDPR.
- Experience with accessibility testing, including Section 508 or WCAG.
- Familiarity with Commercial Real Estate workflows, lease accounting, CAM reconciliation, or document-extraction systems.
- Experience using LLMs to generate test cases, adversarial payloads, or synthetic documents.
- Experience testing high-concurrency background processing within agentic systems.
- Experience with qTest or similar test-management platforms.
- Remote work opportunity based in India.
- Opportunity to work on an enterprise-grade AI and Agent Development Platform.
- High-impact role where automated testing directly influences release gates and client milestone decisions.
- Exposure to advanced AI/LLM testing, agent evaluation, distributed workflows, security testing, and observability.
- Opportunity to work with technologies such as Playwright, TypeScript, Temporal, Langfuse, OpenTelemetry, CI/CD platforms, and modern cloud infrastructure.
- Strong professional community and collaborative, open-door working environment.
- Access to internal meetups, conferences, workshops, Udemy, language courses, and company-paid certifications.
- Opportunities for internal mobility and exposure to diverse technology domains and large-scale projects.
- Company-paid medical insurance.
- Mental health support.
- Financial and legal consultation support.
- Opportunity to collaborate with global engineering teams and work on technology with international impact.
Requirements:
The ideal candidate is an experienced automation engineer with strong framework-building capabilities, advanced knowledge of AI/LLM testing, and the ability to assess complex systems from both reliability and adversarial perspectives.
Preferred qualifications include:
Benefits:
Skills
As published by lever
Resume/CV, Full name, Email, Phone, Current location, Current company