Point your AI agent at freehire and let it find you a job.

Get the CLI →

Data Engineering Lead - Data Quality Systems

NewBe an early applicant

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Data Engineering Lead - Data Quality Systems based in India.

Lead the engineering of data quality systems that determine whether billions of records can be trusted by customers.
You will combine deep hands-on engineering with technical leadership, spending approximately 80% of your time building and 20% leading a small team.
Your scope will include verification pipelines, anomaly detection, scoring frameworks, LLM evaluation, and automated release gates.
You will tackle complex data-quality challenges across multiple markets and large-scale production environments.
The role offers significant ownership, with greenfield opportunities to establish frameworks and engineering standards from the ground up.
You will work in an AI-native environment where agentic development, evaluation, observability, and automation are core engineering practices.
This is an ideal opportunity for a technically strong leader who wants direct ownership of a critical data trust layer while remaining deeply involved in the code.

Accountabilities

  • Architect and build continuous data-quality systems, including verification, sampling, scoring, and reconciliation pipelines operating across billions of company and people records.
  • Design reusable frameworks, abstractions, and technical specifications that allow engineers to create quality checks efficiently, reliably, and consistently.
  • Build evaluation harnesses for LLM-powered validation and extraction, including labeled evaluation sets, precision/recall measurement, judge calibration, prompt versioning, and model-drift detection.
  • Establish pre- and post-production release gates that identify and prevent poor-quality data from reaching customers, supported by effective failure analysis and triage tooling.
  • Investigate large-scale data-quality incidents, identify root causes, implement corrective solutions, and convert recurring failures into permanent automated checks.
  • Lead a team of 3–5 Applied AI Engineers through technical direction, code reviews, pairing, mentoring, and development of end-to-end ownership.
  • Set and maintain a high technical standard while remaining approximately 80% hands-on in engineering and architecture.
  • Apply sound judgment when choosing between deterministic rules and LLM-based validation, using structured rules where appropriate and semantic models where they add value.
  • Operate LLM-based quality systems as production infrastructure, with appropriate evaluation, traceability, prompt and model versioning, cost controls, and performance monitoring.
  • Contribute to an AI-native engineering culture based on agentic development, automated evaluation, logged traces, AI-assisted review, and reusable workflow specifications.
  • Establish scalable engineering practices in a lean environment characterized by high ownership, minimal process overhead, and frequent production releases.
  • Requirements

    • 7+ years of experience building production-grade data systems in business-critical environments, including systems that operate reliably at significant scale.
    • Demonstrated experience working with billions of data rows and designing quality controls that remain performant and dependable at large scale.
    • Proven track record of building data-quality systems and frameworks, such as validation engines, anomaly detection, scoring systems, sampling strategies, or reconciliation mechanisms against trusted data.
    • Experience designing evaluation or test harnesses that are used by other engineers and can support systematic measurement of quality.
    • Previous experience providing technical leadership to engineers, including code reviews, technical direction, pairing, mentoring, and hands-on delivery.
    • Strong Python development skills and advanced SQL expertise, with an understanding of performance optimization, concurrency, and large-scale data transformations.
    • Practical experience operating LLMs as production systems, including evaluation sets, versioned prompts, trace logging, cost controls, and debugging model judges against precision and recall.
    • Proven experience using agentic development environments such as Claude Code, Cursor, or equivalent tools to build and ship production software.
    • Strong technical judgment regarding when to use deterministic rules versus LLM-based semantic evaluation, with the ability to clearly justify architectural decisions.
    • Experience with B2B data, including firmographics, people data, entity resolution, or registry matching across multiple markets, is highly valued.
    • Familiarity with cloud data platforms such as Snowflake, Databricks, or Redshift, together with AWS-based pipeline deployment, is advantageous.
    • Production-scale experience with Airflow or an equivalent orchestration platform is a plus.
    • Knowledge of vector databases, embeddings, retrieval patterns, matching, or deduplication is desirable.
    • Startup or scaleup experience, particularly in environments where engineering standards and frameworks had to be established from the ground up, is highly valued.
    • Strong ownership, judgment, adaptability, and communication skills suited to a fast-moving, autonomous, and highly collaborative engineering environment.
    • Benefits

      • Fully remote position based in India.
      • Competitive base salary aligned with the seniority and technical scope of the role.
      • Meaningful equity participation and the opportunity to share in the organization's long-term growth.
      • Significant technical ownership over a critical data-quality and trust layer.
      • Greenfield engineering opportunities to define frameworks, standards, validators, evaluation systems, and release gates.
      • Exposure to frontier engineering challenges involving LLM evaluation, model drift, agentic development, anomaly detection, and large-scale data quality.
      • Opportunity to lead a small, senior engineering team while remaining deeply hands-on technically.
      • Lean, high-autonomy environment with minimal management layers and strong end-to-end ownership.
      • AI-native engineering practices, with agentic development, evaluations, traces, and AI-powered review integrated into everyday workflows.
      • Fast release cycles and the opportunity to make visible contributions across multiple international markets.
      • Equal-opportunity environment that values diverse perspectives and inclusive collaboration.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available