Point your AI agent at freehire and let it find you a job.

Get the CLI →

jobgether

NewBe an early applicant

Data Engineer - Data Platform

Discussion

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Data Engineer - Data Platform based in India.

This is a hands-on data engineering role focused on building the pipelines and infrastructure behind a large-scale B2B data platform. You will transform billions of raw records into clean, reliable, and modelled datasets that power products, analytics, and AI systems. A key part of the role is integrating LLMs directly into production data pipelines for extraction, enrichment, entity resolution, and semantic validation. You will help make probabilistic AI components reliable through measurable evaluations, regression testing, observability, cost controls, and robust fallback mechanisms. Working with modern data technologies such as Python, SQL, dbt, Airflow, Snowflake, and AWS, you will own systems end to end. This is an AI-native, fast-moving environment with small senior teams, significant autonomy, and a strong emphasis on shipping high-quality engineering work.

Accountabilities:

  • LLM-Powered Data Pipelines: Build and operate production systems that use LLMs for extraction, enrichment, entity resolution, and semantic validation across millions of records, with structured outputs, retries, and human-review fallbacks.
  • Evaluation & Quality: Develop labelled evaluation datasets, scoring systems, and regression suites that measure LLM performance and gate prompt or model changes using metrics such as precision and recall.
  • Rules vs. AI: Determine when deterministic approaches such as SQL, regex, dbt tests, and data contracts are more appropriate than LLM-based decisions, applying sound engineering judgment to balance reliability and flexibility.
  • Cost & Model Management: Establish token budgets, route workloads between models based on complexity and cost, and monitor model behavior for drift following vendor or model updates.
  • Observability: Ensure production LLM calls and data pipelines are fully observable, capturing prompt versions, models, costs, latency, decisions, logs, metrics, and alerts to identify issues before they cascade.
  • Core Data Engineering: Build and maintain Airflow orchestration, modular dbt models across staging and mart layers, and high-performance Snowflake data solutions operating across billions of records.
  • Infrastructure: Support scalable AWS infrastructure and infrastructure-as-code practices for reliable data platform operations.
  • Entity Resolution: Develop matching and deduplication capabilities using embeddings, retrieval, and other techniques to resolve company and people entities across multiple international markets.
  • Reliability & Ownership: Own production pipelines end to end, considering upstream dependencies, downstream impact, reliability, performance, and cost before making changes.
  • Engineering Innovation: Contribute to an AI-native development environment where AI coding tools, automated testing, and AI-assisted engineering practices are part of the standard workflow.
  • Requirements

    • Data Engineering Experience: 3+ years of experience in data engineering or data infrastructure, with demonstrated ownership of production pipelines from end to end and 2+ years of hands-on Snowflake experience.
    • Python & SQL: Strong production-grade Python and SQL skills, with the ability to write efficient, performance-aware solutions at very large scale.
    • Production LLM Experience: Proven experience integrating LLMs into production data pipelines or quality systems for extraction, enrichment, validation, or similar use cases, including structured outputs and labelled evaluation datasets.
    • Evaluation Harnesses: Experience building or maintaining evaluation and testing harnesses for LLM outputs, with the ability to explain how they measure quality and identify model failures.
    • Engineering Judgment: Strong ability to distinguish between deterministic rules and LLM-based approaches, selecting the simplest reliable solution for each problem.
    • dbt & Airflow: Production experience with dbt and Airflow, including modular and tested data models and resilient DAGs that recover gracefully from failures.
    • Cloud Data Warehousing: Strong experience with cloud data warehouses, preferably Snowflake, including schema design, query optimization, performance management, and cost control.
    • AI-Assisted Development: Daily experience using AI coding tools such as Claude Code, Cursor, or equivalent, with demonstrable examples of shipped production work.
    • Systems Thinking: Strong ownership mindset and ability to assess upstream dependencies, downstream effects, reliability, and operational risks when designing or modifying systems.
    • Advanced AI & Data Skills: Experience with LLM evaluation frameworks, tracing, embeddings, vector search, fuzzy matching, fine-tuning, or model distillation is highly valued.
    • Cloud & Distributed Systems: Experience with AWS services such as S3, Lambda, Glue, ECS, or RDS, as well as PostgreSQL, is advantageous.
    • Big Data & Streaming: Experience with Spark/PySpark or streaming technologies such as Kafka, Kinesis, or Snowpipe Streaming is a plus.
    • Security & Compliance: Exposure to data privacy and compliance requirements such as GDPR, SOC 2, or CCPA is beneficial.
    • Communication & Autonomy: Comfortable working independently in a lean, senior environment with minimal process, high ownership, and distributed collaboration.
    • Benefits

      • Fully Remote: Work remotely from India with flexibility and autonomy in your working environment.
      • Competitive Compensation: Receive a competitive base salary along with meaningful equity participation.
      • Equity Opportunity: Share in the long-term upside of the business through an equity component.
      • AI-Native Environment: Work with LLMs as production infrastructure rather than experimental tools, solving challenging problems involving evaluation, observability, cost, and model reliability.
      • Technical Ownership: Take significant ownership of systems that operate at scale across billions of records and multiple international markets.
      • Greenfield Engineering: Help build evaluation frameworks, drift detection, cost controls, and other foundational capabilities for production AI data pipelines.
      • High Impact: Work in a small senior team where your engineering decisions can directly influence products, analytics, AI systems, and customers.
      • Continuous Delivery: Operate in a fast-moving environment with frequent releases, high autonomy, and a strong focus on practical engineering outcomes.
      • Additional Benefits: Healthcare, leave, and other employee benefits may be provided according to the partner company’s applicable policies.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

Skills

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available