Point your AI agent at freehire and let it find you a job.

Get the CLI →

Dataeconomy

NewBe an early applicant

Data Engineer — ML Training Data Pipeline

Posted
Discussion

Summary

Builds and maintains AWS data pipelines that transform raw production traces into high-quality training datasets for LLM fine-tuning — handling ingestion, deduplication, format conversion, quality filtering, and train/test splitting at scale. Core stack is Python (pandas, pyarrow), JSONL processing, HuggingFace Datasets, and AWS S3/EC2, based in Hyderabad or Pune.

Job Title:Data Engineer - ML Training Data Pipeline
Notice period: 0-30 Days
Experience : 5+ Years
Location: Hyderabad OR Pune

We are looking for Data Engineer - ML Training Data Pipeline who can Build and maintain the data pipeline that transforms raw production traces into high-quality training datasets for LLM fine-tuning-ingestion, deduplication, format conversion, quality filtering, and train/test splitting at scale on AWS.


What We Expect:

  • Build end-to-end data pipelines: raw trace ingestion → dedup → format conversion → quality gating → training-ready datasets
  • Process large-scale JSONL data on AWS S3 (tens of thousands of traces per batch)
  • Convert between chat-completion formats (e.g., OpenAI → Llama 3.1 tool-calling format)
  • Implement smart deduplication and sampling to balance training distribution
  • Design identity-aware train/test splits that measure true generalization
  • Build data validation gates to detect schema drift and format anomalies
  • Create a continuous pipeline that auto-processes new production traces for retraining


Requirements

  • Experience: 6+ years data engineering focused on ML data pipelines
  • Python: Strong — pandas, pyarrow, JSONL processing at scale
  • ML Data Libraries: HuggingFace Datasets, Arrow-based storage
  • Data Formats: Multi-turn conversation/chat data structures and tokenizer-specific formatting
  • Deduplication: Content hashing, identity-based grouping strategies
  • AWS: S3, EC2, batch processing workflows

Preferred (Not Required): LLM training data prep (chat templates, tool-calling schemas); Axolotl or similar dataset formats; data versioning (DVC, LakeFS); browser-automation trace data or Playwright.



Benefits

  • Comprehensive Medical Coverage:
    Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind.
  • Robust Protection Plans:
    Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones.
  • Retirement Benefits:
    PF and Gratuity provided as per standard government regulations.
  • Flexible Work Options:
    Enjoy hybrid work arrangements & flexible working hours.
  • Generous Leave Policy:
    21 days of annual leave, in addition to 10 company-declared holidays.
  • Employee Well-being Spaces:
    Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.


Skills

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available