Point your AI agent at freehire and let it find you a job.

Get the CLI →

jobgether

NewBe an early applicant

Senior Data Engineer (AI/ML)

Posted
Discussion

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Data Engineer (AI/ML) based in India.

As a Senior Data Engineer (AI/ML), you will build the data and AI infrastructure powering next-generation intelligent products and experiences.
You will combine modern data engineering with Generative AI, working across LLMs, RAG, embeddings, vector search, and AI agents.
The role involves designing scalable data platforms and production-grade pipelines for training, inference, evaluation, and retrieval workloads.
You will collaborate closely with data scientists, ML engineers, software engineers, and product teams to turn innovative prototypes into reliable systems.
Your work will help establish strong foundations for data quality, governance, observability, performance, and cost efficiency.
This is an opportunity to shape AI-enabled data products while working with large-scale distributed and streaming technologies in a global environment.

Accountabilities:

  • Design and build AI/LLM data pipelines supporting training, inference, evaluation, embeddings, and retrieval workloads.
  • Build production-grade RAG systems covering ingestion, chunking, embedding generation, indexing, retrieval, reranking, and context construction.
  • Develop AI applications using LLMs, structured outputs, function and tool calling, and agentic workflows.
  • Build and optimize semantic search and vector retrieval systems.
  • Develop frameworks for LLM evaluation, monitoring, tracing, quality measurement, latency analysis, and cost optimization.
  • Design scalable batch and streaming pipelines using Databricks, Apache Spark, Delta Lake, Snowflake, and Airflow.
  • Build data products and platforms that make structured and unstructured enterprise data accessible to AI applications.
  • Develop reliable ETL/ELT pipelines and optimize large-scale distributed workloads for performance and cost.
  • Establish data quality, governance, lineage, security, and observability practices.
  • Partner with ML and application engineering teams to transition AI prototypes into reliable, production-ready systems.
  • Support large-scale data platforms, real-time processing, event-driven architectures, and complex orchestration workflows.
  • Contribute to AI evaluation datasets and pipelines that measure quality, accuracy, relevance, latency, and cost.
  • Monitor production AI systems, including token usage, model performance, failures, latency, and overall system health.
  • Requirements:

    • 5+ years of experience in data engineering, software engineering, distributed systems, or a related field.
    • Strong programming skills in Python and/or Scala/Java, combined with advanced SQL capabilities.
    • Hands-on experience with Databricks, Snowflake, Apache Spark, Delta Lake, and Airflow.
    • Strong experience working with cloud-based data platforms and scalable data architectures.
    • Practical experience building applications using LLMs or Generative AI.
    • Strong understanding of RAG architectures, embeddings, vector databases, semantic search, and retrieval systems.
    • Familiarity with prompting, structured outputs, tool calling, model evaluation, and other modern LLM concepts.
    • Experience designing scalable, reliable, observable production data systems.
    • Strong knowledge of large-scale data platforms, distributed processing, and complex data workflows.
    • Experience with real-time and streaming architectures using technologies such as Kafka or Spark Structured Streaming.
    • Experience designing low-latency pipelines and event-driven architectures.
    • Strong experience with multi-stage ETL/ELT and data orchestration workflows using Airflow or similar platforms.
    • Experience optimizing Spark or Databricks workloads through partitioning, clustering, caching, joins, and compute optimization.
    • Experience supporting both batch and real-time AI/ML workloads.
    • Experience with LLM/AI evaluation frameworks, automated evaluations, experimentation, quality metrics, and evaluation datasets.
    • Familiarity with AI observability and tracing, including token usage, model performance, latency, failures, and production monitoring.
    • Experience with LangGraph, LangChain, LlamaIndex, or similar AI orchestration frameworks is preferred.
    • Experience with vector databases such as Qdrant, Pinecone, Weaviate, or Databricks Vector Search is preferred.
    • Familiarity with Kafka, MLflow, Unity Catalog, Databricks Mosaic AI, or model-serving platforms is a plus.
    • Strong understanding of distributed systems, cloud architecture, APIs, CI/CD, data governance, and production operations.
    • Benefits:

      • 100% remote position across India.
      • Work from almost anywhere for up to 20 days per year.
      • Generous paid vacation and time off for your birthday.
      • Paid parental leave.
      • Company-paid therapy sessions through SpringHealth.
      • Company-paid Headspace subscription.
      • Annual company-wide week off to support rest and well-being.
      • Generous health insurance and pension fund.
      • Tax optimization options.
      • Development Dollars and leadership development opportunities.
      • Access to thousands of on-demand learning resources.
      • Paid volunteer time.
      • Travel discounts.
      • Employee Resource Groups.
      • Quarterly team offsites.
      • Global and collaborative working environment.
      • Opportunities to work with large-scale data engineering, Generative AI, distributed systems, and modern AI infrastructure.
      • Flexible collaboration across international teams and time zones, with local laws and regulations taken into consideration.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

Skills

What Senior Data Engineering jobs ask for — and how much of it you have →

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available