freehire launches on Product Hunt on 26 August.

Follow →

AI Engineer (Data Guardrails & LLM Ingestion Pipelines)

Open 26d

Summary

Design and build scalable data ingestion, cleaning, and LLM-powered pipelines to transform raw data into AI-ready datasets with quality guardrails, using Python, LLMs, and vector databases.

This is a remote position.

About the Role

We are looking for a highly skilled AI Engineer to design and build robust data ingestion, cleaning, validation, and LLM enhancement pipelines that power our AI applications. You will transform raw, unstructured data into high-quality, AI-ready datasets while implementing guardrails that ensure accuracy, consistency, and reliability.

Key Responsibilities

· Design and develop scalable data ingestion pipelines for structured and unstructured data.

· Build automated data cleaning, normalization, and preprocessing workflows.

· Develop AI-powered enrichment pipelines using LLMs (OpenAI, Claude, Gemini, etc.).

· Implement data quality validation and AI guardrails.

· Develop prompt engineering workflows for data transformation.

· Build document processing pipelines for PDFs, Word documents, CSVs, websites, and APIs.

· Develop Retrieval-Augmented Generation (RAG) pipelines.

· Create evaluation frameworks for LLM quality and accuracy.

· Build ETL/ELT workflows for AI-ready datasets.

· Integrate vector databases for semantic search.

· Monitor pipeline performance, cost, latency, and data quality.

· Collaborate with cross-functional teams to deliver production AI systems.

Required Technical Skills

Programming

· Python (Expert)

· SQL

· Git

AI & LLMs

· OpenAI API

· Anthropic Claude API

· Google Gemini API

· Prompt Engineering

· Function Calling

· Structured Outputs

AI Frameworks

· LangChain

· LlamaIndex

· DSPy (Preferred)

· PydanticAI (Nice to Have)

Data Engineering

· Pandas

· Polars

· ETL/ELT Pipelines

· Apache Airflow (Preferred)

· Data Validation Frameworks

Vector Databases

· Pinecone

· Weaviate

· Qdrant

· ChromaDB

· FAISS

Cloud & Infrastructure

· Docker

· Kubernetes (Preferred)

· AWS / Azure / GCP

· Linux

Databases

· PostgreSQL

· MongoDB

· Redis


Requirements

Preferred Qualifications

· Experience building production-grade AI systems.

· Strong understanding of RAG architectures.

· Experience implementing AI guardrails and hallucination mitigation.

· Experience with OCR and document parsing.

· Experience with embedding models and semantic search.

· Knowledge of data governance and security best practices.

Success Metrics

· Build scalable ingestion pipelines.

· Deliver automated data cleaning and LLM enhancement workflows.

· Implement AI guardrails to improve output quality.

· Develop evaluation pipelines for LLM performance.

· Contribute to a production-ready AI platform.


See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available