AI Engineer (Data Guardrails & LLM Ingestion Pipelines)
Summary
Design and build scalable data ingestion, cleaning, and LLM-powered pipelines to transform raw data into AI-ready datasets with quality guardrails, using Python, LLMs, and vector databases.
This is a remote position.
About the Role
We are looking for a highly skilled AI Engineer to design and build robust data ingestion, cleaning, validation, and LLM enhancement pipelines that power our AI applications. You will transform raw, unstructured data into high-quality, AI-ready datasets while implementing guardrails that ensure accuracy, consistency, and reliability.
Key Responsibilities
· Design and develop scalable data ingestion pipelines for structured and unstructured data.
· Build automated data cleaning, normalization, and preprocessing workflows.
· Develop AI-powered enrichment pipelines using LLMs (OpenAI, Claude, Gemini, etc.).
· Implement data quality validation and AI guardrails.
· Develop prompt engineering workflows for data transformation.
· Build document processing pipelines for PDFs, Word documents, CSVs, websites, and APIs.
· Develop Retrieval-Augmented Generation (RAG) pipelines.
· Create evaluation frameworks for LLM quality and accuracy.
· Build ETL/ELT workflows for AI-ready datasets.
· Integrate vector databases for semantic search.
· Monitor pipeline performance, cost, latency, and data quality.
· Collaborate with cross-functional teams to deliver production AI systems.
Required Technical Skills
Programming
· Python (Expert)
· SQL
· Git
AI & LLMs
· OpenAI API
· Anthropic Claude API
· Google Gemini API
· Prompt Engineering
· Function Calling
· Structured Outputs
AI Frameworks
· LangChain
· LlamaIndex
· DSPy (Preferred)
· PydanticAI (Nice to Have)
Data Engineering
· Pandas
· Polars
· ETL/ELT Pipelines
· Apache Airflow (Preferred)
· Data Validation Frameworks
Vector Databases
· Pinecone
· Weaviate
· Qdrant
· ChromaDB
· FAISS
Cloud & Infrastructure
· Docker
· Kubernetes (Preferred)
· AWS / Azure / GCP
· Linux
Databases
· PostgreSQL
· MongoDB
· Redis
Requirements
Preferred Qualifications
· Experience building production-grade AI systems.
· Strong understanding of RAG architectures.
· Experience implementing AI guardrails and hallucination mitigation.
· Experience with OCR and document parsing.
· Experience with embedding models and semantic search.
· Knowledge of data governance and security best practices.
Success Metrics
· Build scalable ingestion pipelines.
· Deliver automated data cleaning and LLM enhancement workflows.
· Implement AI guardrails to improve output quality.
· Develop evaluation pipelines for LLM performance.
· Contribute to a production-ready AI platform.