RAG / AI Data Engineer
Many AI initiatives succeed or fail based on the quality, accessibility and governance of the underlying data. We work with organisations building the foundations required to support retrieval, search, agentic systems and production AI applications.
Role Overview
- Design and build data pipelines to support AI applications
- Develop Retrieval-Augmented Generation (RAG) architectures
- Create and maintain vector databases and knowledge repositories
- Structure, clean and prepare data for AI consumption
- Improve retrieval accuracy, relevance and performance
- Build scalable data foundations for AI products and agents
- Collaborate with architects, engineers and product teams to enable AI delivery
Tools & Technologies (required)
- Python
- SQL
- Vector databases (Pinecone, Weaviate, Qdrant, Chroma or similar)
- Embedding models and retrieval frameworks
- LangChain, LlamaIndex or equivalent
- Data pipeline and ETL tooling
Experience (required)
- Building data pipelines for AI or ML applications
- Designing or implementing RAG architectures
- Working with vector databases
- Managing structured and unstructured datasets
- Optimising retrieval quality and search performance
- Agent-based AI systems and workflows
About You
- Strong understanding of data architecture and retrieval systems
- Able to balance accuracy, performance and scalability
- Comfortable working across structured and unstructured datasets
- Interested in practical AI implementation rather than theoretical research
- Focused on creating reliable foundations for AI systems
- Strong problem-solving and analytical skills