Senior AI Data Engineer
Summary
Senior AI Data Engineer at ComplySci (a RegTech firm) in London, building the company's semantic layer: turning ontological models into production knowledge graphs, vector search infrastructure, and RAG/LLM-powered data pipelines for AI-ready compliance data products. Core stack: Python, JSON-LD/RDF/OWL/SKOS, graph DBs (Neo4j, Neptune, Fuseki), vector DBs (Pinecone, Weaviate, Qdrant), and AWS/Azur
In this role you will implement Comply’s semantic layer, turning ontological models into production-ready knowledge graphs, vector search infrastructure, and LLM-powered pipelines. You’ll own semantic layer delivery and collaborate with application and data teams to ensure AI-ready data products are reliable and scalable. You’ll join a new Data and Analytics team to advance our AI ambitions and enable future data capabilities in the compliance domain. This role offers hands-on work at the intersection of knowledge representation and AI infrastructure with meaningful impact on regulatory programs.
Responsibilities- Implement JSON-LD semantic models into production data systems
- Build and maintain knowledge graph structures and graph DB schemas
- Develop data ingestion pipelines and ensure semantic consistency with downstream products
- Design embedding pipelines and operate vector DB infrastructure for semantic search
- Implement RAG architectures grounding LLM outputs in proprietary data
- Evaluate and integrate suitable LLM tooling and frameworks
- Build reliable, observable data pipelines from upstream sources
- Apply DataOps practices including testing, monitoring, lineage, and SLAs
- Collaborate with Ontologist to reflect domain intent in models
- Assist application teams in adopting AI-ready data products
- Strong hands-on data engineering with a focus on semantic or AI data infrastructure
- Experience building/operating knowledge graphs or graph databases (e.g. Jena Fuseki, Neo4j, Amazon Neptune)
- Experience with vector databases and embedding pipelines (e.g. Pinecone, Weaviate, Qdrant, pgvector)
- Practical experience implementing RAG architectures or LLM-integrated data pipelines
- Familiarity with semantic web standards — JSON-LD, RDF, OWL, SKOS
- Strong Python skills and data pipeline framework experience
- Experience with cloud-native data platforms (AWS, Azure, or GCP)
- Desirable: domain-driven design (DDD) and bounded contexts
- Experience working with ontologists or knowledge engineers is a plus
- Familiarity with data contracts and data product frameworks is a plus
- Experience with DataOps tooling, data reliability, or data observability is desirable
- Background in financial services, RegTech, or compliance data is a plus
- Cross-functional collaboration
- Strong communication with technical and non-technical stakeholders
- Problem-solving mindset with attention to data quality
- JSON-LD, RDF, OWL, SKOS
- Graph databases: Jena Fuseki, Neo4j, Amazon Neptune
- Vector databases: Pinecone, Weaviate, Qdrant, pgvector