AI Data Engineering
Summary
An AI data engineer owning end-to-end enterprise data solutions in a hybrid-cloud environment: building batch and real-time pipelines with Databricks, Spark, Kafka, Python and SQL, and productionising GenAI data capabilities such as knowledge bases and RAG architectures for analytics and AI applications.
Responsibilities
- Own the end-to-end engineering of enterprise data solutions, from ingestion and processing through to data consumption, across a hybrid-cloud environment. Build solutions that are scalable, resilient, secure and aligned with the organisation’s technology architecture and governance framework.
- Engineer high-volume batch and real-time data pipelines using platforms such as Databricks, Apache Spark and Kafka, with emphasis on performance, reliability, monitoring and long-term maintainability.
- Establish and enhance data foundations for analytics, ML and AI, ensuring data is accessible, trusted and fit for downstream use cases.
- Drive the engineering and productionisation of GenAI data capabilities, including knowledge bases and Retrieval-Augmented Generation (RAG) architectures supporting enterprise AI and agentic applications.
- Build data processing and transformation logic using Python, PySpark and SQL, including data cleansing, validation and enrichment based on defined business and technical requirements.
- Design appropriate ingestion approaches for data originating from APIs, databases, files, event streams and other enterprise systems, collaborating with upstream and downstream teams to establish effective integration patterns.
- Engineer the underlying capabilities required for knowledge retrieval and AI applications, including knowledge storage, document/data lifecycle management, embedding generation, vectorisation and related components.
- Take ownership of data pipeline health and operational performance, proactively identifying data quality issues, failures, bottlenecks and opportunities for optimisation.
- Establish engineering standards and provide technical direction to engineers and implementation partners, covering architecture patterns, reusable frameworks, coding practices, deployment standards and production support.
- Ensure data and AI components are production-ready, with appropriate monitoring, alerting, incident response, troubleshooting, root-cause analysis, release processes and operational documentation.
- Work across the broader data ecosystem, integrating solutions with platforms including Microsoft Fabric, Databricks and Delta Lake, as well as other relevant enterprise technologies.
- Improve engineering efficiency through automation and modern software delivery practices, including source control, CI/CD and repeatable deployment processes.
- Maintain clear technical documentation, metadata and lineage to support governance, transparency, troubleshooting and ongoing platform management.
- Incorporate security, access management, data governance and technology risk controls throughout the development lifecycle, ensuring solutions comply with enterprise policies and regulatory requirements.
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, Information Technology or a related technical discipline.
- 5–8 years of professional experience spanning data engineering, data platforms, cloud data solutions or large-scale analytics engineering, with experience taking solutions into and supporting production environments.
- Demonstrated ability to independently deliver robust data pipelines at scale, covering areas such as orchestration, fault handling, monitoring, performance optimisation and production operations.
- Strong programming and data manipulation capabilities in Python and SQL.
- Practical experience with Apache Spark / PySpark and distributed data processing at scale.
- Experience developing knowledge management, RAG or retrieval-based data solutions for GenAI, LLM or agentic AI applications.
- Exposure to modern data engineering ecosystems, particularly Databricks, Kafka, Delta Lake and/or Microsoft Fabric.
- Good understanding of data platform architecture, cloud environments, security controls, identity and access management, CI/CD and production release practices.
- Strong analytical and troubleshooting capabilities, with a structured approach to resolving complex technical problems.
- Comfortable taking ownership of technical deliverables and driving discussions with architects, engineers, product teams, business stakeholders and upstream/downstream system owners.
- Strong written and verbal communication skills, with the ability to document technical solutions clearly and translate complex requirements into practical engineering outcomes.
- A strong focus on engineering quality, scalability, reliability and operational excellence, with the ability to work effectively in a fast-moving technology environment.