Point your AI agent at freehire and let it find you a job.

Get the CLI →

re-zoo-me

AI Data Engineering

Posted 5 views
Discussion

Summary

An AI data engineer owning end-to-end enterprise data solutions in a hybrid-cloud environment: building batch and real-time pipelines with Databricks, Spark, Kafka, Python and SQL, and productionising GenAI data capabilities such as knowledge bases and RAG architectures for analytics and AI applications.

Responsibilities

  • Own the end-to-end engineering of enterprise data solutions, from ingestion and processing through to data consumption, across a hybrid-cloud environment. Build solutions that are scalable, resilient, secure and aligned with the organisation’s technology architecture and governance framework.
  • Engineer high-volume batch and real-time data pipelines using platforms such as Databricks, Apache Spark and Kafka, with emphasis on performance, reliability, monitoring and long-term maintainability.
  • Establish and enhance data foundations for analytics, ML and AI, ensuring data is accessible, trusted and fit for downstream use cases.
  • Drive the engineering and productionisation of GenAI data capabilities, including knowledge bases and Retrieval-Augmented Generation (RAG) architectures supporting enterprise AI and agentic applications.
  • Build data processing and transformation logic using Python, PySpark and SQL, including data cleansing, validation and enrichment based on defined business and technical requirements.
  • Design appropriate ingestion approaches for data originating from APIs, databases, files, event streams and other enterprise systems, collaborating with upstream and downstream teams to establish effective integration patterns.
  • Engineer the underlying capabilities required for knowledge retrieval and AI applications, including knowledge storage, document/data lifecycle management, embedding generation, vectorisation and related components.
  • Take ownership of data pipeline health and operational performance, proactively identifying data quality issues, failures, bottlenecks and opportunities for optimisation.
  • Establish engineering standards and provide technical direction to engineers and implementation partners, covering architecture patterns, reusable frameworks, coding practices, deployment standards and production support.
  • Ensure data and AI components are production-ready, with appropriate monitoring, alerting, incident response, troubleshooting, root-cause analysis, release processes and operational documentation.
  • Work across the broader data ecosystem, integrating solutions with platforms including Microsoft Fabric, Databricks and Delta Lake, as well as other relevant enterprise technologies.
  • Improve engineering efficiency through automation and modern software delivery practices, including source control, CI/CD and repeatable deployment processes.
  • Maintain clear technical documentation, metadata and lineage to support governance, transparency, troubleshooting and ongoing platform management.
  • Incorporate security, access management, data governance and technology risk controls throughout the development lifecycle, ensuring solutions comply with enterprise policies and regulatory requirements.

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, Information Technology or a related technical discipline.
  • 5–8 years of professional experience spanning data engineering, data platforms, cloud data solutions or large-scale analytics engineering, with experience taking solutions into and supporting production environments.
  • Demonstrated ability to independently deliver robust data pipelines at scale, covering areas such as orchestration, fault handling, monitoring, performance optimisation and production operations.
  • Strong programming and data manipulation capabilities in Python and SQL.
  • Practical experience with Apache Spark / PySpark and distributed data processing at scale.
  • Experience developing knowledge management, RAG or retrieval-based data solutions for GenAI, LLM or agentic AI applications.
  • Exposure to modern data engineering ecosystems, particularly Databricks, Kafka, Delta Lake and/or Microsoft Fabric.
  • Good understanding of data platform architecture, cloud environments, security controls, identity and access management, CI/CD and production release practices.
  • Strong analytical and troubleshooting capabilities, with a structured approach to resolving complex technical problems.
  • Comfortable taking ownership of technical deliverables and driving discussions with architects, engineers, product teams, business stakeholders and upstream/downstream system owners.
  • Strong written and verbal communication skills, with the ability to document technical solutions clearly and translate complex requirements into practical engineering outcomes.
  • A strong focus on engineering quality, scalability, reliability and operational excellence, with the ability to work effectively in a fast-moving technology environment.

Skills

Apply

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available