freehire launches on Product Hunt on 26 August.

Follow →

AI Data Engineer

Summary

Designs and maintains AI-ready data pipelines, feature stores, and curated datasets to power machine learning models, generative AI, and real-time analytics using cloud-native tools.

Job Purpose

The Manager: AI Data Engineer is responsible for designing, building, and managing data pipelines and AI-ready data products that support artificial intelligence and advanced analytics initiatives across the organization.

The role focuses on enabling feature engineering, curated datasets, real‑time and batch data pipelines, and data quality controls that support machine learning models, generative AI solutions, and AI-enabled business services.

The position works closely with AI Platform Engineering, Data Platform teams, MLOps, AI Application Engineers, Security & Risk, and business stakeholders to ensure AI solutions are powered by accurate, secure, compliant, and production‑ready data.

Key Performance Areas (KPAs)

1. AI Data Architecture & Pipelines

  • Design and implement scalable data pipelines to support:
    • Model training
    • Feature generation
    • Inference and real‑time decision-making
  • Build and maintain:
    • Batch data pipelines
    • Streaming data pipelines
    • Data ingestion from internal and external sources
  • Ensure data architectures align with:
    • Enterprise AI reference architectures
    • Organizational data and platform standards

2. Feature Stores & AI Data Products

  • Design, build, and manage feature stores that support:
    • Reuse of engineered features
    • Consistency between model training and inference
  • Create curated AI-ready datasets for:
    • Data Scientists
    • AI Engineers
    • Product Teams
    • Analytics Teams
  • Improve discoverability and reuse of AI data assets across the organization.

3. Data Quality, Lineage & Observability

  • Implement automated validation for:
    • Data accuracy
    • Completeness
    • Timeliness
    • Consistency
  • Maintain end-to-end data lineage and traceability.
  • Build observability into AI data pipelines to:
    • Detect data drift
    • Identify anomalies
    • Support root cause analysis
  • Collaborate with Site Reliability Engineering (SRE) and Platform teams to ensure operational stability.

4. Privacy, Security & Regulatory Compliance

  • Enforce data privacy, sovereignty, and protection requirements.
  • Implement appropriate access controls, masking, and encryption where required.
  • Ensure AI datasets comply with:
    • Applicable regulatory requirements
    • Organizational security and governance standards
  • Support internal and external audit activities related to AI data usage.

5. Enablement of AI & Analytics Use Cases

  • Partner with AI Application and MLOps teams to:
    • Enable rapid experimentation
    • Accelerate production deployment of AI solutions
    • Reduce data preparation effort for AI initiatives
  • Support:
    • Generative AI solutions
    • Machine Learning workloads
    • Advanced Analytics initiatives
  • Balance innovation with strong governance and operational discipline.

6. Continuous Improvement & Standardization

  • Standardize AI data engineering practices, patterns, and tooling.
  • Contribute to enterprise AI and data platform roadmaps.
  • Drive continuous improvement in:
    • Pipeline reliability
    • Data freshness
    • Scalability
    • Performance
    • Reusability of AI data products

Job Requirements

Education

  • Master’s Degree in Computer Science, Data Science, Artificial Intelligence, Big Data, Information Systems, Engineering, or a related discipline.

Experience

  • Minimum of 5 years’ experience in Data Engineering or Data Platform roles.
  • Hands‑on experience designing and implementing:
    • Large‑scale batch and streaming data pipelines
    • Feature engineering pipelines
  • Experience working with:
    • Cloud data platforms (Azure is essential)
    • Structured and unstructured data
  • Exposure to Artificial Intelligence, Machine Learning, or Advanced Analytics environments is preferred.
  • Experience working within highly regulated industries is advantageous.

Technical Competencies

  • Data pipeline architecture and orchestration
  • Feature Store design and implementation
  • Streaming and batch data processing
  • Data quality frameworks
  • Data lineage and observability
  • Cloud-native data platforms
  • Data modelling techniques
  • Security and privacy‑by‑design principles
  • AI data lifecycle management

Skills

  • Strong analytical and problem‑solving abilities
  • Excellent data modelling capabilities
  • Technical documentation and communication skills
  • Cross‑functional collaboration with AI, platform, engineering, and business teams
  • Ability to balance innovation with governance
  • Continuous improvement mindset

Behavioural Competencies

  • Detail‑oriented with a strong focus on quality
  • Accountable and delivery‑driven
  • Structured and methodical approach to work
  • Curious with a passion for learning emerging technologies
  • Collaborative and respectful team player
  • Comfortable working across multiple business units and large-scale environments
#J-18808-Ljbffr

See also