freehire launches on Product Hunt on 26 August.

Follow →

Senior Data Engineer

Summary

Senior Data & ML Engineer builds automated data pipelines, a central SSOT database, and AI/ML-ready data models to support reporting, analytics, and machine learning across the business.

The Senior Data & ML Engineer is responsible for executing the technical implementation of the Data Transformation. The role focuses on designing and building the data ingestion and processing platform, including automated data pipelines, an integrated database, cross-functional system integrations and AI/ ML-ready data structures.

The Senior Data & ML Engineer must translate strategic direction into concrete technical solutions, make sound architectural recommendations, and deliver scalable, robust, production-grade data capabilities that support reliable reporting, advanced analytics and machine learning use cases across the business.

This role will form part of our technical shared services team, contributing to the development of digital management systems, O&M projects implementations and ongoing development within the Data Transformation Project.

Key Responsibilities

Design, build and maintain end-to-end automated data pipelines from internal and external sources into a central data platform:

  • Develop an integrated and scalable Single Source of Truth (SSOT) database, consolidating data from ERP, OT/IoT, SharePoint ensuring scalable data flows.
  • Develop modular and scalable data ingestion developing API integrations, SQL Stored Procedures and ETL frameworks, maintaining reliable, automated data ingestion between internal and external platforms to the SSOT database.
  • Own end to end data development solutions.

Design and develop scalable data models (staging, core, marts, feature sets) that support strategic reporting, advanced analytics and ML.

  • Design scalable data models across clearly defined layers, including staging (raw landed data), core (cleaned and standardised single source of truth), marts (business-ready views for specific domains), and feature sets (model-ready tables for machine learning and advanced analytics).
  • Implement MLOps practices (versioning of data and models, CI/CD for models, monitoring, retraining strategies) as ML use cases mature.
  • Own end-to-end development and processing (data algorithms and ML solutions)
  • Ensure all data models, pipelines and storage approaches are AI/ML-ready.

Technical platform development, data orchestration and data management

  • Apply and enforce data management, security and governance standards in line with the Data Governance Policy.
  • Implement structured change management (version control, release processes, approvals, rollback plans).
  • Work with divisional data owners to reduce data silos, standardise data flows and ensure adherence to agreed standards and timelines.

Asset Lifecycle Management & Platform Integration

  • Oversee the full onboarding and offboarding process for company assets and equipment within the central asset management platforms.
  • Ensure seamless data integration by validating that all asset information is correctly captured and synchronized.
  • Maintain data integrity and platform compliance by routinely reviewing asset entries, resolving discrepancies.

Skills and Competencies

  • ETL/ELT orchestration and job scheduling (Automated workflows)
  • Data modelling (staging, core, marts, feature sets); production operations (monitoring, alerting, incident response)
  • Strong SQL; proficiency in Python for data engineering and ML-enabling tasks; and solid programming foundations in Python, SQL and/ or C#
  • Ability to make scalable architectural decisions and prepare data for ML and model integration into workflows.
  • Solid ML foundations (feature engineering, evaluation, overfitting, drift) and ability to design data pipelines that are fit for ML.

Senior, hands‑on engineering mindset; comfortable owning technical direction.

  • Proven ability to design and lead data platform or data product builds.
  • Self‑directed and proactive: identifies problems, proposes solutions and drives implementation without detailed step‑by‑step direction.
  • Clear, structured communication skills. can explain technical options and trade‑offs to non‑technical stakeholders and leadership.

Strong systems thinking and architecture skills: designs for scalability, maintainability and AI/ML-readiness from the outset.

  • Strong engineering discipline: version control, testing, deployment processes, documentation and incident handling.
  • Enjoys building automation, integrations and ML‑ready datasets.
  • Cross‑team coordination and influence – able to work with multiple divisions, follow up with stakeholders and enforce agreed standards and timelines.

Experience, Knowledge and Qualifications

  • Degree in Computer Science, Information Systems, Engineering, Mathematics or a related field. Proficient in data engineering.
  • 5+ years in data engineering/ data platform development at a senior/ lead level.
  • Programming foundations in Python, SQL and/or C#.
  • Hands‑on responsibility for API integration (REST/JSON, auth, pagination, error handling); ETL/ELT orchestration and job scheduling; data modelling (staging, core, marts/ feature sets); and production operations (monitoring, alerting, incident response).
  • Strong SQL skills (DDL/DML, performance tuning, stored procedures, views, functions).
  • Proficiency in Python for data engineering and ML-enabling tasks, plus experience with ETL.
  • Strong foundations in ML concepts and practical experience preparing data for ML and integrating models into data workflows (even if not a pure data scientist).

Advantageous:

  • Hands‑on experience with ML models in production (e.g. forecasting, classification, anomaly detection) and associated MLOps tooling.
  • Exposure to neural networks/ deep learning (e.g. TensorFlow, PyTorch) and modern ML pipelines.
  • Design and support workflow automation and lightweight data applications using tools such as Power Apps and Power Automate, integrating these solutions with the core data platform to enable efficient business processes.
  • Experience with data lakes/ big data architectures and orchestration tools (e.g. Airflow, Prefect, Azure Data Factory or similar).
  • Familiarity with AI/ ML governance, model risk and secure data handling.
  • Experience working with industrial/IoT or energy sector data.

This role is a weekly hybrid role in Cape Town with a CTC package ranging from ZAR 616 - ZAR 822.

See also