Associate Data & ML Engineer
Summary
Associate-level engineer joining green energy producer Globeleq in Cape Town to build and maintain data pipelines (API, SQL, ETL/ELT), consolidate industrial sources (ERP, IoT, CAMs) into a Single Source of Truth, and prepare feature sets for ML such as anomaly detection and predictive analytics, using Python, SQL and AVEVA PI.
A green energy supplier is looking for an Associate Data & ML Engineer to join their team in Cape Town, WC.
Purpose of the role:
The Associate Data & ML Engineer supports the technical delivery of Globeleq’s Data Transformation initiative under the direction of the Data Engineering Manager. The role focuses on building and maintaining assigned data pipelines, onboarding assets and preparing data ready for AI and machine learning. Working within the technical shared services team, the Associate translates clear requirements into reliable, production grade components, applying sound engineering practice under the guidance of senior engineers.
Responsibilities
Data Engineering
- Build, monitor, and maintain assigned data pipelines across API, SQL, and ETL/ELT workflows, resolving routine issues and escalating complex matters where required.
- Develop and extend the integrated Single Source of Truth (SSOT), consolidating data from sources including CAMs, ERP, OT, IoT systems, and SharePoint into a layered architecture comprising staging, core, marts, and feature sets.
- Apply sound data engineering practices, including version control, testing, documentation, and quality assurance.
Machine Learning
- Develop and implement machine learning algorithms for anomaly detection, predictive analytics, and neural network use cases, leveraging established libraries and frameworks under guidance.
- Structure, prepare, and validate feature sets to ensure data is suitable for machine learning applications, including validation of model inputs and outputs.
- Leverage AI coding assistants to improve development efficiency while thoroughly reviewing and validating all generated code before implementation.
Governance & Ownership
- Take ownership of assigned tasks and deliverables, ensuring reliable and timely completion while seeking guidance on scope and priorities where required.
- Work within defined requirements, established development patterns, and agreed standards, proactively identifying and escalating risks.
- Adhere to Data Governance Policies, security requirements, and IT standards, while maintaining accurate and up-to-date change records.
Platform Ownership
- Onboard assets and equipment and map tags into the AVEVA PI Asset Framework and Central Asset Management System (CAMs) using established templates and processes.
- Support the development, maintenance, and continuous improvement of calculations within CAMs.
- Configure alarms and PI Vision displays according to specifications, monitor assigned PI interfaces, and resolve routine data quality issues with appropriate support.
Skills & Competencies
- Programming: Solid programming foundations in Python and SQL, with C# advantageous. Developing proficiency in SQL, including DDL, DML, stored procedures, and views.
- Data Engineering: Working knowledge of API integration, ETL/ELT orchestration, and data modelling.
- Machine Learning: Solid understanding of machine learning fundamentals, including feature engineering, model evaluation, overfitting, and data drift, with the ability to develop simple models using established libraries.
- Industrial Data & Engineering: Interest in AVEVA PI and industrial data, with sound engineering practices across version control, testing, documentation, and data quality.
- Communication & Ownership: Clear and effective communicator with a strong sense of accountability and ownership for assigned work.
Requirements
- Education: Degree in Computer Science, Information Systems, Engineering, Mathematics, or a related field.
- Experience: 2–4 years’ experience in data engineering, data platform development, or a similar technical role, with demonstrated hands‑on delivery.
- Data Engineering: Hands‑on experience with API integration, ETL/ELT scheduling, data modelling, and basic production operations, including monitoring and incident support.
- Technical Skills: Strong proficiency in SQL and Python, particularly for data engineering and machine learning enablement.
Advantageous
- Exposure to AVEVA PI, including PI Data Archive, Asset Framework, and PI Vision.
- Experience within power engineering, renewable energy, industrial data, or IoT environments.
- Familiarity with machine learning in production, MLOps concepts, and deep learning tools such as TensorFlow and PyTorch.