Data Engineer (FMCG / Food / Retail industry experience)
Summary
Builds and maintains the data foundation for an FMCG/retail organisation in Bellville: ETL/ELT pipelines and integrations across ERP, CRM and supply-chain systems, data quality controls, and BI-ready datasets, with growing use of AI-assisted and agentic tooling. Core stack: SQL (MS/Postgres), Python/JavaScript, REST APIs, data warehousing.
Data Integration & Pipeline Development
- Design, develop and maintain reliable ETL/ELT data pipelines across multiple business systems.
- Integrate data from systems such as ERP, TMS, WMS, finance, CRM, supply chain, HR and other business applications.
- Reduce reliance on manual data extraction and manipulation through effective automation.
- Ensure data pipelines are scalable, efficient and appropriate for the organisation's future data requirements.
- Work closely with Data & Systems Architecture to ensure engineering solutions align with broader technology and architecture standards.
Data Quality, Integrity & Reliability
- Implement automated data-quality checks and validation controls.
- Assist with the implementation of MDM across the organisation.
- Identify and investigate inconsistencies, missing data, duplication and other data-quality issues.
- Work with system owners and business stakeholders to address root causes of data-quality problems.
- Monitor data pipelines and integration processes to identify failures or performance issues.
- Troubleshoot and resolve data-processing and integration errors.
- Ensure that data used for reporting and analysis is accurate, complete, consistent and available when required.
- Support the development of a trusted organisational data environment.
Analytics & Business Intelligence Enablement
- Develop and maintain analysis-ready datasets for Data Analysts, Analyst Developers and other authorised business users.
- Support the development of reliable data models for dashboards, management reporting and business intelligence solutions.
- Assist analysts with complex data extraction, transformation and integration requirements.
- Ensure that commonly used business information is sourced from controlled and consistent datasets.
- Enable greater self-service, democratized analytics through well-structured and governed data.
Automation & Continuous Improvement
- Review existing data flows and recommend improvements.
- Optimise database queries, pipelines and processing routines to improve performance.
- Proactively identify opportunities where improved data integration can simplify business processes.
- Contribute to continuous improvement initiatives within Business Transformation.
- Migrate data warehouse, data lake or similar data structures where applicable.
Business Transformation & Systems Projects
- Participate in business transformation, systems implementation and process-improvement projects.
- Assess data requirements when new systems, processes or technologies are introduced.
- Support data migration, cleansing, mapping and validation during system implementations.
- Work with Business Process Management to understand how information moves through business processes and identify opportunities for improved automation.
- Provide technical data expertise during solution design and project implementation.
- Support testing and implementation of new data solutions.
- Assist with post-implementation troubleshooting and optimisation.
Data Governance, Security & Compliance
- Apply organisational data governance standards across data engineering solutions.
- Ensure appropriate access controls are implemented for sensitive and confidential information.
- Work with IT and relevant stakeholders to ensure data solutions comply with information-security requirements such as POPIA and ISO27001.
- Maintain appropriate controls around the extraction, transfer, storage and use of organisational data.
- Support the establishment and maintenance of data stewardship, definitions and governance practices.
Documentation & Technical Support
- Maintain accurate technical documentation for data pipelines, integrations and data structures.
- Develop and maintain data-flow diagrams, integration specifications and data dictionaries where appropriate.
- Maintain appropriate change records for data solutions.
- Provide technical support and troubleshooting for data-related issues.
- Ensure that critical data processes are sufficiently documented to reduce dependency on individual knowledge.
Artificial Intelligence & Agentic Capability
- The Data Engineer is expected to be a user, implementer and support resource for the organisation's artificial-intelligence and agent-assisted capabilities, applying them to data engineering, data quality and business-process outcomes within approved governance and security controls.
- Use AI-assisted and agentic development tooling in day-to-day engineering work pipeline and integration development, code generation and review, test creation, documentation and troubleshooting and validate all generated code, SQL and configuration before it reaches a governed environment.
- Apply AI and machine-assisted techniques to business-process analysis, ETL/ELT development and data-quality optimisation, including anomaly detection, record matching and de-duplication, classification and rule suggestion, with human review before production use.
- Build and maintain the semantic and metadata foundations that make organisational data usable by AI and agent-based tools business glossaries, data dictionaries, ontologies and taxonomies, and governed data-product contracts so that natural-language questions return consistent, explainable and permission-aware answers.
- Support the implementation of AI-enabled platform capabilities together with Data & Systems Architecture and external technology partners, including configuration, integration, testing, user enablement and post-implementation optimisation.
- Apply AI within the organisation's data-protection and information-security requirements: confidential, personal or business-sensitive data may only be processed by approved AI services, consistent with the governance requirements in 2.6.
- Understanding of code harnesses and agentic coding tools, large language model behaviour, prompt engineering and retrieval-based approaches will be favourable.
KEY PERFORMANCE AREAS (KPAS):
CORE COMPETENCIES:
- Analytical Thinking Able to understand complex data structures, identify relationships and systematically resolve data problems.
- Problem Solving Approaches technical and data-related challenges logically and focuses on identifying sustainable solutions rather than temporary fixes.
- Attention to Detail Maintains a high level of accuracy when working with large datasets, integrations and business-critical information.
- Business Understanding Develops an understanding of how the business operates and ensures technical data solutions support practical business requirements.
- Collaboration Works effectively across technical and non-technical teams and is able to translate business requirements into appropriate data solutions.
- Ownership Takes accountability for the reliability, quality and performance of assigned data solutions.
- Continuous Learning Keeps abreast of developments in data engineering, cloud technology, automation and analytics.
- Communication Able to explain technical concepts clearly to stakeholders with different levels of technical knowledge.
- Futuristic and Open Mindset
- Ability to rapidly change and pivot
QUALIFICATIONS & EXPERIENCE:
- A relevant tertiary qualification in one of the following or a related field: Computer Science / Information Systems / Information Technology / Data Engineering / Software Engineering
- Relevant industry certifications in data engineering, cloud platforms or database technologies would be advantageous.
Experience
- Approximately 5-7 years' relevant experience in data engineering, database development, systems integration or a similar technical data role. At least two of these years should include systems or data integration, or the exchange of data between internal and external systems.
- Experience developing and maintaining ETL/ELT pipelines.
- Proven experience in developing data projects using Python.
- Strong practical experience working with SQL and relational databases.
- Experience with data modelling and data warehouse concepts.
- Experience working in a BI, analytics or reporting environment.
- Experience designing, building or consuming API-based integrations (REST/JSON), including authentication, pagination, error handling, retry and idempotency behaviour.
- Experience with file-based and batch integration to and from external parties, including scheduled transfers, secure file transfer (SFTP), and file validation, quarantine and reconciliation.
- Experience publishing or consuming versioned, governed datasets or data products for downstream and third-party consumers, including schema change management and backward compatibility.
- Experience delivering data or integration artefacts through source control (Git) and promoting them across development, test and production environments.
- Experience monitoring and supporting production integrations, including failure detection, error triage, reprocessing or replay of failed records, and root-cause resolution.
- Experience in a multi-site Retail, FMCG, Wholesale, Logistics or Supply Chain environment would be advantageous.
- Experience supporting system implementations or business-transformation projects would be advantageous.
- Experience with event- or message-driven integration (webhooks, publish/subscribe or streaming platforms such as Kafka), including at-least-once delivery, idempotency, replay and dead-letter handling, would be advantageous.
- Experience with an integration or middleware runtime for example Apache Camel, Azure Integration Services, Boomi, MuleSoft or a comparable iPaaS/ESB would be advantageous.
- Experience with master data management, data cataloguing or business glossary tooling would be advantageous.
- Experience exchanging data with external trading partners, suppliers or service providers under defined data-sharing, security and privacy controls would be advantageous.
TECHNICAL SKILLS & KNOWLEDGE
Essential
- Advanced SQL (MS, Postgres)
- ETL/ELT development
- Data modelling
- Data validation and quality controls
- Python/JavaScript
- REST API design and consumption (JSON payloads, authentication, pagination, error handling, retries and idempotency)
- DevOps or automated deployment practices including environment promotion across development, test and production
- Microsoft 365 & Co-Pilot
- Data integration patterns: batch, file-based, API-based and event-driven integration; incremental loads and change data capture; source-to-target reconciliation.
- API and data-contract concepts: interface specifications (OpenAPI/Swagger), schema definition, versioning and backward compatibility.
- Secure file transfer (SFTP) and structured file handling (CSV, fixed-width, JSON, XML).
- Credential, key and secret management, and least-privilege access to data and integration endpoints.
- Integration and pipeline observability: logging, monitoring, alerting, error handling, retry, dead-letter and replay.
- Master data management, data cataloguing and metadata concepts, including business glossaries, data dictionaries and lineage.
Advantageous
- Microsoft Foundry
- Microsoft Fabric
- Pascal Script
- OData 4.01 and other standards-based data-access protocols
- Containerisation and orchestration fundamentals (Docker, Kubernetes)
- Distributed SQL query engines and lakehouse/object-storage concepts (for example Trino, S3-compatible storage, Parquet/Avro)
- Observability tooling and standards (for example OpenTelemetry)
- EDI and trading-partner exchange formats used in Retail, FMCG and Logistics
- Ontology, taxonomy and semantic-layer modelling concepts
ONLY SHORTLISTED CANDIDATES WILL BE CONTACTED
Skills
- Agentic AI
- AI
- Analytics
- Anomaly Detection
- API
- Api Design
- Authentication
- Automation
- Azure
- Cloud
- Containerization
- CRM
- Data Engineering
- Data Governance
- Data Lake
- Data Modeling
- Data Pipelines
- Data Quality
- Data Warehousing
- DevOps
- Docker
- EDI
- ELT
- ERP
- ETL
- Event Driven Architecture
- Git
- ISO 27001
- JavaScript
- JSON
- Kafka
- Kubernetes
- Lakehouse
- LLM
- Master Data Management
- MDM
- Microsoft Fabric
- Mulesoft
- Observability
- OpenAPI
- OpenTelemetry
- Parquet
- PostgreSQL
- Process Improvement
- Prompt Engineering
- Python
- REST
- S3
- SFTP
- Solution Design
- SQL
- Swagger
- Trino
- Webhooks
- XML