Junior Data Engineer
Summary
A junior data engineer owns and develops the company's Microsoft Fabric data environment: building and maintaining ETL/ELT pipelines, data warehouses and Power BI-ready datasets using SQL Server/T-SQL, Python/PySpark, OneLake/Lakehouse, Dataflows Gen2, and Git/Azure DevOps CI/CD.
Minimum Education (essential)
- Bachelor's degree in Computer Science, Information Systems, Data Engineering or a related field.
- Relevant Microsoft Fabric and/or Azure data certification.
- 0-1 years of practical data engineering experience, including strong recent hands-on experience with Microsoft Fabric.
- Strong hands-on experience with Microsoft Fabric, including OneLake, Lakehouse, Warehouse, Data Pipelines/Data Factory, notebooks and Dataflows Gen2.
- Strong SQL Server and T-SQL capability, including complex query development, schema design, indexing and performance optimisation.
- Practical experience developing, maintaining and supporting production ETL/ELT pipelines.
- Experience integrating and extracting data from REST/SOAP APIs, databases, flat files and other structured or unstructured data sources.
- Proficiency in data transformation using SQL and Python/PySpark, with an understanding of scalable data processing practices.
- Practical experience in data warehousing, dimensional modelling, incremental loading, orchestration and schema evolution.
- Experience with troubleshooting pipeline failures, data quality issues and performance bottlenecks, including the ability to restore service efficiently.
- Experience with source control, CI/CD and deployment practices using Git, Azure DevOps or equivalent tools.
- Experience supporting Power BI and other downstream analytical or reporting requirements.
- Demonstrated ability to take ownership of an existing technical environment with limited hand-holding.
- Strong documentation, communication and stakeholder engagement skills.
- Microsoft Fabric: OneLake, Lakehouse, Warehouse, Data Pipelines/Data Factory, notebooks and Dataflows Gen2.
- SQL Server / T-SQL
- Python / PySpark
- REST/SOAP APIs and structured/unstructured data ingestion.
- ETL/ELT, incremental loading, orchestration and scheduling.
- Dimensional modelling, medallion architecture, schema evolution and data warehousing.
- Power BI integration and understanding of downstream analytical requirements.
- Git / Azure DevOps, CI/CD and environment deployment practices.
- Monitoring, data quality, performance optimisation, security and operational support.
- Proficient in Afrikaans and English.
- Own transport and valid drivers license.
KEY PERFORMANCE AREAS AND OBJECTIVES
Fabric Data Engineering and Pipeline Development
- Take ownership of the existing Microsoft Fabric data engineering environment and become productive quickly following handover.
- Design, develop, maintain and orchestrate reliable batch and near-real-time data ingestion pipelines using appropriate Microsoft Fabric capabilities.
- Extract and ingest data from structured and unstructured sources, including REST APIs, SOAP APIs, databases and flat files.
- Develop robust data transformation logic using SQL, Python/PySpark, Fabric notebooks and Dataflows Gen2, as appropriate.
- Implement incremental loading, retry mechanisms, logging, monitoring and alerting to support data integrity and pipeline reliability.
- Troubleshoot and resolve pipeline failures and data processing issues efficiently.
- Optimise data pipelines and processing workloads for performance, scalability and cost-effectiveness.
- Design, manage and evolve scalable data architectures using Microsoft Fabric, OneLake, Lakehouse, Warehouse and SQL Server.
- Maintain appropriate data-layering and medallion architecture principles, where applicable, with clear movement from raw to curated data.
- Develop and maintain robust schema designs, indexes, partitioning and query strategies to support analytical and operational workloads.
- Manage schema evolution and version control to maintain consistency and minimise disruption to downstream consumers.
- Maintain metadata, data dictionaries, architecture documentation and technical documentation to improve supportability and reduce key-person dependency.
- Define and maintain appropriate role-based access and security controls.
- Build and maintain analytical data stores using Microsoft Fabric Warehouse and/or Lakehouse patterns.
- Apply appropriate data-loading, partitioning, storage optimisation and query-performance practices.
- Develop and maintain stable, well-modelled datasets for Power BI and other analytical consumers.
- Work with reporting and analytical teams to investigate and resolve data-related issues.
- Ensure data structures and outputs support downstream reporting and business intelligence requirements.
- Develop and maintain conceptual, logical and physical data models.
- Apply dimensional modelling techniques, including star and snowflake schemas, to support analytics and reporting.
- Apply appropriate normalisation and relational modelling techniques for operational and analytical workloads.
- Ensure consistency of data models across systems.
- Manage schema versioning and evolution without unnecessarily disrupting downstream consumers.
- Apply agreed data engineering standards and modelling principles consistently.
- Work independently and take end-to-end ownership of assigned data engineering deliverables, incidents and production issues.
- Provide clear and timely updates regarding progress, risks, dependencies and blockers.
- Engage directly with technical and business stakeholders to clarify requirements and agree practical solutions.
- Explain technical concepts and trade-offs in a manner appropriate to the relevant stakeholder.
- Maintain practical technical documentation, including runbooks, architecture notes, change logs and release notes.
- Take accountability for the successful delivery and operational support of assigned solutions.
- Automate recurring data engineering and operational activities where practical.
- Implement monitoring and alerting to identify data quality issues, pipeline failures and abnormal processing behaviour.
- Analyse and optimise query, notebook and pipeline performance across SQL Server and Microsoft Fabric.
- Monitor capacity and resource utilisation and contribute to scalability and cost-control decisions.
- Deploy solutions using appropriate CI/CD and controlled deployment practices.
- Apply data security best practices, including secure authentication, least-privilege access and appropriate encryption.
- Ensure data engineering solutions comply with applicable data governance policies and regulatory requirements.
- Apply sound engineering practices relating to recoverability, auditability, supportability and controlled change.
- Protect confidential and sensitive business information.
- Collaborate with developers, data analysts, data scientists and business stakeholders to understand requirements and deliver practical solutions.
- Support effective handover and knowledge transfer to reduce key-person dependency within the data environment.
- Share technical knowledge and contribute to continuous improvement of team practices and the data environment.
- Provide guidance and support to junior team members where required.
- Remain accountable for the quality, reliability and timeliness of own deliverables.
- Document data processes, transformations, dependencies and architectural decisions.
- Validate data outputs through reconciliation, data quality checks and appropriate testing before production deployment.
- Maintain high standards of engineering quality by following agreed development, code review, testing, deployment, backup and archival practices.
- Ensure changes are appropriately tested, documented and controlled before implementation.
- Safeguard confidential information and data.
- Support compliance with applicable organisational policies, standards and regulatory requirements.
Market related
Skills
- Analytics
- API
- Authentication
- Automation
- Azure
- Azure DevOps
- CI/CD
- Data Engineering
- Data Governance
- Data Ingestion
- Data Modeling
- Data Pipelines
- Data Quality
- Data Warehousing
- DevOps
- Dimensional Modeling
- ELT
- ETL
- Git
- Lakehouse
- Microsoft Fabric
- Power BI
- PySpark
- Python
- REST
- Snowflake
- SOAP
- SQL
- SQL Server
- Version Control