Junior Data Engineer
Summary
Build and maintain ETL/ELT pipelines and PySpark jobs to feed AI and analytics platforms; write SQL, validate data, and troubleshoot issues in a cloud-based data stack.
This is a fast-growing technology company that is building scalable data platforms to power AI, analytics, and next-generation digital products. As part of its continued expansion, the organization is seeking to appoint a Junior Data Engineer to support the development of reliable data pipelines and modern data infrastructure.
This is an excellent opportunity for an early-career data professional to work alongside experienced data engineers, software engineers, and data scientists, gaining hands‑on experience building enterprise‑scale data platforms and enabling AI‑driven products.
ResponsibilitiesYou will design, develop, and maintain ETL/ELT data pipelines that ingest, transform, and deliver data from multiple sources into enterprise data platforms. Working closely with data engineers, data scientists, and application development teams, you will ensure high-quality, reliable, and scalable data pipelines that support analytics, reporting, and AI initiatives.
You will write and optimize SQL queries, develop and optimize PySpark data processing jobs, perform data validation and quality checks, troubleshoot pipeline issues, and contribute to improving data reliability and performance. You will also assist with data modelling, documentation, pipeline monitoring, and automation while adopting modern data engineering practices and cloud technologies.
RequirementsWe are looking for a Junior Data Engineer with 1-3 years of experience in data engineering, database development, or ETL development within a technology or data-driven environment.
You should have strong SQL skills and hands‑on experience designing and maintaining ETL/ELT pipelines. Hands‑on experience with PySpark is required, including developing and optimizing distributed data processing pipelines and ETL workflows for large‑scale datasets.
Experience with Python, relational databases, and modern data platforms such as Databricks, Snowflake, Azure Data Factory, Apache Airflow, AWS Glue, or Apache Spark is highly desirable. Familiarity with cloud environments (AWS, Azure, or Google Cloud), data modelling, and data orchestration tools will be advantageous.
Candidates should possess strong analytical and problem‑solving skills, attention to data quality, and the ability to work collaboratively with data engineers, software engineers, and data scientists to build reliable, scalable data solutions.