Data Engineer (PySpark, Apache Spark, Delta Lake, ETL, Oracle, DB2, AWS, TWS, Lambda, IAM)
NOVACLOUD SYSTEMS PTE. LTD. Data Engineer (PySpark, Apache Spark, Delta Lake, ETL, Oracle, DB2, AWS, TWS, Lambda, IAM)
Responsibilities
- Design, develop, and maintain scalable data pipelines for ingesting, transforming, validating, and delivering structured and semi-structured data.
- Develop data engineering solutions using Databricks, PySpark, and Apache Spark for large-volume data processing and transformation.
- Build and manage ETL/ELT workflows using modern data engineering platforms as well as enterprise ETL technologies.
- Develop data transformation and processing logic using SQL and Python, with a focus on performance, reliability, and data accuracy.
- Work with Delta Lake and Delta Live Tables to develop and maintain reliable data pipelines and curated data layers.
- Develop and optimize ETL workflows using IBM DataStage and Informatica PowerCenter where required.
- Work with relational databases including Oracle and IBM DB2 for data extraction, transformation, loading, querying, and performance optimization.
- Implement data integration solutions across cloud and enterprise data platforms.
- Develop and support data pipelines and associated services on AWS, including data storage, processing, monitoring, and related cloud services.
- Perform data validation, reconciliation, quality checks, and troubleshooting to ensure the accuracy and completeness of datasets.
- Optimize Spark jobs, SQL queries, ETL workflows, and data pipelines to improve processing efficiency and overall performance.
- Collaborate with technical and business teams to understand data requirements and translate them into scalable data solutions.
- Follow established development, deployment, documentation, data governance, and SDLC practices.
- Monitor production pipelines, investigate failures, and support timely resolution of data processing issues.
Requirements
- 5+ years of experience in data engineering, ETL development, data integration, or a related field.
- Strong hands-on experience with Databricks and data engineering workloads.
- Hands-on experience in PySpark and Apache Spark for distributed data processing.
- Solid experience developing ETL/ELT pipelines and data engineering solutions.
- Strong experience SQL skills, including complex queries, joins, aggregations, optimization, and data transformation.
- Hands on experience with AWS cloud services used for data engineering and data processing.
- Experience with Delta Lake, Delta Live Tables (DLT) is preferred.
- Hands-on experience with IBM DataStage and Informatica PowerCenter.
- Experience working with relational databases such as Oracle and IBM DB2.
- Hand on experience in Python for data processing, automation, or pipeline development.
- Experience with data modelling, data warehousing, data quality, and data integration concepts.
- Hands on experience in Lambda, Redshift, IAM, CloudWatch, Glue, EC2.
- Experience with production scheduling, pipeline monitoring, troubleshooting, and deployment processes.
- Experience in version control and CI/CD practices such as TWS, GIT, Jenkins ect.
- Ability to work effectively in an Agile/SDLC environment and collaborate with cross-functional teams.
- Data brick certification would be preferred.