Senior Data Engineer
Summary
Build and optimize cloud ETL pipelines (AWS Glue/PySpark) and legacy integrations (IBM DataStage, SAS, SAP) to power machine learning and analytics at a data-driven company.
Jora Malaysia will close on 9th September 2026. Thank you for being with us, we are cheering you on as you continue your career journey.
We are looking for a Senior Data Engineer to manage, modernize, and scale our core data infrastructure. In this role, you will bridge legacy enterprise data architectures and modern cloud environments . Managing pipelines across IBM DataStage, SAS ETL, SAP Data Integration, and AWS Glue. You will also build the data pipelines and feature stores required to support our growing machine learning and predictive analytics initiatives.
Key Responsibilities
- Cloud Pipeline Development: Architect, build, and optimize scalable ETL/ELT pipelines in AWS Glue (PySpark/Python), AWS Lambda, and Redshift/Athena.
- Legacy & ERP Integration: Maintain and optimize existing IBM DataStage and SAS ETL jobs, and orchestrate data extractions from SAP Data Integration environments into our cloud data platform.
- Data Pipelines for ML: Collaborate with data science and AI teams to design, build, and maintain data pipelines for machine learning model training, feature stores, and automated model scoring.
- Performance & Optimization: Identify pipeline bottlenecks, optimize complex SQL queries, and ensure efficient data movement across production databases and cloud storage.
- Cross-Team Collaboration: Partner closely with the Enterprise BI & Data Governance Specialist to ensure pipeline outputs are structured, documented, and migration- and Cognos-ready for downstream reporting and analytics use.
Key Qualifications
- Experience: 5+ years in Data Engineering, ETL Architecture, and Database Management.
- Cloud & Spark: 3+ years of hands-on experience with AWS Glue, PySpark/Spark, and AWS data ecosystems.
- Legacy Enterprise Tools: Hands-on experience with IBM DataStage, SAS ETL (Base SAS / SAS Data Integration Studio), and SAP Data Integration (BODS / SAP Data Services).
- Machine Learning Support: Proficient in Python (Pandas, PySpark, Scikit-Learn) with experience preparing data for ML workflows and MLOps pipelines.
- Database Mastery: Advanced SQL skills, database optimization, indexing strategies, and schema design (relational, columnar, star schema).
- Nice to Have: Exposure to IBM Cognos Analytics or other BI reporting layers, and familiarity with data governance/archiving concepts.
- Adaptability: Willingness and enthusiasm to learn new tools, platforms, and techniques as our data engineering and cloud stack evolves.