Data Engineer
Summary
A founding-team Data Engineer in Selangor who owns the end-to-end data architecture: building ETL/ELT pipelines, a centralized data store, and data quality/monitoring, while refactoring Python/Jupyter workflows into production code. Core stack is Python (pandas/PySpark), SQL, and orchestrators like Airflow, scaling toward cloud warehousing.
In this role you will take ownership of the end-to-end data architecture from ingestion and transformation to storage and governance. You’ll work closely with analysts and stakeholders to design efficient data models, automate recurring workflows, and introduce best practices in data management, version control, and reproducibility.
This role is ideal for someone who thrives in a startup-like environment, enjoys solving complex data challenges, and wants to make a direct impact by shaping how data is managed and used across the organization.
Key Responsibilities
Design, build, and maintain ETL/ELT pipelines for automated data ingestion, transformation, and validation.
Establish a centralized data storage system, starting with lightweight solutions and scaling toward cloud-based infrastructure.
Collaborate with analysts and business teams to translate reporting needs into scalable, reusable data models.
Implement data quality checks, logging, and monitoring to ensure reliability and transparency of data flows.
Refactor existing Python/Jupyter workflows into structured, production-ready codebases.
Evaluate and recommend tools for data warehousing, orchestration, and cloud migration.
Define and promote best practices in version control, documentation, and reproducibility.
Requirements
Strong proficiency in Python (pandas, PySpark, or similar libraries for data processing).
Solid understanding of SQL and experience with relational databases.
Hands-on experience designing or maintaining ETL pipelines and workflow orchestration tools (e.g., Airflow, Prefect, Luigi).
Good grasp of data modeling, schema design, and data governance principles.
Experience handling large datasets and optimizing data performance.
Comfortable working in environments with minimal existing infrastructure and building systems from scratch.
Strong analytical thinking, problem-solving skills, and ability to work with both technical and non-technical stakeholders.
Nice to Have
Experience with cloud platforms (AWS, GCP, or Azure) and cloud-native data services.
Knowledge of data warehousing technologies (Snowflake, BigQuery, Redshift, or Databricks).
Familiarity with CI/CD pipelines, Docker, or containerized workflows.
Exposure to business intelligence tools (Power BI, Tableau, or similar).
What’s In It For You
13th Month Salary & Performance Bonuses
Birthday Leave + Birthday gift – Your day, your way
Growth Opportunities – Join the founding team and grow with us
Startup Culture – Fast-paced, collaborative, and packed with good vibes
Your application will include the following questions:
- How many years' experience do you have as a Data Engineer?
- Which of the following programming languages are you experienced in?
- How many years' experience do you have using SQL queries?
- What's your expected monthly basic salary?
- Which of the following languages are you fluent in?
- How much notice are you required to give your current employer?