Senior Data Engineer
Posted Updated
Summary
Senior Data Engineer in Singapore designing, building, and managing production data pipelines and lakehouse architecture on-premise. Core stack includes SQL Server, Python, Kubernetes, Apache Spark/PySpark, Airflow, and Iceberg data lakes with CI/CD and data governance practices.
Bachelor's degree in Computer Science, IT, Engineering, or a related field with a demonstrated continuous learning ethos.
Must have a minimum of 8+ years IT experience with at least 5+ years hands-on data engineering or datapipeline development
Expert-level SQL proficiency with strong expertise in SQL Server, including queryoptimization, indexing, and performance tuning
Advanced Python programming skills for data processing, automation, and production-grade pipeline development
Kubernetes expertise
– Design, deploy, andmanage containerized data pipelines in on-premise environments Strong data modellingexpertise – Both relational and non-relational concepts Proven experience with flexible lakehouse/data lake architecture – Multi-layer datalakes, partitioning strategies, and metadata management, Iceberg tables, and optimization CI/CD and DevOps practices–
Setting up CI/CD pipelines, Git, automated testing, andinfrastructure-as-code tools ETL/ELT orchestration experience—Apache
Airflow or similar tools for scheduling and monitoring batch and real-time jobs Hands-on experience with at least one NoSQL database (MongoDB, Cassandra, etc.) Hands-on experience with Apache Spark and PySpark for distributed data processing andperformance optimization Data security andgovernance – Role-based access control, data masking, and compliance frameworks Proven ability to work autonomously on complex projects while maintaining high codequality standards Excellent problem-solving, communication, and cross-functional collaboration skills
– Design, deploy, andmanage containerized data pipelines in on-premise environments Strong data modellingexpertise – Both relational and non-relational concepts Proven experience with flexible lakehouse/data lake architecture – Multi-layer datalakes, partitioning strategies, and metadata management, Iceberg tables, and optimization CI/CD and DevOps practices–
Setting up CI/CD pipelines, Git, automated testing, andinfrastructure-as-code tools ETL/ELT orchestration experience—Apache
Airflow or similar tools for scheduling and monitoring batch and real-time jobs Hands-on experience with at least one NoSQL database (MongoDB, Cassandra, etc.) Hands-on experience with Apache Spark and PySpark for distributed data processing andperformance optimization Data security andgovernance – Role-based access control, data masking, and compliance frameworks Proven ability to work autonomously on complex projects while maintaining high codequality standards Excellent problem-solving, communication, and cross-functional collaboration skills