Senior Data Engineer - PySpark, Databricks & SQL
Summary
Build and optimize PySpark/Databricks data pipelines and SQL queries to transform raw data into clean, governed datasets for analytics and dashboards in a government-focused tech environment.
We are looking for an experienced Senior Data Engineer to join our dynamic team in Singapore. The ideal candidate will have a strong background in data engineering with expertise in PySpark, Databricks, and SQL to build robust data pipelines and enable data-driven decision-making.
What You'll Do:
- Design and implement scalable data processing frameworks using PySpark and Databricks.
- Optimize existing ETL processes for performance and efficiency.
- Collaborate with cross-functional teams to understand data requirements and deliver solutions.
- Monitor data quality and troubleshoot issues within the data pipeline.
- Mentor junior engineers and contribute to best practices in data engineering.
- Stay updated with the latest industry trends and technologies related to big data.
What We're Looking For:
- Minimum of 5 years of experience in data engineering or related fields.
- Proficiency in PySpark, Databricks, and SQL is essential.
- Strong analytical skills with the ability to solve complex problems.
- Excellent communication skills for effective collaboration with stakeholders.
- Experience with cloud platforms (e.g., AWS, Azure) is a plus.
Detailed Job Description:
Data Transformation
- Query, clean, and transform datasets using SQL and Python on the Databricks platform.
- Ensure data quality and consistency across systems in accordance with government Instruction Manual (IM8) standards.
- Build, enhance, and maintain dashboards for data monitoring, performance reporting, and user behaviour analytics.
- Partner with business stakeholders to understand data requirements and translate them into impactful visual insights.
Documentation & Knowledge Management
- Develop clear and comprehensive user guides for dashboards, charts, and data workflows in Databricks.
- Maintain technical documentation (e.g., on Confluence) for analytics processes, standards, and best practices.
Additional Opportunities
- Explore and prototype automation solutions to streamline repetitive or manual workflows.
- Creating user stories and acceptance criteria
- Managing JIRA tickets and sprint tasks
- Writing and executing QA test cases
Required Skills
- Proficiency in PySpark, SQL and Python for data processing and automation.
- Hands-on experience in dashboard tools such as Tableau, Power BI, or equivalent visualisation tools.
- Familiarity with Databricks or similar big data/analytics platforms.
- Ability to produce high-quality technical documentation.
AdditionalSkillsPreferred
- Experience in system testing, SIT/UAT, or quality assurance processes.
- Knowledge of Spark for distributed data processing.
- Understanding of Git or other version control systems.
Soft Skills
- Strong analytical mindset and problem-solving abilities.
- Good communication skills, especially in explaining technical concepts to non-technical users.
- Ability to work independently as well as collaboratively in cross-functional teams.
- Attention to detail and commitment to delivering high-quality work.