Data Engineer (PH)
Summary
Build and maintain scalable data pipelines on Databricks and AWS, focusing on ETL, real-time streaming with Spark/Kafka, and cloud infrastructure optimization.
The Role
We are seeking an experienced Data Engineer to join our Data & AI & Risk team. This is a full-time, on-site position based in Taguig City, Metro Manila, Philippines (with opportunities in Malaysia and Indonesia).
Responsibilities
- Design, develop, and maintain highly scalable, reliable, and efficient data processing systems with a strong emphasis on code quality and performance.
- Collaborate closely with data analysts, software developers, and business stakeholders to deeply understand data requirements and architect robust solutions to address their needs.
- Focus on the development and maintenance of ETL pipelines, ensuring seamless extraction, transformation, and loading of data from diverse sources into our data warehouse based on Databricks platform.
- Spearhead the development and maintenance of real-time data processing systems utilizing cutting‑edge big data technologies such as Spark Streaming and Kafka.
- Establish and enforce rigorous data quality and validation checks to uphold the accuracy and consistency of our data assets.
- Act as a point of contact for troubleshooting and resolving complex data processing issues, collaborating with cross‑functional teams as necessary to ensure timely resolution.
- Proactively monitor and optimize data processing systems to uphold peak performance, scalability, and reliability standards, leveraging advanced AWS operational knowledge.
- Utilize AWS services such as EC2, S3, Glue, and Databricks to architect, deploy, and manage data processing infrastructure in the cloud.
- Implement robust security measures and access controls to safeguard sensitive data assets within the AWS environment.
- Stay abreast of the latest advancements in AWS technologies and best practices, incorporating new tools and services to continually improve our data processing capabilities.
Requirements
- Bachelor's or Master's degree in Computer Science or a related field.
- Minimum of 5 years of hands‑on experience as a Data Engineer, demonstrating a proven track record of designing and implementing sophisticated data processing systems.
- Good understanding of Databricks platform and Delta Lake.
- Familiar with data job scheduler tools such as Dagster.
- Proficiency in one or more programming languages such as Scala, Java, or Python.
- Deep expertise in big data technologies including Apache Spark for ETL processing and optimization.
- Proficient in utilizing BI tools such as Metabase for data visualization and analysis.
- Advanced understanding of data modeling, data quality, and data governance best practices.
- Outstanding communication and collaboration skills, with the ability to effectively engage with diverse stakeholders across the organization.
- Extensive experience in AWS operational management, including deployment, configuration, and optimization of data processing infrastructure within the AWS cloud environment.
- Strong understanding of AWS services such as EC2, S3, Glue, and EMR, with the ability to architect scalable and resilient data solutions leveraging these services.
- Proficiency in AWS security best practices, with experience implementing robust security measures and access controls to protect sensitive data assets.
- Hands‑on experience with automation and DevOps tools such as Terraform for infrastructure as code and automation purposes.
- Ability to read and write in English.