freehire launches on Product Hunt on 26 August.

Follow →

Data Engineer

Summary

Build and maintain cloud data pipelines in Python, using AWS services and IaC to ingest, process, and model data for analytics and operations.

Overview

Design, build, and maintain cloud-based data pipelines and workflows that support analytics and operational systems. Integrate data from various sources using APIs and cloud services. Develop clean, efficient, and test-driven code in Python for data ingestion and processing. Optimize data storage and retrieval using big data formats like Apache Parquet and ORC. Implement robust data models, including relational, dimensional, and NoSQL models. Collaborate with cross-functional teams to gather and refine requirements and deliver high-quality solutions. Deploy infrastructure using Infrastructure as Code (IaC) tools such as AWS CloudFormation or CDK. Monitor and orchestrate workflows using Apache Airflow or Dagster. Follow best practices in data governance, quality, and security.

Responsibilities

  • Design, build, and maintain cloud-based data pipelines and workflows to support analytics and operational systems.
  • Integrate data from various sources using APIs and cloud services.
  • Develop clean, efficient Python code for data ingestion and processing with test-driven development.
  • Optimize data storage and retrieval using big data formats (Parquet, ORC).
  • Implement robust data models, including relational, dimensional, and NoSQL models.
  • Collaborate with cross-functional teams to gather requirements and deliver high-quality solutions.
  • Deploy infrastructure using IaC tools (AWS CloudFormation or CDK).
  • Monitor and orchestrate data workflows using Apache Airflow or Dagster.
  • Follow best practices in data governance, quality, and security.

Qualifications & Experience

  • Experience: At least 5 years in a cloud data engineering role working on data integration, processing, and transformation using open-source languages (e.g., Python) and cloud technologies.
  • Strong programming skills in Python for API integration and data libraries, with emphasis on quality and test-driven development.
  • Proficiency with big data storage formats (Apache Parquet, ORC) and knowledge of optimization strategies.
  • Proficiency with SQL.
  • Experience with data modeling: relational, dimensional, and NoSQL modeling.
  • Working knowledge of IaC on AWS (CloudFormation or CDK).
  • Working knowledge of AWS Services: Glue, IAM, Lambda, DynamoDB, Step Functions, S3, CloudFormation or CDK. Nice-to-have: Athena, Kinesis, MSK, MWAA, SQS.
  • Experience with data orchestration of pipelines: Apache Airflow or Dagster. Nice-to-have: data streaming (Kinesis, Kafka), Apache Spark.
  • Client-facing experience, multi-cultural team experience, technical leadership, or team leadership is a plus.

Working Arrangements

  • Must be willing to work in a hybrid setup, with onsite reporting to UP Ayala Technohub, Quezon City.
  • Available ASAP.
  • Engagement is project-based for 6 months, with a possibility of extension for good performance.
  • Work schedule is mid-shift and graveyard rotation.
  • The role observes U.S. holidays instead of Philippine holidays.

See also