freehire launches on Product Hunt on 26 August.

Follow →

Cloud Data Engineer

Summary

Build and maintain cloud data pipelines in Python, integrating APIs and AWS services to feed analytics and operational systems with clean, test-driven code.

Responsibilities

  • Design, build, and maintain cloud-based data pipelines and workflows that support analytics and operational systems.
  • Integrate data from various sources using APIs and cloud services.
  • Develop clean, efficient, and test-driven code in Python for data ingestion and processing.
  • Optimize data storage and retrieval using big data formats like Apache Parquet and ORC.
  • Implement robust data models, including relational, dimensional, and NoSQL models.
  • Collaborate with cross-functional teams to gather and refine requirements and deliver high-quality solutions.
  • Deploy infrastructure using Infrastructure as Code (IaC) tools such as AWS

Qualifications

  • At least 3 years in a data engineering role working on data integration, processing, and transformation use cases with open-source languages (i.e. Python) and cloud technologies.
  • Strong programming skills in Python specifically for API integration and data libraries, with emphasis on quality and test-driven development.
  • Demonstrated proficiency with big data storage formats (Apache Parquet, ORC) and practical knowledge of pitfalls and optimization strategies.
  • Demonstrated proficiency with SQL.
  • Experience with data modeling: relational modeling; dimensional modeling; NoSQL modeling.
  • Working knowledge of IaC on AWS (CloudFormation or CDK).
  • Working knowledge of AWS Services: Glue, IAM, Lambda, DynamoDB, Step Functions, S3, CloudFormation or CDK.
  • Must be willing to work on a hybrid setup, with onsite reporting to Quezon City.
  • Must be available to start by June 1, 2025.
  • Engagement is project-based for 6 months, with a possibility of extension.
  • Experience with Athena, Kinesis, MSK, MWAA, SQS.
  • Experience with orchestration of data flows/pipelines: Apache Airflow or Dagster.
  • Experience with data streaming (Kinesis, Kafka).
  • Experience with Apache Spark.
  • Client-facing experience, multi-cultural team experience, technical leadership, team leadership.

See also