freehire launches on Product Hunt on 26 August.

Follow →

Data Engineer

Summary

Build and operate cloud-native data platforms using Python, Spark, and AWS/GCP, designing scalable ETL/ELT pipelines and modern data warehouses.

Role Summary

We are seeking engineers who combine strong software engineering fundamentals with modern data engineering expertise. Candidates should be capable of designing, building, deploying and operating cloud-native data platforms while applying software engineering best practices throughout the delivery lifecycle.

Strong SQL and data warehousing experience remain important but should complement broader engineering capability rather than define the candidate's profile.

Minimum Technical Requirements (Non-Negotiable)

Candidates must demonstrate practical project experience with:

Software Engineering

  • Python as a primary programming language
  • Software engineering principles and clean coding practices
  • Object-oriented programming
  • Testing and code quality practices
  • Software Development Life Cycle (SDLC)
  • Git and collaborative development workflows
  • API development and integration
  • CI/CD pipelines
  • Production software deployment
  • ETL
  • AWS - native data services
  • Data stores
  • Workflow systems

Engineering Mindset

Candidates should demonstrate the ability to:

  • Solve business problems through code
  • Design scalable solutions
  • Work within engineering teams
  • Contribute to production systems
  • Follow engineering standards and best practices

Core Data Engineering Requirements

Candidates should have hands-on experience with several of the following:

Data Processing

  • Spark
  • PySpark
  • Databricks
  • Data Lake architectures
  • Batch processing
  • Streaming architectures
  • Data transformation frameworks
  • ETL/ELT design

Data Platform Technologies

  • Kafka
  • Flink
  • Delta Lake
  • Iceberg
  • Airflow
  • Modern orchestration platforms

Data Storage & Analytics

  • SQL
  • Data modelling
  • Data warehousing concepts
  • Relational databases
  • Analytical data platforms

Cloud Engineering Requirements

Candidates should have practical experience delivering solutions on at least one major cloud platform:

Preferred Order

  1. Google Cloud Platform (GCP)
  2. Amazon Web Services (AWS)
  3. Microsoft Azure

Typical Technologies

GCP

  • BigQuery
  • Dataflow
  • Dataproc
  • Pub/Sub
  • GKE
  • Cloud Storage

AWS

  • Glue
  • EMR
  • Redshift
  • Kinesis
  • EKS
  • S3

Azure

  • Data Factory
  • Synapse
  • Databricks
  • Event Hubs
  • AKS
  • Azure Storage

Cloud experience should reflect real project delivery rather than certifications alone.

DevOps & Platform Engineering

Strong candidates should also demonstrate exposure to:

Containerisation & Deployment

  • Docker
  • Kubernetes

Infrastructure

  • Terraform
  • Infrastructure as Code
  • Environment management

Operations

  • Monitoring
  • Observability
  • Logging
  • Production support

Data Modelling Requirements

Candidates should possess working knowledge of:

Data Architecture

  • Conceptual modelling
  • Logical modelling
  • Physical modelling

Data Warehousing

  • Dimensional modelling
  • Relational modelling
  • Data warehouse design
  • Data governance concepts

Data modelling should support modern platform development, not exist in isolation.

See also