freehire launches on Product Hunt on 26 August.

Follow →

Data Engineer

Summary

Build and maintain scalable data pipelines and infrastructure to power analytics and ML systems, ensuring high-quality, production-ready datasets.

We are looking for a skilled and passionate Data Engineer to design, build, and maintain scalable data infrastructure that powers our analytics, machine learning, and operational systems. You will work closely with software engineers, and business stakeholders to turn raw, complex data into reliable, production-ready pipelines and datasets. This is a hands-on, high-impact role for someone who thrives in a fast-paced environment and takes pride in data quality, reliability, and engineering excellence.

Essential Duties and Responsibilities

Pipeline Design & Data Ingestion

  • Design and build robust ETL/ELT pipelines for both batch and real-time processing, ensuring high throughput and fault tolerance.
  • Ingest and process high-frequency data from diverse sources including REST/GraphQL APIs, relational and NoSQL databases, IoT/sensor streams, and event queues.
  • Transform raw, messy, and heterogeneous data into clean, validated, and production-ready datasets for downstream consumers.
  • Architect and maintain data lake and data warehouse structures, including partitioning strategies, schema evolution, and data versioning.
  • Evaluate, select, and integrate appropriate data tools and frameworks based on project requirements and scale.

ML & Analytics Support

  • Support and deploy machine learning-ready pipelines, including feature engineering, data preprocessing, and preparation of model inputs and training datasets.
  • Collaborate with data scientists to productionize ML workflows, ensuring reproducibility and scalability of pipelines.
  • Build and maintain feature stores and data marts that support analytical dashboards, reporting, and model serving.
  • Develop and maintain data catalogues and metadata management systems to improve data discoverability and lineage.

Data Quality & Reliability

  • Implement comprehensive data quality checks, validation frameworks, and alerting mechanisms across all pipelines.
  • Monitor pipeline health, throughput, and latency; proactively diagnose and resolve data issues before they impact downstream systems.
  • Design for scalability and reliability in production, applying best practices for idempotency, retry logic, and graceful failure handling.
  • Conduct root cause analysis on data incidents and implement corrective and preventive measures.

Infrastructure & Orchestration

  • Deploy and manage pipeline orchestration using tools such as Apache Airflow, Prefect, or Dagster.
  • Work with cloud platforms (AWS, GCP, or Azure) to provision and manage data infrastructure including storage, compute, and streaming services.
  • Implement CI/CD practices for data pipelines, including automated testing, version control, and deployment pipelines.
  • Collaborate with DevOps and platform teams to containerize and deploy data workloads using Docker and Kubernetes.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available