freehire launches on Product Hunt on 26 August.

Follow →

Senior Data Engineer

Summary

Designs and maintains scalable data pipelines using Airflow, Spark, Kafka, and Iceberg tables, containerized via Docker/Kubernetes and deployed through GitLab CI/CD.

Follow us on LinkedIn to get job related updates.

A technology-driven software company delivering tailored digital solutions to startups, SMEs, and large enterprises. Their expertise spans product development, system integration, and scalable software engineering.

Role Introduction

Builds and operates end-to-end data pipelines , ingestion, transformation, and orchestration , across the lakehouse stack, deploying through CI/CD onto containerized infrastructure.

Features

  • Onsite
  • Fulltime

Requirements

  • Design and build batch and streaming ingestion pipelines using Airflow, Kafka, and Spark.
  • Develop transformation logic and ETL/ELT workflows using Informatica IDMC alongside custom Spark jobs where needed.
  • Containerize pipeline code and deploy via GitLab CI/CD onto Dockerized/Kubernetes infrastructure.
  • Write and maintain Airflow DAGs with proper dependency management, retries, and SLA monitoring.
  • Implement data quality checks and validation logic at each stage of the pipeline.
  • Optimize Spark jobs for performance and cost (partitioning, caching, shuffle management) writing into Iceberg tables.
  • Collaborate with the Data Modeler and Data Architect to ensure pipeline output matches target schema and table-format requirements.
  • Troubleshoot production pipeline issues and participate in on-call/SLA support rotations as needed.

Specifications

  • 6+ years of hands-on data engineering experience building production pipelines, ideally with 3+ years specifically on Spark.
  • Strong working knowledge of Apache Airflow for orchestration , DAG design, sensors, and operational troubleshooting.
  • Production experience with Kafka , producers/consumers, schema registry, partitioning strategy.
  • Direct experience with Informatica IDMC (or PowerCenter/IICS background actively transitioning to IDMC) for managed ETL/ELT.
  • Comfort with GitLab CI/CD pipelines and Docker/Kubernetes-based deployment of data workloads.
  • Strong Python and SQL skills; ability to read/write Spark (PySpark) jobs independently.
  • Experience writing to and reading from open table formats (Iceberg, Delta Lake, or Hudi) in a lakehouse setup.

Expertise

Skills: Airflow, Docker/Kubernetes, GitLab CI/CD, Informatica IDMC, Kafka, Spark

About TalentHue

TalentHue provides scalable, reliable Tech Recruitment, Corporate Recruitment and Consulting (Strategy, Operations, Performance) services. Our Recruitment and HR consultants will work alongside your team to meet the unique needs of your business.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available