Senior Data Engineer
Summary
Designs and maintains scalable data pipelines using Airflow, Spark, Kafka, and Iceberg tables, containerized via Docker/Kubernetes and deployed through GitLab CI/CD.
Follow us on LinkedIn to get job related updates.
A technology-driven software company delivering tailored digital solutions to startups, SMEs, and large enterprises. Their expertise spans product development, system integration, and scalable software engineering.
Role Introduction
Builds and operates end-to-end data pipelines , ingestion, transformation, and orchestration , across the lakehouse stack, deploying through CI/CD onto containerized infrastructure.
Features
- Onsite
- Fulltime
Requirements
- Design and build batch and streaming ingestion pipelines using Airflow, Kafka, and Spark.
- Develop transformation logic and ETL/ELT workflows using Informatica IDMC alongside custom Spark jobs where needed.
- Containerize pipeline code and deploy via GitLab CI/CD onto Dockerized/Kubernetes infrastructure.
- Write and maintain Airflow DAGs with proper dependency management, retries, and SLA monitoring.
- Implement data quality checks and validation logic at each stage of the pipeline.
- Optimize Spark jobs for performance and cost (partitioning, caching, shuffle management) writing into Iceberg tables.
- Collaborate with the Data Modeler and Data Architect to ensure pipeline output matches target schema and table-format requirements.
- Troubleshoot production pipeline issues and participate in on-call/SLA support rotations as needed.
Specifications
- 6+ years of hands-on data engineering experience building production pipelines, ideally with 3+ years specifically on Spark.
- Strong working knowledge of Apache Airflow for orchestration , DAG design, sensors, and operational troubleshooting.
- Production experience with Kafka , producers/consumers, schema registry, partitioning strategy.
- Direct experience with Informatica IDMC (or PowerCenter/IICS background actively transitioning to IDMC) for managed ETL/ELT.
- Comfort with GitLab CI/CD pipelines and Docker/Kubernetes-based deployment of data workloads.
- Strong Python and SQL skills; ability to read/write Spark (PySpark) jobs independently.
- Experience writing to and reading from open table formats (Iceberg, Delta Lake, or Hudi) in a lakehouse setup.
Expertise
Skills: Airflow, Docker/Kubernetes, GitLab CI/CD, Informatica IDMC, Kafka, Spark
About TalentHue
TalentHue provides scalable, reliable Tech Recruitment, Corporate Recruitment and Consulting (Strategy, Operations, Performance) services. Our Recruitment and HR consultants will work alongside your team to meet the unique needs of your business.