Senior Data Engineer
Summary
Designs and maintains data pipelines using Spark, Airflow, Kafka, and Iceberg tables, deploying via GitLab CI/CD on Kubernetes.
Role Introduction
Builds and operates end-to-end data pipelines, ingestion, transformation, and orchestration across the lakehouse stack, deploying through CI/CD onto containerized infrastructure.
Features
- Onsite
Requirements
- Design and build batch and streaming ingestion pipelines using Airflow, Kafka, and Spark.
- Develop transformation logic and ETL/ELT workflows using Informatica IDMC alongside custom Spark jobs where needed.
- Containerize pipeline code and deploy via GitLab CI/CD onto Dockerized/Kubernetes infrastructure.
- Write and maintain Airflow DAGs with proper dependency management, retries, and SLA monitoring.
- Implement data quality checks and validation logic at each stage of the pipeline.
- Optimize Spark jobs for performance and cost (partitioning, caching, shuffle management) writing into Iceberg tables.
- Collaborate with the Data Modeler and Data Architect to ensure pipeline output matches target schema and table‑format requirements.
- Troubleshoot production pipeline issues and participate in on‑call/SLA support rotations as needed.
Specifications
- 6+ years of hands‑on data engineering experience building production pipelines, ideally with 3+ years specifically on Spark.
- Strong working knowledge of Apache Airflow for orchestration, DAG design, sensors, and operational troubleshooting.
- Production experience with Kafka, producers/consumers, schema registry, partitioning strategy.
- Direct experience with Informatica IDMC (or PowerCenter/IICS background actively transitioning to IDMC) for managed ETL/ELT.
- Comfort with GitLab CI/CD pipelines and Docker/Kubernetes‑based deployment of data workloads.
- Strong Python and SQL skills; ability to read/write Spark (PySpark) jobs independently.
- Experience writing to and reading from open table formats (Iceberg, Delta Lake, or Hudi) in a lakehouse setup.