Senior Data Engineer
Posted Updated 1
view
Summary
Senior Data Engineer designing and managing containerized data pipelines, lakehouse architectures, and ETL/ELT workflows using SQL Server, Python, Kubernetes, Apache Spark, and Airflow in an on-premise environment.
. Bachelor's degree in Computer Science, IT, Engineering, or a related field with a demonstrated continuous learning ethos.
. Must have a minimum of 8+ years IT experience with at least 5+ years hands-on data engineering or datapipeline development
. Expert-level SQL proficiency with strong expertise in SQL Server, including queryoptimization, indexing, and performance tuning
. Advanced Python programming skills for data processing, automation, and production-grade pipeline development
.
Kubernetes expertise
- Design, deploy, andmanage containerized data pipelines in on-premise environments .
Strong data modellingexpertise - Both relational and non-relational concepts . Proven experience with flexible lakehouse/data lake architecture - Multi-layer datalakes, partitioning strategies, and metadata management, Iceberg tables, and optimization .
CI/CD and DevOps practices-
Setting up CI/CD pipelines, Git, automated testing, andinfrastructure-as-code tools .
ETL/ELT orchestration experience-Apache
Airflow or similar tools for scheduling and monitoring batch and real-time jobs . Hands-on experience with at least one NoSQL database (MongoDB, Cassandra, etc.) . Hands-on experience with Apache Spark and PySpark for distributed data processing andperformance optimization .
Data security andgovernance - Role-based access control, data masking, and compliance frameworks . Proven ability to work autonomously on complex projects while maintaining high codequality standards . Excellent problem-solving, communication, and cross-functional collaboration skills
Kubernetes expertise
- Design, deploy, andmanage containerized data pipelines in on-premise environments .
Strong data modellingexpertise - Both relational and non-relational concepts . Proven experience with flexible lakehouse/data lake architecture - Multi-layer datalakes, partitioning strategies, and metadata management, Iceberg tables, and optimization .
CI/CD and DevOps practices-
Setting up CI/CD pipelines, Git, automated testing, andinfrastructure-as-code tools .
ETL/ELT orchestration experience-Apache
Airflow or similar tools for scheduling and monitoring batch and real-time jobs . Hands-on experience with at least one NoSQL database (MongoDB, Cassandra, etc.) . Hands-on experience with Apache Spark and PySpark for distributed data processing andperformance optimization .
Data security andgovernance - Role-based access control, data masking, and compliance frameworks . Proven ability to work autonomously on complex projects while maintaining high codequality standards . Excellent problem-solving, communication, and cross-functional collaboration skills