Senior Data Engineer
Summary
The Senior Data Engineer will design and manage autonomous, production-grade ETL/ELT pipelines using Python, PySpark, and SQL Server within a Kubernetes-based on-premise environment. The role focuses on building scalable lakehouse architectures, enforcing data governance, and mentoring junior team members.
Design and develop autonomous, production-grade ETL/ELT data pipelines using Python and PySpark that ingest, transform, and deliver high-quality data while maintaining integrity and performance standards.
Implement and manage flexible lakehouse architecture across raw, curated, and consumption layers, including data partitioning, cataloging, and metadata management.
Deploy and manage data pipelines using Kubernetes and Docker to ensure scalability, reliability, and efficient resource utilization in on-premise environments.
Leverage strong SQL Server expertise to design optimal data models, write complex queries, and perform query optimization across the data platform.
Establish and maintain robust CI/CD practices for data pipeline deployment, including automated testing, version control, and continuous monitoring.
Enforce security, governance, and role-based access controls across all data layers while ensuring compliance and auditability.
Mentor junior engineers, conduct code reviews, and establish best practices across the team.
Collaborate with Data Scientists, Business Analysts, and stakeholders to deliver datasets aligned with operational and analytical needs.
Provide L3 support and expert consultation for complex data challenges; evaluate and recommend new tools and practices to improve agility and performance.
Requirements
Must-have qualifications:
- 8+ years IT experience; 5+ years hands-on data engineering or data pipeline development.
- Expert-level SQL proficiency with strong expertise in SQL Server, including query optimization, indexing, and performance tuning.
- Advanced Python programming skills for data processing, automation, and production-grade pipeline development.
- Kubernetes expertise - Design, deploy, and manage containerized data pipelines in on-premise environments.
- Strong data modeling expertise - Both relational and non-relational concepts.
- Proven experience with flexible lakehouse/data lake architecture - Multi-layer data lakes, partitioning strategies, and metadata management, Iceberg tables and optimization.
- CI/CD and DevOps practices - Setting up CI/CD pipelines, Git, automated testing, and infrastructure-as-code tools.
- ETL/ELT orchestration experience - Apache Airflow or similar tools for scheduling and monitoring batch and real-time jobs.
- Hands-on experience with at least one NoSQL database (MongoDB, Cassandra, etc.).
- Hands-on experience with Apache Spark and PySpark for distributed data processing and performance optimization.
- Data security and governance - Role-based access control, data masking, and compliance sframeworks.
- Proven ability to work autonomously on complex projects while maintaining high code quality standards.
- Excellent problem-solving, communication, and cross-functional collaboration skills.
- Bachelor's degree in Computer Science, IT, Engineering, or related field with demonstrated continuous learning ethos.
Preferred qualifications:
- Experience with on-premise data virtualization or logical data warehouse concepts.
- Understanding of data mesh or data fabric architecture patterns.
- Real-time streaming technologies (Kafka, Apache Flink).
- Metadata management and data lineage tools.
- Experience mentoring junior engineers or leading technical initiatives.
- Agile delivery methodologies and product-oriented data architecture.
Other Professional Skills and Mind-set:
- Autonomous Work Ethic - Work independently on complex problems while proactively seeking collaboration.
- Continuous Learning - Committed to staying current with data engineering trends and best practices.
Skills
- Agile
- Airflow
- Automation
- Cassandra
- CI/CD
- Data Engineering
- Data Lake
- Data Lineage
- Data Modeling
- Data Pipelines
- Data Warehousing
- DevOps
- Docker
- ELT
- ETL
- Flink
- Git
- Infrastructure as Code
- Kafka
- Kubernetes
- Lakehouse
- Metadata Management
- MongoDB
- NoSQL
- PySpark
- Python
- Spark
- SQL
- SQL Server
- Test Automation
- Version Control
- Virtualization