Data Engineer/Developer with Python and SQL
Summary
A data engineer who designs, builds, and optimizes ETL pipelines using Python, Spark, and SQL, migrates legacy SSIS/SQL Server jobs to modern data platforms, and ensures data quality through testing while working with architects, scientists, and analysts in an agile team.
Key Responsibilities
- Design, develop, and implement new ETL (Extract, Transform, Load) jobs using Python and/or Spark to support various data initiatives.
- Migrate existing ETL processes from SSIS and SQL Server to modern data platforms, ensuring data integrity and performance.
- Maintain and optimize existing data pipelines and ETL processes for efficiency, reliability, and scalability.
- Develop and implement robust testing strategies for all ETL jobs to ensure data quality and accuracy.
- Collaborate with data architects, data scientists, and business analysts to understand data requirements and translate them into technical specifications.
- Participate actively in an agile development environment, including stand-ups, sprint planning, and retrospectives.
- Communicate effectively with users, stakeholders, and team members to gather requirements, provide updates, and resolve issues.
- Troubleshoot and resolve data-related issues and performance bottlenecks in a timely manner.
- Strong proficiency in Python and/or Apache Spark for data processing and ETL development.
- Strong SQL knowledge, with proven experience in writing complex queries, stored procedures, and optimizing database performance.