Data Engineer
Summary
Data Engineer building and maintaining scalable data pipelines, debugging production data issues, optimizing SQL queries, and implementing data ingestion/storage strategies using Python, Pandas, SQL, and cloud platforms.
Key Responsibilities
- Diagnose and resolve data issues, production outages, and performance bottlenecks to ensure system reliability.
- Debug data job failures, including pipeline breakdowns and unexpected changes in row-level data.
- Optimize database queries and monitor database health to enhance efficiency and scalability.
- Develop automated workflows for data regeneration to maintain accuracy and consistency.
- Enable secure and efficient data access for analysis, balancing openness with compliance requirements.
- Stay updated on industry trends and recommend best-in-class data engineering technologies and practices.
- Collaborate with the Data Lead and engineering team to implement optimal data ingestion and storage strategies.
- Design and build scalable, robust, and maintainable data pipelines using cutting-edge technologies in consultation with the Data Lead.
Requirements
- Minimum of 4 years of relevant working experience preferred.
- Strong proficiency in Python, with extensive hands-on experience using Pandas for data manipulation, transformation, and analysis of large datasets.
- Expertise in SQL performance tuning, database optimization, and query efficiency.
- Hands-on experience developing, deploying, and debugging robust data pipelines and ETL/ELT workflows in production environments.
- Experience with at least one cloud platform (AWS, GCP, or Azure) and data warehousing solutions.
- Strong problem-solving skills, particularly in navigating ambiguous or unknown data issues.
- Excellent communication skills with the ability to collaborate effectively with engineering and product teams.