Junior Data Engineer
Summary
The Junior Data Engineer will assist in developing and maintaining ETL/ELT pipelines, writing complex SQL queries, and ensuring data quality. The role requires strong proficiency in Python and SQL, with a focus on collaborating with cross-functional teams to build robust data architectures.
- Assist in developing, testing, and deploying robust ETL/ELT pipelines integrating data from various source systems into centralized data repositories
- Write efficient SQL queries, create views, and maintain database schemas
- Implement automated data validation routines and quality checks for data accuracy, consistency, and completeness
- Partner with Data Scientists, Software Engineers, and Business Analysts to translate data requirements into efficient data structures
- Document data models, pipeline architectures, and data flows
- Assist in monitoring running jobs and troubleshooting operational issues
Requirements
- Bachelor's degree in Computer Science, Information Technology, Data Engineering, Software Engineering, or a related quantitative field, or equivalent hands-on experience
- Strong proficiency in Python for data manipulation, scripting, and pipeline automation
- Familiarity with pandas, PySpark, or NumPy
- Advanced proficiency in SQL, including complex queries, joins, aggregations, and performance-tuned scripts
- Solid understanding of software engineering principles, version control, data modeling, and relational and non-relational database architectures
- Strong analytical mindset, attention to detail, and proactive troubleshooting approach
- Good verbal and written communication skills
- Hands-on exposure to AWS preferred
- Experience with Databricks and Apache Spark (PySpark) preferred
- Familiarity with DAG-based orchestration tools preferred
- Familiarity with dbt, Snowflake, or BigQuery preferred
Core Competencies
Proficient in developing and deploying ETL/ELT pipelines, with strong expertise in SQL and Python for data manipulation and automation. Capable of collaborating with cross-functional teams to ensure data accuracy and integrity while documenting data architectures and flows.
Highest-signal resume keywords
- ETL/ELT Pipeline Development
- Advanced SQL Proficiency
- Python Scripting for Data Automation
- Data Modeling and Database Architecture
- AWS Exposure
ATS Optimization Keywords
Hard Skills
- SQL
- Python
- Data Modeling
- ETL/ELT Development
- Data Validation
- Database Schema Maintenance
- Data Pipeline Automation
- Complex Query Writing
- Performance Tuning
- Data Analysis
Soft Skills
- Analytical Mindset
- Attention to Detail
- Proactive Troubleshooting
- Verbal Communication
- Written Communication
Industry Keywords
- Data Engineering
- Data Integration
- Data Repositories
- Data Quality Checks
- Software Engineering Principles
Tools & Technologies
- Pandas
- PySpark
- NumPy
- Databricks
- Apache Spark
- DAG-based Orchestration Tools
- Dbt
- Snowflake
- BigQuery