freehire launches on Product Hunt on 26 August.

Follow →

Data Engineer: Python Pipelines, Airflow & AWS

Summary

Build and maintain Python-based ETL pipelines using Airflow and AWS, optimizing SQL for Snowflake and ensuring data quality with dbt tests.

Python Skills

  • Advanced Proficiency in Python concepts like Code Structures, Modules, Packages, Class, SubClass, Inheritance, Multi-Threading and Functional Programming.
  • Experience in developing reusable Python packages for internal or public usage.
  • Ability to write automating ETL processes and scheduling jobs like Airflow DAG.
  • Ability to track job pipeline runs to reprocess error records.
  • Ability to orchestrate different pipelines to run in sequence or parallel.
  • Troubleshoot data pipeline errors and fix issues.
  • Export or Import data to/from various formats like CSV, JSON, XML etc preferably from S3 or other cloud storage.
  • Experience in using AI IDE tool.

SQL

  • Advanced SQL skills, including complex joins, CTE's and subqueries.
  • Experience in optimizing SQL queries for performance and optimization in data warehouse technologies preferably Snowflake.

Testing and Documentation

  • Proficiency in Python unit, integration and system test.
  • Proficiency in implementing DBT tests for data validation and quality checks.

Code Generation

  • Experience in generating code using configurations using python and jinja templates.

Version Control

  • Experience in GitHub, including implementing CI/CD process from scratch.

AWS Expertise

  • In depth understanding of AWS S3 for data storage, ECS, IAM including best practices for organization and security.
  • Knowledge of AWS security best practices, including IAM roles, encryption standard, secure coding guidelines, DBT profiles access configurations and more.
  • Data Integration (nice to have): Experience with AWS lambda for serverless data processing tasks.

See also