Senior Azure Data Engineer
Summary
Design, develop, and optimize scalable data pipelines and ETL workflows using Python, PySpark, Databricks, and Microsoft Azure, collaborating with cross-functional data teams.
We are looking for an experienced Data Engineer with strong expertise in Python, PySpark, Databricks, and Microsoft Azure. The ideal candidate will be responsible for designing, developing, and optimizing scalable data processing solutions, ETL pipelines, and data engineering workflows while ensuring high standards of code quality, performance, and maintainability.
Responsibilities
- Design, develop, and maintain scalable data pipelines and ETL workflows using Python and PySpark.
- Develop modular, reusable, maintainable, and testable Python code for data engineering applications.
- Build and optimize large-scale batch data processing solutions using Apache Spark/PySpark.
- Develop data ingestion, transformation, cleansing, and integration pipelines.
- Implement data processing solutions using Databricks and Azure cloud services.
- Optimize Spark jobs, data pipelines, queries, and overall processing performance.
- Work with large datasets and implement efficient data transformation and aggregation techniques.
- Monitor, troubleshoot, and resolve issues across data pipelines and processing workflows.
- Implement data quality checks, validation, error handling, and logging mechanisms.
- Collaborate with Data Architects, Data Scientists, BI teams, and other stakeholders to deliver data solutions.
- Follow software engineering best practices including version control, code reviews, unit testing, and CI/CD.
- Contribute to the design and implementation of scalable and reliable cloud-based data platforms.
- Ensure data pipelines meet security, reliability, performance, and scalability requirements.
Requirements
- 6–8 years of professional experience in Data Engineering or a related field.
- Advanced proficiency in Python programming.
- Strong understanding of modular, maintainable, reusable, and testable code development.
- Hands‑on experience with Apache Spark and PySpark for large-scale data processing.
- Strong experience developing batch processing and ETL workflows.
- Hands‑on experience with Databricks.
- Hands‑on experience with Microsoft Azure and Azure data services.
- Strong SQL skills and experience working with relational and large-scale data environments.
- Experience with data pipeline development, transformation, and integration.
- Good understanding of distributed data processing and performance optimization.
- Experience with Git and software development best practices.
- Strong analytical, troubleshooting, and problem‑solving skills.