Data Engineer
Summary
Design, optimize, and maintain scalable data pipelines using Python, SQL, and PySpark, collaborating with Data Science teams to deploy and operationalize ML solutions in production.
We are looking for a Mid-Level Data Engineer to design, optimize, and maintain complex data pipelines, ensuring the scalability, reliability, and performance of data solutions. In this role, you will also collaborate with Data Science teams to integrate and operationalize machine learning models in production environments.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines.
- Optimize data processing workflows for large-scale datasets.
- Build and manage data ingestion solutions from multiple sources.
- Ensure data quality, reliability, and performance across platforms.
- Collaborate with Data Science and Engineering teams to deploy and operationalize ML solutions.
- Contribute to software engineering best practices and continuous improvement initiatives.
- Proven experience of 3+ years in Data Engineering.
- Advanced proficiency in Python and SQL.
- Demonstrated ability to design robust, maintainable, and highly scalable solutions for large-scale data processing and performance optimization.
- Hands-on experience with in-memory data processing libraries such as Polars (or equivalent).
- Solid experience with distributed data processing frameworks, particularly PySpark.
- Experience developing data ingestion pipelines leveraging REST APIs.
- Strong knowledge of software development best practices.
- Experience with Git, CI/CD pipelines, and software testing methodologies.
- Experience deploying, monitoring, and maintaining Machine Learning pipelines in production environments.
Nice to have:
- Experience with modern data platforms such as Databricks.
- Experience orchestrating workflows using tools such as Apache Airflow.
- Knowledge of cloud platforms (AWS, Azure, or GCP).