Python Developer – Data Engineering
Summary
Python developer on a data engineering team in Toronto (hybrid, 4 days in office), building and optimizing high-performance data pipelines for large-scale datasets using pandas, polars, Dask, Docker/Kubernetes, ClickHouse, and NATS, with tests written in pytest.
Python Developer – Data Engineering | Pandas, Polars, Docker, Kubernetes, Dask
We are seeking a skilled Python Developer to join our data engineering team. You will design, develop, and maintain high-performance data processing pipelines using modern Python frameworks and tools. In this role, youll work with large-scale datasets, containerized systems, and distributed computing platforms to deliver robust data solutions.
Key Responsibilities
Required Skills and Experience
Python & Data Processing: Advanced proficiency in pandas and polars for data manipulation, transformation, and analysis. Experience optimizing code performance for large datasets.
Containerization & Orchestration: Hands-on experience with Docker for building container images and composing multi-container applications. Knowledge of Kubernetes for container orchestration and deployment management.
Data Infrastructure: Working knowledge of ClickHouse or similar columnar databases for OLAP workloads and analytical queries.
Messaging & Streaming: Familiarity with NATS.io for building message-driven systems and asynchronous workflows.
Testing & Quality Assurance: Proficiency with pytest for writing unit tests, integration tests, and maintaining code coverage standards.
Distributed Computing: Experience with Dask for parallel processing and handling out-of-core computations.
Version Control: Strong command of Git workflows, branching strategies, and collaborative development practices.
Preferred Qualifications