Data Engineering Specialist
Responsibilities
- Design, develop, maintain, and optimize robust and scalable ETL/ELT pipelines for structured, semi-structured, and unstructured data from multiple sources, including databases, REST APIs, files, documents, and external systems.
- Develop reliable and resilient pipelines by applying leading data engineering models, including incremental processing, idempotency, deduplication, backfill, schema change management, and disaster recovery.
- Orchestrate data pipelines using Dagster and dlt by applying principles of modularity, reusability, dependency management, and observability.
- Integrate data from various sources into centralized platforms such as data lakes, lakehouse architectures, and data warehouses.
- Contribute to the design and evolution of modern data architectures in cloud environments, mainly AWS and Azure.
- Design and maintain data models that meet analytical, operational, business intelligence, and artificial intelligence needs.
- Collaborate with data scientists, artificial intelligence specialists, and architecture consultants to develop the data components and pipelines required for analytics and artificial intelligence solutions.
- Ensure the availability, quality, traceability and reproducibility of data used by analytics and artificial intelligence solutions.
- Optimize the performance of pipelines and data platforms by considering volumes, access patterns, compute resources, and costs.
- Implement quality, testing and observability mechanisms to detect anomalies, pipeline breaks, schema changes and data freshness issues.
- Apply software development best practices to data engineering, including Git, code reviews, automated testing, dependency management, and CI/CD practices.
- Participate in the operation of data solutions in production, the diagnosis of incidents and the implementation of patches to improve their reliability.
- Contribute to the technical design of end-to-end data solutions, from source integration to production.
- Evaluate technology options and participate in design choices based on project needs and constraints.
- Create reusable components, APIs, connectors, and automation scripts to reduce manual intervention and standardize practices.
- Ensure the application of good data security and confidentiality practices, including the management of access, secrets and sensitive data.
- Document pipelines, data models, architectures, and technical decisions.
- Participate in technical reviews, knowledge sharing, and continuous improvement of data engineering practices.
- Work closely with data scientists, artificial intelligence specialists, analysts, developers, and business teams to transform their needs into robust data solutions.
Profile
- Bachelor's degree in computer science, software engineering, or a related field, or an equivalent combination of education and work experience.
- 3 to 5 years of relevant experience in data engineering or a related role, acquired in several projects or technological environments.
- Proficient in Python applied to data engineering, including pipeline development, API integration, automation, and testing.
- Possess an excellent command of SQL as well as a good understanding of relational and analytical databases.
- Design, develop, deploy, and operate ETL/ELT pipelines in a production environment.
- Understand key data engineering models, including incremental processing, idempotency, deduplication, backfill, disaster recovery, and schema evolution.
- Use Dagster or a comparable orchestrator for data pipeline development and orchestration.
- Leverage dlt or a comparable tool for developing data ingestion pipelines.
- Integrate data from a variety of sources, including REST APIs, databases, structured and semi-structured files, and external systems.
- Apply best practices for quality, monitoring, and observability of data pipelines in production.
- Implement software development practices appropriate for data solutions, including Git, automated testing, dependency management, continuous delivery, and continuous deployment.
- Know the principles of data modeling.
- Collaborate with more senior profiles or architectural consultants while contributing independently to the realization of data solutions.
- Use Linux and containerized environments, including Docker.
- Understand the lifecycle of artificial intelligence and machine learning solutions and the data needs associated with developing, training and operating models.
- Demonstrate hands‑on experience with at least one major cloud platform, ideally AWS or Azure.
- Experience with a second cloud platform such as AWS, Azure or GCP (an asset).
- Implement framework as code using Terraform, Ansible, or equivalent tools (an asset).
- Use dbt or other modern data transformation and modeling tools (an asset).
- Leverage Spark, PySpark, Databricks, or other distributed processing technologies (an asset).
- Deploy and administer Kubernetes or equivalent containerized workload orchestration platform (an asset).
- Use MLflow, DVC, or other tools related to model lifecycle and MLOps practices (an asset).
- Design data architectures for generative AI solutions, including document ingestion, metadata, embeddings, vector bases, and RAG pipelines (an asset).
- Knowledge of modern data storage formats and technologies such as Parquet, Delta Lake or lakehouse architectures (an asset).
- Leverage data dissemination or event integration systems such as Kafka, Event Hubs, Kinesis or equivalent solutions (an asset).
- Participate in the modernization or migration of data platforms to modern cloud architectures (an asset).
- Hold an AWS, Azure or GCP certification (an asset).
- Demonstrate autonomy in technical implementation and know how to request the required expertise when necessary.
- Analyze, diagnose and solve complex problems efficiently.
- Rapidly develop new technical skills through strong learning ability and sustained curiosity.
- Translate business or technical needs into concrete, sustainable and maintainable solutions.
- Communicate effectively with technical and non-technical audiences.
- Adapt quickly to different projects, client contexts and technological environments.
- Foster collaboration, customer centricity and quality of delivered solutions.
- To evolve effectively in contexts with a high level of uncertainty.
Skills
- AI
- Analytics
- Ansible
- API
- Automation
- AWS
- Azure
- CI/CD
- Cloud
- Dagster
- Data Engineering
- Data Ingestion
- Data Modeling
- Data Pipelines
- Databricks
- dbt
- Delta Lake
- Docker
- ELT
- Embeddings
- ETL
- GCP
- Generative AI
- Git
- Kafka
- Kinesis
- Kubernetes
- Lakehouse
- Linux
- Machine Learning
- MLflow
- MLOps
- Observability
- Parquet
- PySpark
- Python
- REST
- Spark
- SQL
- Terraform
- Test Automation