Point your AI agent at freehire and let it find you a job.

Get the CLI →

Data Engineering Specialist

Discussion

Summary

The Data Engineering Specialist designs, develops, and maintains scalable ETL/ELT pipelines using Python, Dagster, and dlt. The role involves integrating data into cloud-based lakehouse architectures on AWS or Azure to support analytics and AI initiatives.

Responsibilities

  • Design, develop, maintain, and optimize robust and scalable ETL/ELT pipelines for structured, semi-structured, and unstructured data from multiple sources, including databases, REST APIs, files, documents, and external systems.
  • Develop reliable and resilient pipelines by applying leading data engineering models, including incremental processing, idempotency, deduplication, backfill, schema change management, and disaster recovery.
  • Orchestrate data pipelines using Dagster and dlt by applying principles of modularity, reusability, dependency management, and observability.
  • Integrate data from various sources into centralized platforms such as data lakes, lakehouse architectures, and data warehouses.
  • Contribute to the design and evolution of modern data architectures in cloud environments, mainly AWS and Azure.
  • Design and maintain data models that meet analytical, operational, business intelligence, and artificial intelligence needs.
  • Collaborate with data scientists, artificial intelligence specialists, and architecture consultants to develop the data components and pipelines required for analytics and artificial intelligence solutions.
  • Ensure the availability, quality, traceability and reproducibility of data used by analytics and artificial intelligence solutions.
  • Optimize the performance of pipelines and data platforms by considering volumes, access patterns, compute resources, and costs.
  • Implement quality, testing and observability mechanisms to detect anomalies, pipeline breaks, schema changes and data freshness issues.
  • Apply software development best practices to data engineering, including Git, code reviews, automated testing, dependency management, and CI/CD practices.
  • Participate in the operation of data solutions in production, the diagnosis of incidents and the implementation of patches to improve their reliability.
  • Contribute to the technical design of end-to-end data solutions, from source integration to production.
  • Evaluate technology options and participate in design choices based on project needs and constraints.
  • Create reusable components, APIs, connectors, and automation scripts to reduce manual intervention and standardize practices.
  • Ensure the application of good data security and confidentiality practices, including the management of access, secrets and sensitive data.
  • Document pipelines, data models, architectures, and technical decisions.
  • Participate in technical reviews, knowledge sharing, and continuous improvement of data engineering practices.
  • Work closely with data scientists, artificial intelligence specialists, analysts, developers, and business teams to transform their needs into robust data solutions.

Profile

  • Bachelor's degree in computer science, software engineering, or a related field, or an equivalent combination of education and work experience.
  • 3 to 5 years of relevant experience in data engineering or a related role, acquired in several projects or technological environments.
  • Proficient in Python applied to data engineering, including pipeline development, API integration, automation, and testing.
  • Possess an excellent command of SQL as well as a good understanding of relational and analytical databases.
  • Design, develop, deploy, and operate ETL/ELT pipelines in a production environment.
  • Understand key data engineering models, including incremental processing, idempotency, deduplication, backfill, disaster recovery, and schema evolution.
  • Use Dagster or a comparable orchestrator for data pipeline development and orchestration.
  • Leverage dlt or a comparable tool for developing data ingestion pipelines.
  • Integrate data from a variety of sources, including REST APIs, databases, structured and semi-structured files, and external systems.
  • Apply best practices for quality, monitoring, and observability of data pipelines in production.
  • Implement software development practices appropriate for data solutions, including Git, automated testing, dependency management, continuous delivery, and continuous deployment.
  • Know the principles of data modeling.
  • Collaborate with more senior profiles or architectural consultants while contributing independently to the realization of data solutions.
  • Use Linux and containerized environments, including Docker.
  • Understand the lifecycle of artificial intelligence and machine learning solutions and the data needs associated with developing, training and operating models.
  • Demonstrate hands‑on experience with at least one major cloud platform, ideally AWS or Azure.
  • Experience with a second cloud platform such as AWS, Azure or GCP (an asset).
  • Implement framework as code using Terraform, Ansible, or equivalent tools (an asset).
  • Use dbt or other modern data transformation and modeling tools (an asset).
  • Leverage Spark, PySpark, Databricks, or other distributed processing technologies (an asset).
  • Deploy and administer Kubernetes or equivalent containerized workload orchestration platform (an asset).
  • Use MLflow, DVC, or other tools related to model lifecycle and MLOps practices (an asset).
  • Design data architectures for generative AI solutions, including document ingestion, metadata, embeddings, vector bases, and RAG pipelines (an asset).
  • Knowledge of modern data storage formats and technologies such as Parquet, Delta Lake or lakehouse architectures (an asset).
  • Leverage data dissemination or event integration systems such as Kafka, Event Hubs, Kinesis or equivalent solutions (an asset).
  • Participate in the modernization or migration of data platforms to modern cloud architectures (an asset).
  • Hold an AWS, Azure or GCP certification (an asset).
  • Demonstrate autonomy in technical implementation and know how to request the required expertise when necessary.
  • Analyze, diagnose and solve complex problems efficiently.
  • Rapidly develop new technical skills through strong learning ability and sustained curiosity.
  • Translate business or technical needs into concrete, sustainable and maintainable solutions.
  • Communicate effectively with technical and non-technical audiences.
  • Adapt quickly to different projects, client contexts and technological environments.
  • Foster collaboration, customer centricity and quality of delivered solutions.
  • To evolve effectively in contexts with a high level of uncertainty.

Skills

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available