freehire launches on Product Hunt on 26 August.

Follow →

data engineer for an educational project

Summary

Build and maintain Preply’s scalable data lake and real-time ingestion pipelines, ensuring high-quality, governed data assets for analytics and ML across 180+ countries.

Описание

Preply creates personalized language-learning experiences by connecting learners with tutors. Its human-led, technology-enabled platform serves learners in 180 countries, with over 100,000 tutors teaching more than 90 languages. The Data Ingestion and Enrichment team provides a trusted, scalable data foundation for analytics, machine learning, and product features through governed, production-grade data assets.

Задачи

  • Build and own Preply’s data lake and trusted ingestion and enrichment foundations
  • Develop and operate scalable batch and streaming ingestion pipelines for real-time and analytical use cases
  • Design raw, standardized, and consumption data layers with clear responsibilities, lineage, and retention strategies
  • Define and implement data contracts covering schemas, freshness, volume, and quality guarantees
  • Embed validation, anomaly detection, and quality checks early in the ingestion lifecycle
  • Standardize how quality metrics are measured, monitored, and surfaced
  • Build enrichment logic that joins, standardizes, and contextualizes data across domains
  • Support historical tracking, point-in-time correctness, and dataset versioning
  • Instrument ingestion pipelines with freshness, latency, data quality, and cost metrics
  • Contribute to SLOs, alerting, and incident response playbooks
  • Apply access control, classification, privacy protections, masking, minimization, and anonymization at ingestion time
  • Contribute to standardized ingestion templates, shared libraries, and platform tooling
  • Improve dataset discoverability, documentation, and metadata
  • Collaborate with Product, Backend, Analytics, and ML partners on ingestion requirements, trade-offs, and priorities
  • Promote shared ownership of data quality and platform standards

Требования

  • Experience building architectural patterns for large, high-scale applications, including well-designed APIs, high-volume data pipelines, or efficient algorithms
  • Solid experience in platform or data engineering teams, or equivalent impact
  • Evidence of leading multi-stakeholder deliveries
  • Familiarity with AWS, GCP, or equivalent cloud platforms
  • Familiarity with modern DevOps practices
  • Hands‑on experience designing and implementing real‑time and batch data processing infrastructure
  • Exceptional problem‑solving skills and a proactive, innovative mindset focused on continuous improvement
  • Strong communication and cross‑functional collaboration skills
  • English at B2+ level
  • Nice to have: Spark, Flink, Spark Streaming, Kafka, Debezium, Airflow, dbt, or similar tools

Условия

  • A monthly allowance for lessons on Preply.com
  • Learning & Development budget
  • Time off for self-development
  • Attractive relocation package to join the Preply Barcelona Hub
  • Competitive financial package with equity, leave allowance, and health insurance
  • Free mental health support platforms
  • Access to Gympass-partnered wellness and gym centers throughout Spain

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available