Senior Data Engineer - Data Ingestion, Enrichment
Responsibilities
- Design, build, and own Preply’s data lake.
- Develop and operate scalable, reliable batch and streaming ingestion pipelines.
- Define and implement data contracts between producers and consumers.
- Build enrichment logic that joins, standardizes, and contextualizes data across domains.
- Instrument ingestion pipelines with strong observability: freshness, latency, data quality, and cost metrics.
- Apply consistent access control, classification, and privacy protections at ingestion time.
- Contribute to standardized ingestion templates, shared libraries, and platform tooling that enable teams to onboard new data sources independently.
- Work closely with Product, Backend, Analytics, and ML partners to align on ingestion requirements.
Requirements
- Exposure to and experience building architectural patterns of a large, high-scale application (e.g., well-designed APIs, high-volume data pipelines, efficient algorithms).
- Solid experience working in platform or data engineering teams (or equivalent impact) with evidence of leading multi-stakeholder deliveries.
- Familiarity with cloud platforms (AWS/GCP or equivalent) and modern DevOps practices.
- Hands‑on experience designing and implementing real‑time and batch data processing infrastructures using modern frameworks like Spark, Flink, Spark streaming, Kafka, Debezium, etc.
- Expertise with orchestration tools such as Airflow, dbt, or similar.
- Exceptional problem‑solving skills paired with a proactive, innovative mindset focused on continuous improvement.
- Strong communication and cross‑functional collaboration skills (English level B2+)
Core Competencies
Demonstrates expertise in designing and implementing scalable data ingestion pipelines and architectures, with a strong focus on data quality, observability, and cross‑functional collaboration. Proficient in utilizing cloud platforms and orchestration tools to enhance data processing capabilities.
Highest-signal resume keywords
- Data Lake Design
- Batch And Streaming Ingestion Pipelines
- Real-Time Data Processing
- Cloud Platforms (AWS/GCP)
- Orchestration Tools (Airflow, dbt)
ATS Optimization Keywords
Hard Skills
- Data Ingestion
- Data Quality Metrics
- Data Contracts
- Data Standardization
- Architectural Patterns
- High-Volume Data Pipelines
- Efficient Algorithms
- Real-Time Processing
- Batch Processing
- Data Engineering
Soft Skills
- Problem-Solving
- Proactive Mindset
- Innovative Thinking
- Communication
- Cross-Functional Collaboration
Industry Keywords
- Data Engineering
- DevOps Practices
- Data Privacy
- Access Control
- Data Classification
Tools & Technologies
- Spark
- Flink
- Kafka
- Debezium
- Airflow
- Dbt