Senior Data Engineer (Spark)
Summary
Senior Data Engineer builds and runs large-scale batch data pipelines with Apache Spark, tuning performance and cost as data grows. They model analytics data in a lakehouse (Apache Iceberg, Trino), handle scheduling, monitoring, and data quality checks, and work with platform and product teams on end-to-end data flows.
Description
Builds and runs large scale batch data pipelines, keeping jobs fast and affordable as data volumes grow. Works with Apache Spark to build and manage large scale data processing jobs, including performance tuning to maintain efficient and reliable pipeline execution. Supports data modelling for analytics and organises data in a lakehouse so it is usable downstream. Handles automated scheduling, monitoring, and data quality checks, while working with platform and product teams on end to end data flows.
Requirements
Requirements
- Strong hands on Apache Spark including performance tuning, not Spark usage through a managed notebook only.
- Data modelling for analytics and organising data in a lakehouse so it is usable downstream.
- Automated scheduling, monitoring, and data quality checks.
- Works with platform and product teams on end to end data flows.
- Apache Iceberg or other open table formats.
- Trino or similar query engines.
- On premises or self managed cluster experience.
Skills
As published by workable · 7 questions
Basics
First name, Last name, Email, Headline, Phone, Address, Photo, Pronouns, Education, Experience, Summary, Resume, Upload additional documents, Notice Period, Current Salary, Expected Salary
Short answers (4)
- How would you rate your Apache Spark proficiency out of 10?
- How would you rate your proficiency in Python, Java, and Scala out of 10?
- How would you rate your hands-on proficiency with Apache Kafka, Apache Airflow, Docker, and Kubernetes out of 10?
- How would you rate your English proficiency out of 10?
Pick from a list (3)
- Do you have 5+ years of experience as a Data Engineer?
- Do you have strong hands-on experience with Apache Spark?
- Do you have experience building large-scale batch data pipelines using Spark?