Data Engineer
Summary
Build and migrate data pipelines for a large-scale lakehouse platform, moving legacy Hadoop workloads to Spark 3/Iceberg and optimizing complex SQL queries.
OdoCore is looking for an experienced Data Engineer to join a large-scale data platform modernization engagement. This is a hands‑on role focused on migrating a legacy Hadoop‑based data warehouse to a modern lakehouse architecture. You'll be working directly on production‑critical pipelines that power core business reporting and analytics, so real‑world experience with the tools below — not just theoretical familiarity — is essential.
The Engagement
You'll be part of a team modernizing a large-scale data platform, covering:
- Offloading IBM Netezza and IBM DataStage workloads onto a Spark 3 / Iceberg Lakehouse
Required Hard Skills
- Strong SQL, with real experience reading and re-optimizing complex, poorly‑written legacy queries
- Apache Spark (Spark 2 and Spark 3) — PySpark or Scala
- ETL / data warehousing fundamentals — medallion (bronze/silver/gold) architecture, dimensional modeling
- Cloudera CDP platform exposure (Impala, Ranger) — or a fast ability to ramp up on it
Experience Level
3+ years in data engineering, with at least one prior migration, ETL modernization, or lakehouse build under your belt. Mid‑to‑senior individual contributors preferred.
Responsibilities
- Migrate existing Spark 2 workloads to Spark 3, ensuring performance and stability across pipelines
- Re-platform Hive tables to Apache Iceberg, handling partitioning and schema evolution
- Rebuild Oozie workflows in Apache Airflow for improved orchestration and monitoring
- Offload legacy IBM Netezza and DataStage jobs onto the Spark 3/Iceberg Lakehouse
- Read, debug, and re-optimize complex legacy SQL queries for the new platform
Must Have
- Strong SQL skills, with hands‑on experience optimizing complex legacy queries
- Practical experience with Apache Spark (Spark 2 and Spark 3) using PySpark or Scala
- Working knowledge of Apache Hive and Apache Iceberg, including partitioning and schema evolution
- Experience with orchestration tools such as Oozie and/or Apache Airflow
- Comfortable working in Linux/Unix environments with Git version control
Nice to have
- Prior experience with IBM DataStage or IBM Netezza, including reading job designs and appliance‑based DW concepts
- Exposure to Cloudera CDP tools such as Impala and Ranger
- Experience using AI‑assisted code migration or refactoring tools
- Prior involvement in a large‑scale legacy‑to‑lakehouse migration project
- Familiarity with Python scripting for automation and pipeline tooling
What's great in the job?
- Great team of smart people, in a friendly and open culture
- No dumb managers, no stupid tools to use, no rigid working hours
- No waste of time in enterprise processes, real responsibilities and autonomy
- Expand your knowledge of various business industries
- Create content that will help our users on a daily basis
- Real responsibilities and challenges in a fast evolving company
Each employee has a chance to see the impact of his work.You can make a real contribution to the success of the company.
Several activities are often organized all over the year, such as weeklysports sessions, team building events, monthly drink, and much more
A full-time position
Attractive salary package.
Trainings
12 days / year, including
6 of your choice.
Sport Activity
Play any sport with colleagues,
the bill is covered.