Data Engineer Scala/Spark
NewBe an early applicantAt DAC.digital, we are continuously expanding our business and strengthening our position in the market. As part of our growth strategy, we are launching a strategic partnership with one of the world’s leading providers of financial market infrastructure.
Key information:
- 27 000 – 32 000 PLN net/month – pure B2B contract
- 23 500 –28 000 PLN net/month – B2B contract (days off included)
It is vital that you have:
- strong experience with Scala and Apache Spark for designing and developing distributed data processing solutions;
- hands-on experience with Databricks and Zeppelin Notebooks in data engineering and analytics projects;
- proficiency in leveraging Google Cloud Platform (GCP) services for scalable cloud-based data solutions;
- advanced SQL knowledge for complex querying, data modeling, and performance tuning;
- experience in building, scheduling, and monitoring ETL pipelines using Apache Airflow;
- practical knowledge of Jira and Confluence for agile delivery, project tracking, and documentation management;
- knowledge of English (min. B2);
- proven leadership experience in technical teams;
- experience managing stakeholders and aligning technical solutions with business needs;
- strong communication, collaboration, and presentation skills;
- working in agile methodologies (Scrum, Kanban);
- high communication skills;
- eager to learn and share knowledge.
Nice to have:
- Oracle
- Informatica (just to analyse the current System and Workflows, you don’t need to develop any new pipelines)
- Scala
- Spark
- GCP
- SQL
- Databricks
- Airflow
- Jira
- Confluence
You will be responsible for supporting our team in:
- designing, developing, and maintaining scalable data processing solutions using Scala and Apache Spark;
- building and optimizing data pipelines and ETL workflows in cloud-based environments on Google Cloud Platform (GCP);
- developing and maintaining data engineering solutions using Databricks and Zeppelin Notebooks;
- writing, optimizing, and troubleshooting complex SQL queries to support business and analytical requirements;
- orchestrating, scheduling, and monitoring data workflows using Apache Airflow;
- ensuring data quality, reliability, and performance across data platforms and processing pipelines;
- collaborating with cross-functional teams to gather requirements and deliver data-driven solutions;
- participating in code reviews, testing, and continuous improvement of data engineering best practices;
- documenting technical solutions and project deliverables in Confluence;
- supporting agile delivery processes through effective use of Jira for task management and collaboration.
What do we offer:
- possibility to work 100% remotely or on-site at our office in Gdańsk;