Summary
Builds and maintains global data pipelines and lakes using Python, Spark, and Azure tools to ensure reliable, high-quality data for analytics and operations.
MUNKAVÉGZÉS HELYE
7622 Pécs, Bajcsy-Zsilinszky utca 33.
TEVÉKENYSÉGI TERÜLET
IT
MUNKAVÉGZÉS KEZDETE
Megegyezés szerint
FOGLALKOZTATÁS MÉRTÉKE
Teljes munkaidő
My responsibilities:
Maintain, develop, and optimize the global datalake environment, ensuring high data quality, reliability, and overall performance
Design and implement high-performance data processing pipelines and workflows utilizing Python, Spark, and Databricks
Manage scalable and efficient data integration and processing solutions using Azure Data Factory, Databricks, and Event Hub
Establish robust data management processes and execute advanced data querying and quality control with SQL and big data technologies
Automate operational tasks and support infrastructure maintenance by leveraging shell scripting and foundational UNIX knowledge
Support data modeling initiatives for coherent structure design and document all technical processes to ensure operational continuity
The knowledge I own:
Strong background and understanding of database data management (general SQL knowledge, data querying, data quality control)
Strong programming knowledge in Python (pandas, numpy)
Advanced level Big Data / data lake framework (Spark, Databricks)
Strong cloud knowledge and handling of Azure Data Factory, Databricks, Event Hub
Basic familiarity with UNIX operating system, especially shell scripting
Basic understanding of network level problems and connectivity requirements
Basic Understanding data modelling principes
Data-centric mindset
Structured, analytical thinking
Strong communication skills
Well-organized
Can work in a multi-shift operation (Monday to Friday, 07:00 AM - 22:00 PM) + on-call (weeknights, weekends and public holidays)
The offer that would convince me:</