freehire launches on Product Hunt on 26 August.

Follow →

Data Engineer (Azure)

Summary

Build and maintain Azure-based data pipelines using Python/PySpark and Databricks, automating ETL processes and integrating data sources for a global background-screening platform.

We support recruitment for a US-based company that is a provider of mission-critical background screening solutions. They work with Fortune 100 clients helping them manage risk and hire the best talent. This role will provide you with an outstanding opportunity to work for an industry-leading company. With over 4500 employees from 30+ different nationalities, you will be working with a diverse bunch of creatives redefining the world of digital background check and verification services across the globe.

We’re seeking a self-motivated Data Engineer with strong Python/PySpark skills to join the Data Engineering Team and help build the Azure Data Analytics Platform. The ideal candidate is an independent, collaborative team player who leads projects, identifies process gaps, and continuously develops expertise in Human Capital technology.

The role involves developing reusable, metadata-driven data pipelines, automating platform processes, building data integrations, extending ETL libraries, writing unit tests, creating Databricks monitoring solutions, proactively resolving ETL issues, collaborating on cloud resources, updating documentation, conducting code reviews, and enhancing platform architecture.

Key takeaways:

Stack: Python/PySpark, SQL, Databricks Spark, Knowledge of Azure cloud native solutions

Salary: 130 - 150 PLN net/h on B2B

Working model: 100% Remote

Location in Poland: Krakow

Recruitment process:

  • A call with a Motife recruiter (30 min)

  • An online interview with a technical case (1.5h)

Responsibilities:

  • Build reusable, metadata-driven data pipelines.

  • Automate and optimize data platform processes.

  • Develop integrations with data sources and consumers.

  • Extend shared ETL libraries with transformation methods.

  • Write unit tests.

  • Create monitoring solutions for the Databricks platform.

  • Proactively address ETL performance and quality issues.

  • Collaborate with infrastructure teams on cloud resources.

  • Update data platform wiki and documentation.

  • Conduct code reviews to ensure quality.

  • Initiate and implement architecture improvements.

Requirements:

  • Strong experience with Python/PySpark and SQL.

  • Hands-on experience building robust data pipelines using Databricks Spark.

  • Experience processing large-scale datasets in production environments.

  • Strong knowledge of Kafka or other streaming platforms such as Azure Event Hubs or Amazon Kinesis.

  • Experience building streaming data integrations and event-driven data pipelines.

  • Good understanding of Spark Structured Streaming.

  • Familiarity with CDC (Change Data Capture) concepts; experience with Debezium is a plus.

  • Strong knowledge of Databricks Delta optimization (partitioning, Z-ordering, compaction, etc.).

  • Experience developing reusable Python libraries and packages.

  • Hands-on experience with CI/CD pipelines.

  • Good understanding of networking fundamentals.

  • Familiarity with Agile/Scrum methodologies.

What we offer:

  • 100% Remote work model.

  • Superior co-working and personal development experience in an international setting.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available