freehire launches on Product Hunt on 26 August.

Follow →

Senior Data Engineer

Summary

Senior Data Engineer builds and migrates data pipelines for an investment banking client, rewriting 40% of the code to move from Cloudera to a new data center using Python, SQL, and Hadoop technologies.

Project description

We've been engaged by an Investment Banking client in the FM potfolio . PTS app is an inhouse build system which is silver source of all settlement information originated from GPTM. PTS app is in CDP (Cloudera data platform) platform which bag data ecosystem. It will be migrated to a different datacentre and CDP will be decommissioned . Hence PTS application, we need to rewrite code approx. 40% to rebuild from scratch like develop data ingestion framework, extraction frame work and migration of existing components to new data centre from west to East.

Responsibilities

  • Create and maintain optimal data pipeline architecture, assemble large, complex data sets that meet functional / non-functional business requirements.
  • Identify, design, and implement internal process improvements: automating manual processes, optimizing data delivery, re-designing infrastructure for greater scalability, etc.
  • Build the infrastructure required for optimal extraction, transformation, and loading of data from a wide variety of data sources using SQL and Hadoop technologies.
  • Build analytics tools that utilize the data pipeline to provide actionable insights into customer acquisition, operational efficiency and other key business performance metrics.
  • Work with project team including the Executive, Product, Data and Design teams to assist with data-related technical issues and support their data infrastructure needs.
  • Create data tools for analytics and data scientist team members that assist them in building and optimizing our product into an innovative industry leader.
  • Work with data and analytics experts to strive for greater functionality in our data systems.

SKILLS

Must have

  • Bachelor’s or Master’s degree in Computer Science, Data Engineering, or a related technical field
  • 8+ years of experience in data engineering (or equivalent hands on experience with large scale data platforms)
  • Strong programming skills in Python for data processing, automation, and tooling
  • Hands on experience with Minio, Iceberg and DuckDB for large scale data processing
  • Advanced SQL skills, including complex queries, performance tuning, and working with large relational datasets
  • Experience building and maintaining ETL/ELT pipelines and working with structured and semi structured data
  • Understanding of data modelling concepts (e.g. relational, dimensional, and/or big data modelling)
  • Strong understanding of DevOps/CI/CD concepts and tooling (e.g. Docker, Kubernetes, Git)

Nice to have

• Experience of Unix/Linux environments is plus • Experience of Agile/Scrum development methodologies is a plus • Be nice, respectful, able to work in a team • Willingness to learn

See also