freehire launches on Product Hunt on 26 August.

Follow →

Data Engineer

Summary

Build and maintain scalable cloud data infrastructure, design ETL pipelines with Airflow/Spark, and collaborate with cross-functional teams to ensure data reliability and timeliness.

What you will be working on:

  • Designing, building, and maintaining highly available and scalable data infrastructure on the Cloud
  • Designing, developing, and maintaining reliable data ingestion and transformation pipelines using modern workflow orchestrator
  • Designing, developing, and maintaining robust code repository and reliable CI/CD pipelines
  • Designing, developing, and maintaining software packages that abstract out the repetitive procedures for extracting, loading, and transforming raw data into valuable cleaned data
  • Collaborating with Data Analysts, Data Scientists, ML Engineers, Software Engineers, Product Managers, and Business Stakeholders of varying functions to ensure the veracity and the timeliness of the data processed
  • Publishing high-quality technical documentation in company's internal Wiki and conducting regular knowledge sharing, hands‑on training, and mentoring for the less experienced team members
  • Being the VP's and Manager's technical counsellor/researcher in designing a future-ready data infrastructure

Who you are:

  • You have 5+ years of Data Engineering experience, or equivalent experience of handling 500+ daily data pipelines in Terabytes of scale
  • You enjoy working at the intersection of Data Engineering and DevOps
  • You have hands‑on experience in managing Kubernetes cluster on the Cloud
  • You are adept at setting up data pipeline CI/CD along with its tests using tools such as Jenkins, Gitlab CI, or equivalent
  • You are fluent in Python, SQL, and Bash scripting for building data infrastructure and for developing data ingestion and transformation pipelines — fluency in Java/Go/Rust is a plus
  • You know how to work with modern workflow orchestrators — Airflow is a must
  • You are skillful at developing abstraction layer on top of workflow orchestrator to enable data transformation at scale
  • You are familiar with various flavour of RDBMS and Columnar Query Engines — familiarity with NoSQL is a plus
  • You know how to work with Apache Spark for both batch and stream data processing — familiarity with Apache Flink or Beam is a plus
  • You are naturally inclined into optimising complex systems, both performance-wise and cost-wise

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available