Lead Data engineer

Summary

Lead Data Engineer designing and building scalable batch and real-time data pipelines using Scala, Apache Spark, SQL, Kafka, and cloud platforms (AWS/Azure/GCP).

Job Title: Lead Data Engineer

Experience: 7–12 Years

Location: Bangalore

Employment Type: Full-Time

Work Mode: Work from Office

( Scala experience Must )


Job Summary

We are looking for an experienced Lead Data Engineer to lead the design, development, and implementation of scalable data engineering solutions. The ideal candidate should have strong hands-on expertise in Scala, Apache Spark, SQL, and cloud data platforms, along with experience in leading technical teams and delivering large-scale data pipelines.

Key Responsibilities

  • Lead the design and development of high-volume, scalable data pipelines and data processing frameworks.
  • Develop robust batch and real-time data processing solutions using Apache Spark.
  • Work extensively with Scala, Spark, and SQL to build and optimize data transformation pipelines.
  • Design and implement data solutions on cloud platforms such as AWS, Azure, or GCP.
  • Develop and maintain data pipelines for structured and unstructured data.
  • Perform performance tuning and optimization of Spark jobs, SQL queries, and data pipelines.
  • Work with technologies such as Kafka for real-time/streaming data processing.
  • Ensure data quality, reliability, security, and governance across data platforms.
  • Provide technical leadership, conduct code reviews, and mentor junior and mid-level engineers.

Mandatory Skills

  • 7–12 years of experience in Data Engineering.
  • Strong hands-on experience with Scala + Apache Spark.
  • Excellent knowledge of SQL and database concepts.
  • Strong experience building ETL/ELT data pipelines.
  • Experience with Spark Batch and Spark Streaming.
  • Experience with Kafka or other streaming technologies.
  • Hands-on experience with at least one cloud platform – AWS / Azure / GCP.
  • Strong understanding of data warehousing, data lakes, and distributed data processing.
  • Experience with performance tuning and optimization of Spark applications.
  • Good understanding of CI/CD, Git, and Agile methodologies.
  • Strong problem-solving and communication skills.

Good to Have

  • Experience with Databricks / Delta Lake.
  • Knowledge of Python/PySpark.
  • Experience with Airflow or other workflow orchestration tools.
  • Experience with Docker/Kubernetes.
  • Knowledge of modern Lakehouse architecture.
  • Experience leading a team or mentoring data engineers.


See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available