Data Engineering Summer Intern (May-Aug '27)

Description

We are looking for a motivated Student Engineer to join our Enterprise Systems and Data Engineering team. You will work with experienced engineers to build, enhance, test, and support modern data solutions using Databricks and Microsoft Azure. The internship offers practical exposure to enterprise-scale data pipelines, data quality, cloud storage, governance, and production engineering practices.

This summer internship period is from May/June to August/September 2027.

Responsibilities

What you will work on -

  • Develop and maintain data ingestion and transformation pipelines using Python, SQL, PySpark, and Databricks notebooks.
  • Use Databricks Workflows and jobs to orchestrate, schedule, monitor, and troubleshoot data processing activities.
  • Build and test ETL/ELT solutions for structured and semi-structured data from enterprise systems, databases, files, and APIs.
  • Work with Delta Lake and lakehouse concepts, including Bronze, Silver, and Gold data layers.
  • Apply data profiling, validation, reconciliation, and quality checks to improve data reliability.
  • Assist with onboarding new datasets into Azure Data Lake Storage and Databricks.
  • Support pipeline monitoring, root-cause analysis, defect resolution, and documentation.
  • Use Git and CI/CD practices for version control, peer review, testing, and controlled deployments.
  • Collaborate with data engineers, analysts, platform teams, and business stakeholders to understand requirements and deliver usable data products.

Databricks Learning Focus

  • Hands-on exposure may include: Databricks workspace and notebooks, Apache Spark and PySpark, Delta Lake, Databricks Workflows, SQL Warehouses, Unity Catalog fundamentals, data quality controls, performance basics, and lakehouse architecture.

What You Could Learn

  • Databricks & Spark - Develop notebooks and scalable transformations with SQL, Python, and PySpark.
  • Pipeline Engineering - Understand ingestion, orchestration, testing, monitoring, and operational support.
  • Cloud Data Platforms - Work with Azure-based storage, integration, and data processing patterns.
  • Data Quality & Governance - Apply validation, documentation, access control, lineage, and reliability practices.
  • Engineering Delivery - Gain experience with Git, code reviews, CI/CD, Agile delivery, and stakeholder collaboration.

Required Experience and Skills

  • Currently pursuing or recently completed a Bachelor’s or Master’s degree in Computer Science, Information Technology, Data Science, Electronics Engineering, or a related technical discipline.
  • Ability to work on-site for 3 month internship period
  • Basic programming proficiency in Python and the ability to write clear, testable code.
  • Working knowledge of SQL, relational databases, joins, aggregations, and data manipulation.
  • Understanding of data structures, algorithms, and software engineering fundamentals.
  • Strong analytical and problem-solving skills with attention to detail.
  • Clear written and verbal communication skills, with the ability to collaborate in a team environment.
  • Curiosity, accountability, and willingness to learn new technologies.

Desired Experience and Skills

  • Academic, internship, or personal project experience with Databricks, Apache Spark, or PySpark.
  • Exposure to Microsoft Azure, Azure Data Factory, Azure Data Lake Storage, or Azure SQL.
  • Understanding of ETL/ELT, data warehousing, lakehouse, or medallion architecture concepts.
  • Familiarity with Git, Azure DevOps, CI/CD, Linux, REST APIs, or Power BI.
  • Awareness of data governance, security, access control, or data quality principles.
  • Databricks or Microsoft Azure learning credentials are an advantage but not required.

Ideal Candidate Profile

  • Takes ownership of assigned work and communicates progress or blockers early.
  • Approaches problems methodically and validates results before considering work complete.
  • Can learn independently while seeking guidance at the right time.
  • Values clean code, documentation, data security, and reliable delivery.
  • Is interested in building a long-term career in data engineering and cloud data platforms.

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available