Data Engineering Co-op: SQL, Python & Databricks
Responsibilities
- Support the development and maintenance of data pipelines using SQL, Python, and Spark.
- Assist with Databricks Unity Catalog migration activities, including data validation and basic permission checks.
- Help analyze Databricks usage, cost, and job performance using metrics and AI‑driven recommendations.
- Perform data access cleanup by reviewing and updating user and group permissions.
- Support data quality checks and assist in resolving data issues with guidance from senior engineers.
- Document data assets, data processes, and standard operating procedures.
- Assist the team with ad‑hoc analysis and operational support tasks as required.
Requirements
- Undergraduate / Graduate students with a background in data engineering, computer science, information systems, analytics, or related disciplines.
- Strong SQL skills and basic Python knowledge for querying, transforming, and validating data.
- Basic understanding of data engineering concepts such as ETL / ELT and batch data processing.
- Familiarity with Databricks and Spark, including notebooks, jobs, and tables.
- Understanding of basic data access and governance concepts such as users, groups, and permissions.
- Strong analytical mindset to review job performance, cost metrics, and identify simple improvement opportunities.
- Good communication, documentation, and teamwork skills, with willingness to learn and receive feedback.
- Able to commit to the internship on a full‑time basis, from Aug 2026 – January 2027.
Core Competencies
Proficient in SQL and basic Python for data querying and transformation, with a foundational understanding of data engineering concepts such as ETL/ELT. Capable of supporting data quality checks, documentation, and operational tasks while demonstrating strong analytical and communication skills.