Data Engineer

Summary

Build and maintain data pipelines using Python, PostgreSQL, Apache Airflow, and Apache Flink for a growing data team in Mumbai.

Job Title: Associate Data Engineer / Junior Data Engineer (Fresher)

Job Type: Full-time
Experience Level: Fresher (0-1 Years)

About the Role:

We are looking for an enthusiastic and detail-oriented Associate Data Engineer to join our growing data team. If you are a recent graduate passionate about data and eager to build a career in data engineering, this is the perfect opportunity for you!

In this role, you will work closely with senior engineers to build, maintain, and optimize data pipelines. You will get hands-on experience with modern data infrastructure, transitioning your basic academic or project-level knowledge into production-grade skills. We focus heavily on mentorship, ensuring you have the support you need to grow as a professional.

Key Responsibilities:

  • Pipeline Development: Assist in writing and maintaining ETL (Extract, Transform, Load) scripts to move data across various systems.
  • Database Management: Write basic SQL queries, create tables, and interact with PostgreSQL databases to support data storage and retrieval.
  • Data Orchestration: Help schedule, monitor, and troubleshoot automated data workflows using Apache Airflow.
  • Stream Processing: Collaborate with the team to build and monitor real-time or near-real-time data streaming applications using Apache Flink.
  • Coding & Debugging: Write clean, readable, and well-documented code in Python.
  • Continuous Learning: Actively participate in code reviews, learn best practices from senior team members, and stay updated on the latest data engineering trends.

Must-Have Qualifications:

  • Education: Bachelor’s or Master’s degree in Computer Science, IT, Data Science, Mathematics, or a related field (Recent graduates from batch [Year] are welcome).
  • Python: Solid foundational knowledge of Python programming (data structures, loops, functions, basic libraries like Pandas).
  • PostgreSQL (SQL): Basic understanding of relational databases, table design, and the ability to write standard SQL queries (SELECT, JOINs, GROUP BY).
  • Apache Airflow: Theoretical understanding or project-level exposure to how Airflow works (DAGs, tasks, scheduling).
  • Apache Flink: Basic conceptual knowledge of distributed data processing or stream processing using Flink.
  • Soft Skills: Strong analytical thinking, problem-solving skills, and a genuine eagerness to learn and adapt to new technologies.

Good-to-Have (Bonus Points):

  • Spatial Data: Knowledge of the PostGIS extension for PostgreSQL and a basic understanding of how to write geo/spatial queries.
  • Version Control: Basic understanding of tools like Git/GitHub.
  • Cloud Basics: Familiarity with cloud platforms (AWS, GCP, or Azure).
  • OS: Comfortable with basic Linux/Unix command-line operations.
  • Portfolio: Any personal or academic projects showcasing data pipelines or data processing.