Point your AI agent at freehire and let it find you a job.

Get the CLI →

V2 Solutions

NewBe an early applicant

Data Platform Engineer

Posted Updated
Discussion

Summary

A junior data platform engineer role supporting and maintaining big data environments — Apache Spark clusters, Airflow, Hive/Hadoop, and JupyterHub — while writing Python automation and ETL scripts and troubleshooting Spark jobs under senior engineers' guidance.

Job Scope

We are looking for an enthusiastic Junior Data Platform Engineer to support and manage our Apache Spark, Apache Airflow, and JupyterHub environments. This role is ideal for someone with a strong foundation in Python and Linux, who is eager to build a career in big data engineering and data platform administration. You will work closely with senior engineers to ensure smooth operation, deployment, and optimization of our data processing ecosystem.

Total /Relevant Experience

1+ years experience

Key Responsibilities

  • Assist in the setup, monitoring, and maintenance of Apache Spark clusters, Apache Hive, Hadoop and Airflow environments.
  • Support the development and scheduling of data pipelines using Airflow DAGs and Python scripts.
  • Help manage and configure JupyterHub for multi-user access and integration with Spark.
  • Monitor cluster health and performance under guidance and assist in troubleshooting Spark job failures.
  • Write and maintain Python automation scripts for data workflows, ETL, and process automation.
  • Participate in code reviews, documentation, and deployment activities.
  • Learn and follow best practices for distributed data processing, CI/CD, and DevOps workflows.
  • Collaborate with senior engineers and data scientists to implement improvements and new features.

Must-have skill

  • Basic understanding of Apache Spark, Apache Hive and Hadoop File System.
  • Familiarity with Apache Airflow (understanding of DAGs, scheduling, and task dependencies).
  • Hands-on experience with Python scripting (data processing, automation, or API interaction).
  • Comfortable working in Linux and container environments (command line, system logs, process management).
  • Good understanding of data processing concepts, including ETL and distributed computing.
  • Basic knowledge of Git and version control.

Good-to-Have Skills

  • Exposure to Jupyter / JupyterHub for collaborative notebook environments.
  • Knowledge of Docker or Kubernetes.
  • Knowledge of Hadoop and Apache Spark Cluster.
  • Familiarity with SQL and working with structured/unstructured data.
  • Experience with cloud platforms (AWS, GCP, or Azure) is a plus.
  • Interest in big data (Hadoop), DevOps, and data pipeline automation.

Qualifications Criteria

  • Bachelor's or any relevant Degree.

Certification Criteria

  • NA

Disclaimer

This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

Skills

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available