freehire launches on Product Hunt on 26 August.

Follow →

IT Storage Engineer

Open 25d

Summary

Designs, deploys, and maintains HDFS clusters for enterprise data applications, ensuring high availability, security, and performance while troubleshooting issues and automating operations.

Key Responsibilities:

  1. 1. Design, deploy, and maintain HDFS clusters to support enterprise-scale data applications.
  2. Monitor cluster performance, manage storage capacity, and ensure high availability and fault tolerance.
  3. Implement data security, access controls, and encryption for HDFS data.
  4. Troubleshoot and resolve issues related to HDFS, including data node failures, replication issues, and performance bottlenecks.
  5. Manage data ingestion pipelines and optimize data storage formats (e.g. Parquet, Avro).
  6. Support and work with data engineering and analytics teams to ensure reliable data delivery and transformation workflows.
  7. Automate cluster operations using scripting (e.g., Bash, Python) and orchestration tools.
  8. Conduct upgrades and patching of Hadoop ecosystem components (HDFS, YARN, Hive, etc.).
  9. Maintain documentation of architecture, configurations, and best practices.
  10. Ensure compliance with data governance and data privacy policies.

Qualifications and Requirement:

  • a) Bachelor’s degree in Computer Science, Information Systems, or a related field.
  • 3–5+ years of experience with Hadoop ecosystem, particularly HDFS administration.
  • Strong understanding of HDFS architecture, replication, and fault tolerance.
  • Experience with Cloudera, Hortonworks, or Apache Hadoop distributions.
  • Proficiency in Linux/Unix system administration and scripting (Bash, Python, etc.).
  • Familiarity with related components: YARN, Hive, HBase, Spark, Oozie, and Zookeeper.
  • Experience with monitoring tools like Ambari, Cloudera Manager, or Nagios.

Advantage to have: -

  • Hadoop certification (e.g., Cloudera Certified Administrator for Apache Hadoop - CCAH).
  • Knowledge of cloud-based big data platforms (AWS EMR, Azure HDInsight, GCP Dataproc).
  • Experience with containerization (Docker/Kubernetes) for big data workloads.
  • Exposure to data lake architectures and data governance tools.

    See also

    Tailor your CV for this role?

    We couldn't check your fit for this role — add a CV to your profile to see it next time.

    A new version of freehire is available