Big Data Platform Engineer – L3
Summary
Maintain and optimize enterprise big-data platforms (Hadoop, Kafka, OpenSearch, AWS EMR) and troubleshoot production issues in a 24/7 environment.
Key Responsibilities
- Provide L3 technical support for enterprise Big Data platforms and production environments.
- Administer and maintain Hadoop clusters, including HDFS, YARN, HBase and related components.
- Perform cluster lifecycle activities such as provisioning, scaling, patching and decommissioning.
- Manage and optimise Apache Kafka for high-throughput, real-time data streaming.
- Administer OpenSearch/Elasticsearch clusters and optimise indexing and query performance.
- Support AWS EMR environments for scalable data processing and reconciliation workloads.
- Monitor and tune MapReduce, YARN and Spark workloads for performance and reliability.
- Manage Kerberos authentication, access controls and security across the Hadoop ecosystem.
- Perform capacity planning, performance tuning, failover and disaster recovery activities.
- Support high-severity incidents and drive technical issue resolution within agreed SLAs.
- Develop and maintain runbooks, SOPs, technical documentation and operational best practices.
- Work with architects, development teams and project teams on technology changes and transformation initiatives.
- Review technology changes and identify potential operational and technical risks.
- Ensure new solutions meet production readiness and operational requirements.
- Coach technical team members and partner resources and promote knowledge sharing.
- Identify opportunities for service improvement, automation and operational efficiency.
Key Requirements
- 11–14 years of experience in Big Data, Data Platform Engineering or Infrastructure Engineering.
- Strong hands‑on experience in the Hadoop ecosystem, including HDFS, YARN, Spark, MapReduce and HBase.
- Strong experience in Apache Kafka administration and Kafka internals.
- Experience managing OpenSearch / Elasticsearch clusters.
- Hands‑on experience with AWS EMR and good knowledge of AWS Cloud services.
- Strong Linux system administration and scripting skills using Shell, Python or similar languages.
- Experience with Kerberos, access control, data security and governance.
- Experience supporting high-volume and low-latency production environments.
- Knowledge of Hadoop components such as Storm and other ecosystem technologies.
- Good understanding of JVM and virtual machine environments.
- Knowledge of SQL, Hive or other SQL-on-Hadoop technologies.
- Experience with ETL processes or ETL software is advantageous.
- Strong troubleshooting, analytical and problem‑solving skills.
- Good communication and stakeholder management skills.
- Ability to work under pressure and participate in on‑call support.
Good to Have
- AWS Cloud certification.
- Knowledge of PAM and Kerberos-based access control.
- Experience with hardware configuration, rack setup, disk topology and RAID.
- Knowledge of virtual machine deployment and configuration.
- Proficiency in Python, Java or Scala.
- Experience with ETL processes/software.
- Experience in banking, payments or other high-volume transaction environments.
EA Number: 11C4879