Big Data Engineer
ACCORD INNOVATIONS PTE. LTD. Big Data Engineer
Job Summary
We are looking for an experienced Big Data Administrator / Big Data Engineer – L3 to support and maintain enterprise Big Data infrastructure in a production environment.
The successful candidate will be responsible for the availability, reliability, performance and stability of Big Data platforms, with a strong focus on Hadoop ecosystem administration, Linux infrastructure, HDFS, YARN, Kafka and related technologies.
This is an L3 technical role requiring strong troubleshooting capabilities, hands-on experience with large Hadoop clusters, production incident management, performance tuning, capacity planning and infrastructure improvements.
The role will also involve working closely with architects, developers, project teams, service managers and other technical teams to support technology roadmaps, transformation projects and continuous service improvement.
Key Responsibilities
Big Data Platform Administration
- Build, configure and maintain enterprise Hadoop cluster infrastructure.
- Perform Hadoop installation, upgrades, patches and version migrations.
- Administer large-scale Hadoop clusters and associated ecosystem components.
- Manage HDFS, YARN and MapReduce environments.
- Perform cluster maintenance, deployment and node addition/removal activities.
- Monitor cluster health, availability, capacity and performance.
- Perform HDFS backup and restoration activities.
- Manage file systems, storage capacity and cluster resources.
- Review Hadoop logs and monitoring alerts to identify and resolve issues.
- Ensure Hadoop clusters are properly configured, secured and optimized.
- Perform capacity planning and recommend infrastructure improvements.
- Tune cluster and MapReduce workloads for optimal performance.
- Configure and monitor cluster jobs and workloads.
Security & Access Management
- Manage Hadoop users and access permissions.
- Support Linux user administration related to Hadoop environments.
- Configure and support Kerberos authentication.
- Support PAM and other access-control mechanisms.
- Monitor Hadoop cluster connectivity and security.
- Apply security patches and infrastructure hardening practices.
- Ensure security and operational procedures are followed.
Kafka & Data Platform Operations
- Perform Kafka administration and operational support.
- Monitor Kafka infrastructure and troubleshoot operational issues.
- Support availability and reliability of data-streaming services.
- Work with development and architecture teams to resolve platform issues.
- Support data-quality and data-availability requirements across technology teams.
Linux & Infrastructure Administration
- Perform Linux system administration activities supporting Big Data platforms.
- Manage system monitoring, storage capacity and performance.
- Support OS-level patches and infrastructure changes.
- Coordinate rolling OS changes with cluster administration tools.
- Automate software and application installation/configuration where applicable.
- Troubleshoot system and infrastructure issues using operating-system and application logs.
- Support network and system administration activities within large-scale data-centre environments.
Scripting & Automation
- Develop and maintain automation scripts using Shell, Python or Perl.
- Debug existing scripts and automation processes.
- Automate installation, configuration and operational activities.
- Identify opportunities to improve operational efficiency through automation.
Production Support & Incident Management
- Provide L3 technical support for Big Data production environments.
- Lead the technical investigation of complex and high-severity incidents.
- Perform root-cause analysis and drive issues through to resolution.
- Proactively identify potential risks and service-impacting issues.
- Support incident, problem and change-management activities.
- Maintain system availability and reliability in accordance with agreed SLAs.
- Review technology changes and assess potential operational risks.
- Provide technical guidance to L1/L2 teams and partner resources.
Project & Technology Transformation
- Support Big Data infrastructure projects and transformation initiatives.
- Work with architects, developers, project teams and service managers on technology roadmaps.
- Ensure new solutions are production-ready before operational handover.
- Support technical validation and readiness activities for new projects.
- Evaluate new technologies and identify opportunities for service improvement.
- Recommend improvements to Big Data infrastructure, processes and operational practices.
- Maintain technical documentation, SOPs, knowledge articles and best practices.
- Share technical knowledge and coach other team members where required.
Required Technical Skills
Critical Skills
- Linux Administration
- Hadoop Administration
- HDFS
- YARN
- MapReduce
- Hadoop ecosystem components
- Hortonworks Data Platform (HDP)
- Shell / Perl / Python scripting
- PAM
- Kerberos
- Access control and security mechanisms
Essential Skills
- Kafka Administration
- Hardware configuration and setup
- Rack and disk topology
- RAID
- Virtual Machine deployment and configuration
- JVM administration / troubleshooting
Advantageous Skills
- Elasticsearch Administration
- Python / Java / Scala
- SQL
- Hive / SQL-on-Hadoop technologies
- ETL processes and tools
- Apache Ambari
- Spark
- Hadoop 2.x / YARN-based applications
Required Experience & Qualifications
- Minimum 8 years of experience in system administration or infrastructure operations, including system monitoring, storage management, performance tuning and infrastructure development.
- At least 1–2 years of hands-on experience deploying and administering large Hadoop clusters.
- Strong hands-on experience with the Hadoop ecosystem, including HDFS and YARN.
- Experience supporting Big Data platforms in a production environment.
- Experience troubleshooting Hadoop services using system logs, Hadoop logs and monitoring/alerting platforms.
- Strong Linux administration experience.
- Experience with security and access-control mechanisms such as PAM and Kerberos.
- Strong scripting experience using Shell, Python or Perl.
- Experience with cluster-wide monitoring and performance troubleshooting.
- Experience with OS-level patching and changes in clustered environments.
- Experience in financial services or banking environments is required.
- Bachelor's degree in Engineering, Computer Science, Information Technology or a related discipline.
- Strong analytical and troubleshooting skills.
- Ability to work independently on complex technical issues.
- Ability to work effectively with technical teams across infrastructure, development and architecture.
- Strong communication and interpersonal skills.
- Ability to work under pressure and manage high-severity production incidents.
Preferred Experience
The following would be advantageous:
- Spark development or administration.
- Java, Python, Perl, PHP, HTML or CSS development exposure.
- Apache Ambari.
- Elasticsearch.
- ETL technologies.
- Hadoop 2.x clusters.
- YARN-based applications.
- Experience working with large-scale data-centre environments.
- French language proficiency.
Soft Skills
- Strong analytical and problem-solving ability.
- Ability to prioritize competing technical issues.
- Ability to work autonomously.
- Strong teamwork and collaboration skills.
- Ability to adapt to changing technologies and environments.
- Strong decision-making and incident-management skills.
- Ability to communicate complex technical issues clearly.
- Willingness to mentor and develop junior/partner resources.
- Strong focus on continuous improvement and service quality.
Working Environment
The role will primarily operate during Europe-aligned business hours.
The successful candidate must also be willing to participate in rotational on-call support for production Big Data services.
Key Technology Areas
Hadoop | HDFS | YARN | MapReduce | Hive | Kafka | Hortonworks | Linux | Kerberos | PAM | Python | Perl | Shell | Spark | Elasticsearch | SQL | VMware/Virtual Machines | ETL | JVM