Big Data Infrastructure Engineer - Linux & Cloudera
Summary
Maintain and automate a Cloudera/Hadoop big-data platform on Red Hat Linux, ensuring uptime, security, and performance while troubleshooting issues and onboarding new workloads.
We are looking for an experienced Linux Red Hat System Engineer to join a team responsible for the operation, support, and continuous improvement of a critical Cloudera/Hadoop infrastructure. This role offers the opportunity to work in a complex enterprise environment, combining Linux administration, automation, big data operations, and infrastructure engineering.
Location: Lisbon (Hybrid, 2 days/week onsite)
The Role
As part of the Infrastructure Operations team, you will be responsible for ensuring the availability, performance, security, and stability of Cloudera/Hadoop production environments. You will work closely with platform, data, network, storage, backup, security, and Wintel teams, supporting business-critical services and driving operational excellence through automation.
Key Responsibilities
- Operate, monitor, and support Cloudera/Hadoop-based production environments.
- Provide L3 support, troubleshooting, and performance tuning for Big Data platforms.
- Automate operational tasks using Python and Ansible, including monitoring, alerting, deployments, and configuration management.
- Ensure HDFS data integrity, replication, and capacity management.
- Develop and maintain monitoring and alerting solutions for platform services and infrastructure health.
- Apply patches, upgrades, and configuration changes to maintain security and platform stability.
- Manage Kerberos authentication, Ranger/Sentry authorizations, and TLS encryption.
- Collaborate with platform and data teams to onboard new workloads and optimize resource utilization.
- Work closely with stakeholders to deliver custom infrastructure requirements and improvements.
- Maintain technical documentation and contribute to knowledge-sharing initiatives.
Technical Skills
Mandatory
- Red Hat Linux System Administration and Operations
- Cloudera/Hadoop Infrastructure
- Python
- Ansible
- Linux Server Monitoring and Diagnostics (CPU, Memory, Storage)
- JVM and Middleware Troubleshooting
- HDFS Administration and Capacity Management
- Storage Management (SAN, NAS, RAID)
- Incident Management and Root Cause Analysis
Nice to Have
- Cloudera or other Big Data certifications
- Kafka
- NiFi
- Flume
- ITIL processes for Incident and Change Management
What We're Looking For
- Bachelor's Degree in Computer Science or equivalent.
- Minimum 5 years of experience in a similar role.
- Experience supporting mission-critical production environments.
- Strong troubleshooting and problem-solving skills.
- Ability to diagnose and resolve complex issues in Big Data ecosystems.
- Strong communication and stakeholder management skills.
- English level B2-C1.