Big Data Infrastructure Engineer - Linux & Cloudera
We are looking for an experienced Linux Red Hat System Engineer to join a team responsible for the operation, support, and continuous improvement of a critical Cloudera/Hadoop infrastructure. This role offers the opportunity to work in a complex enterprise environment, combining Linux administration, automation, big data operations, and infrastructure engineering.
Location: Lisbon (Hybrid, 2 days/week onsite)
The Role
As part of the Infrastructure Operations team, you will be responsible for ensuring the availability, performance, security, and stability of Cloudera/Hadoop production environments. You will work closely with platform, data, network, storage, backup, security, and Wintel teams, supporting business-critical services and driving operational excellence through automation.
Key Responsibilities
- Operate, monitor, and support Cloudera/Hadoop-based production environments.
- Provide L3 support, troubleshooting, and performance tuning for Big Data platforms.
- Automate operational tasks using Python and Ansible, including monitoring, alerting, deployments, and configuration management.
- Ensure HDFS data integrity, replication, and capacity management.
- Develop and maintain monitoring and alerting solutions for platform services and infrastructure health.
- Apply patches, upgrades, and configuration changes to maintain security and platform stability.
- Manage Kerberos authentication, Ranger/Sentry authorizations, and TLS encryption.
- Collaborate with platform and data teams to onboard new workloads and optimize resource utilization.
- Work closely with stakeholders to deliver custom infrastructure requirements and improvements.
- Maintain technical documentation and contribute to knowledge-sharing initiatives.
Technical Skills
Mandatory
- Red Hat Linux System Administration and Operations
- Cloudera/Hadoop Infrastructure
- Python
- Ansible
- Linux Server Monitoring and Diagnostics (CPU, Memory, Storage)
- JVM and Middleware Troubleshooting
- HDFS Administration and Capacity Management
- Storage Management (SAN, NAS, RAID)
- Incident Management and Root Cause Analysis
Nice to Have
- Cloudera or other Big Data certifications
- Kafka
- NiFi
- Flume
- ITIL processes for Incident and Change Management
What We're Looking For
- Bachelor's Degree in Computer Science or equivalent.
- Minimum 5 years of experience in a similar role.
- Experience supporting mission-critical production environments.
- Strong troubleshooting and problem-solving skills.
- Ability to diagnose and resolve complex issues in Big Data ecosystems.
- Strong communication and stakeholder management skills.
- English level B2-C1.