Senior Data Platform Engineer
Summary
Build and maintain Tookitaki's cloud-native platform on AWS EKS using GitOps (Argo), Terraform, and Kubernetes operators, enabling self-service infrastructure for application teams.
Location: manila,,
Senior Data Platform Engineer — Cloudera, AWS & KubernetesThe mission We are looking for a battle-tested Data Platform Engineer who can build, operate, troubleshoot, and evolve large-scale data platforms across on-premises Cloudera environments and cloud-native AWS/Kubernetes architectures. This is not a dashboard or SQL-only role. You will work where distributed compute, storage, networking, Kubernetes, and production data pipelines meet. You must be comfortable tracing a failed workload from the application layer through Spark or Flink, Kubernetes operators, HDFS/Hive, infrastructure, and AWS services. The goal is to help us evolve safely from on-premises, VM, and EC2-based platforms into resilient, observable, cloud-native data systems. What you'll do - Build and operate production data platforms across Cloudera on premises, Cloudera cloud environments, AWS EMR, and Kubernetes. - Design batch and streaming pipelines using Spark, Flink, Hive, and related technologies. - Transform large datasets through filtering, sorting, joining, aggregation, partitioning, enrichment, and restructuring. - Work with Parquet, Avro, JSON, CSV, and other delimited or semi-structured formats. - Design storage, partitioning, compression, retention, and lifecycle strategies across HDFS, Hive, and object storage. - Design and maintain Hive schemas, tables, partitions, metadata, and data models. - Administer Cloudera clusters, including installation, upgrades, configuration, scaling, patching, security, backup, and recovery. - Troubleshoot unhealthy services, failed jobs, resource contention, data skew, small-file problems, metadata issues, and storage bottlenecks. - Tune Spark and Flink workloads for memory, CPU, parallelism, shuffle behavior, checkpointing, and recovery. - Operate AWS services such as EMR, S3, IAM, EC2, EKS, CloudWatch, KMS, and supporting networking services. - Deploy and operate data workloads on Kubernetes using operators, Helm, custom resources, and GitOps-based delivery. - Help migrate workloads from physical or virtual machines and EC2 into AWS and Kubernetes-based platforms. - Separate compute from storage where appropriate while accounting for performance, resilience, security, and cost. - Build monitoring, alerting, capacity management, and operational runbooks for critical data services. - Automate platform provisioning and configuration using Terraform, Ansible, scripting, and CI/CD. - Support production incidents involving failed pipelines, delayed data, cluster degradation, storage pressure, or infrastructure failure. - Work with data engineering, infrastructure, security, and application teams to resolve problems across ownership boundaries. What we're looking for - Strong hands-on experience administering Cloudera platforms in production. - Experience with both on-premises Cloudera and cloud-based Cloudera deployments. - Deep working knowledge of Hadoop, HDFS, Hive, YARN, Spark, and the wider distributed-data ecosystem. - Experience building or operating production workloads using Apache Flink. - Strong understanding of distributed data processing, including partitioning, shuffling, serialization, checkpointing, and failure recovery. - Experience transforming large datasets using joins, aggregations, filtering, sorting, and schema evolution. - Practical knowledge of Parquet, Avro, JSON, CSV, compression formats, and serialization tradeoffs. - Experience designing data layouts for query performance, ingestion throughput, retention, and cost. - Strong AWS experience, particularly with EMR, S3, EC2, EKS, IAM, CloudWatch, and KMS. - Strong Kubernetes experience, including operators, controllers, Helm, scheduling, storage, networking, and workload troubleshooting. - Experience migrating data platforms from on-premises or VM-based environments into AWS and Kubernetes. - Ability to troubleshoot Linux, JVM, networking, storage, DNS, certificates, and resource-management issues. - Experience with observability platforms and the ability to correlate infrastructure symptoms with data-pipeline failures. - Ability to automate operational work using Python, Bash, Terraform, Ansible, or equivalent tools. - Strong judgment around production changes, data integrity, access control, rollback, and recovery. Production scenarios you should be able to handle - A Spark job that ran in 40 minutes yesterday now takes four hours. - A join creates severe data skew and repeatedly exhausts executor memory. - HDFS is approaching capacity while NameNode health is degrading. - Hive queries return incomplete results because partitions or metadata are inconsistent. - A Flink job repeatedly fails after checkpoint recovery. - An EMR workload is reliable but significantly more expensive than expected. - A Kubernetes operator reports success while the underlying data workload is unhealthy. - A migrated workload behaves differently on S3 than it did on HDFS. - A certificate, Kerberos, IAM, DNS, or network problem presents as an application failure. - A critical pipeline misses its SLA and ownership is unclear across platform and data teams. Certifications Relevant certifications are useful, particularly: - Cloudera Certified Professional: Data Engineer - Cloudera Administrator certification or equivalent production experience - AWS data, analytics, or architecture certifications - Kubernetes certifications such as CKA or CKAD Certification is supporting evidence. The ability to diagnose and recover a real distributed platform matters more. What success looks like - Data pipelines meet their reliability and processing-time objectives. - Platform failures are detected before downstream consumers report them. - Incidents move quickly from symptoms to an evidence-backed root cause. - Cloudera, AWS, and Kubernetes environments are operated through repeatable automation. - Migrations preserve data correctness while improving scalability and operability. - Storage and compute designs balance performance, resilience, and cost. - Data engineers can ship workloads without becoming accidental platform administrators. - Operational knowledge becomes monitoring, automation, and runbooks—not tribal memory. The person we want You understand that a data platform is a distributed production system, not a collection of product names. You can move from a Hive execution plan to Spark executor logs, Kubernetes events, HDFS health, S3 behavior, IAM permissions, and network telemetry without losing the thread. You know the architectural differences between on-premises Hadoop and cloud- native data platforms, including where a lift-and-shift approach will fail. We need someone who can enter a degraded platform, establish the facts, protect data integrity, restore service, explain the failure chain, and make the system harder to break next time.
Apply to this job