Sr. Data Platform Reliability Engineer
Summary
Senior engineer who runs and improves enterprise data platforms on Kubernetes (on-prem and AWS/GCP, similar to Data on EKS), keeping production workloads like Spark, Airflow, and Jupyter reliable and secure. Day to day: GitOps deployments, monitoring/observability, incident response in 24/7 rotations, troubleshooting, and mentoring junior engineers.
Job Opportunity: Sr. Data Platform Reliability Engineer (Manila)
Location: Manila, Philippines
Employment Type: Full-Time
Work Arrangement: Onsite, with automatic work-from-home arrangements on weekends and holidays
Schedule: Shifting schedules (6:00 AM, 10:00 AM, or 2:00 PM)
About the Role
We are looking for a Senior Data Platform Reliability Engineer to operate, maintain, and continuously improve enterprise data platforms running on Kubernetes, whether on-premises or on AWS/GCP.
The role involves supporting production data workloads and ensuring platform reliability, performance, security, and operational efficiency. You will work with technologies and deployment frameworks similar to Data on EKS (DoEKS) and AI on EKS (AIoEKS).
This position is ideal for an experienced data platform professional with strong expertise in Kubernetes, data engineering, production operations, and incident management.
Key Responsibilities
- Operate, maintain, and continuously improve data platforms running on Kubernetes across on-premises and cloud environments.
- Deploy new releases and configuration changes through GitOps and DevOps practices.
- Monitor platform and service health using logs, metrics, and observability tools.
- Participate in incident response, root cause analysis, and 24/7 operational rotations.
- Improve platform observability, operational tooling, automation, and self-service capabilities to reduce recurring issues.
- Investigate and troubleshoot user concerns by identifying system-related issues, broken integrations, and user-specific errors or misconfigurations.
- Recommend and execute appropriate resolutions to platform and service issues.
- Provide technical mentorship and guidance to junior engineers.
- Advocate for platform standards, security best practices, and operational excellence.
Qualifications and Requirements
- At least 3 years of solid experience supporting production data workloads and platforms, such as Spark, Airflow, and Jupyter.
- At least 5 years of hands-on experience in ETL/ELT pipeline development and data transformations using Python/Java and SQL.
- Practical proficiency in Kubernetes environments, including cloud-managed Kubernetes services such as AWS EKS and GCP GKE.
- Comprehensive knowledge of Linux environments, microservice architectures, and service communication patterns.
- Strong troubleshooting skills, including application crashes, resource contention, service latency, and scaling behavior.
- Well-rounded competency in analyzing logs, metrics, monitoring systems, and service KPIs.
Nice-to-Have Qualifications
- Exposure to other data and AI platforms such as Flink, Trino, Druid, and Ray.
- Hands-on experience with automation and scripting using Bash and Python.
- Relevant Kubernetes or data certifications, such as:
- Certified Kubernetes Application Developer (CKAD)
- AWS Certified Data Engineer
Work Arrangement and Benefits
- Onsite work arrangement in Manila.
- Shifting schedules: 6:00 AM, 10:00 AM, or 2:00 PM.
- Automatic work-from-home arrangement when your scheduled shift falls on a weekend or holiday.
- HMO coverage starting Day 1.
- Transportation allowances provided.
Why Join Us?
Be part of a team responsible for maintaining and improving enterprise data platforms that support critical production workloads. This opportunity allows you to apply your expertise in data engineering, Kubernetes, automation, and reliability engineering while mentoring junior engineers and contributing to operational excellence.