Point your AI agent at freehire and let it find you a job.

Get the CLI →

ATS CONSULTING SERVICES PH INC.

NewBe an early applicant

Senior Data Platform Reliability Engineer

Posted 1 view
Discussion

Summary

Operates, maintains, and continuously improves the company's data platforms on Kubernetes (on-prem and AWS/GCP EKS/GKE), handling GitOps deployments, monitoring, incident response on 24x7 rotations, and troubleshooting user issues while mentoring junior engineers. Core stack includes Kubernetes, Linux, and ETL/ELT tooling like Spark, Airflow, and Python/SQL.

About the role

As a Senior Data Platform Reliability Engineer, you will be responsible for operating, maintaining, and continuously improving the company's data platforms running on Kubernetes (on-premises and/or on AWS/GCP) - similar to the DoEKS (Data on EKS) / AIoEKS (AI on EKS) deployment frameworks.

Key responsibilities

  • Deploy new releases and configuration changes through GitOps/DevOps

  • Monitor platform and service health using logs, metrics, and observability tools

  • Participate in incident response, root cause analysis, and 24x7 operational rotations

  • Improve platform observability, operational tooling/automations, self-service capabilities, and reliability practices to reduce recurring issues

  • Investigate & troubleshoot user concerns by either correlating them to system-related issues, breaking integrations, and/or user-specific errors/misconfigurations up to recommending/executing resolutions

  • Provide technical mentorship to junior engineers

  • Advocate for platform standards, security best practices, and operational excellence

About you

  • 3+ years of solid experience supporting production data workloads/platforms (Spark/Airflow/Jupyter)

  • 5+ years of hands-on experience on ETL/ELT pipeline development & data transformations (Python/Java & SQL)

  • Practical proficiency in Kubernetes environments including Cloud-provider managed Kubernetes flavors (AWS-EKS/GCP-GKE)

  • Comprehensive knowledge on Linux environments, microservice architectures and service communication patterns

  • Strong troubleshooting fundamentals such as application crashes, resource contentions, service latency, and scaling behaviour

  • Well-rounded competency in analysing logs, metrics, monitoring systems, and service KPIs

  • Exposure to other Data/AI platforms such as Flink, Trino, Druid, and Ray

  • Hands-on experience with automation or scripting (Bash, Python)

  • Kubernetes or Data certifications (CKAD, AWS Certified Data Engineer)

Skills

What Senior Data Engineering jobs ask for — and how much of it you have →
Apply

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available