freehire launches on Product Hunt on 26 August.

Follow →

Senior Data Platform Reliability Engineer

Our client is a technology consulting company providing operational and engineering services to the high-tech sector. They support platform and infrastructure teams in managing multi-cloud environments, delivering complex migrations, and enabling reliable application deployments at scale.

The Role

Our client is looking for a Senior Data Platform Reliability Engineer to operate, maintain, and continuously improve production data platforms running on Kubernetes across on-premise, AWS, and GCP environments.

You will work across platform reliability, observability, automation, incident management, and data infrastructure, supporting environments similar to Data on EKS and AI on EKS architectures.


Key Responsibilities

  • Deploy platform releases and configuration updates using GitOps and DevOps practices.
  • Monitor platform and service health through logs, metrics, monitoring, and observability tooling.
  • Participate in incident response, root cause analysis, and 24/7 operational rotations.
  • Improve platform reliability through better observability, automation, operational tooling, and self-service capabilities.
  • Investigate and troubleshoot user and platform issues, including system failures, broken integrations, configuration problems, and application-level errors.
  • Recommend and implement appropriate technical resolutions.
  • Mentor junior engineers and support the development of engineering best practices.
  • Promote strong standards across platform operations, security, reliability, and engineering quality.

Qualifications

  • 3+ years of experience supporting production data platforms or workloads using technologies such as Spark, Airflow, or Jupyter.
  • 5+ years of hands-on experience developing ETL/ELT pipelines and data transformations using Python or Java and SQL.
  • Strong practical experience working with Kubernetes, including managed Kubernetes platforms such as AWS EKS or Google GKE.
  • Strong knowledge of Linux environments, microservices architectures, and service communication patterns.
  • Solid troubleshooting skills across application crashes, resource contention, service latency, performance, and scaling issues.
  • Experience analysing logs, metrics, monitoring systems, and service-level KPIs.

Nice to Have

  • Experience with additional Data or AI platforms such as Flink, Trino, Druid, or Ray.
  • Hands-on scripting and automation experience using Bash or Python.
  • Relevant Kubernetes, cloud, or data certifications such as CKAD or AWS Certified Data Engineer.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available