freehire launches on Product Hunt on 26 August.

Follow →

Data Operations Engineer, X-Layer

The company will be prioritising applicants who have a current right to work in Singapore, and do not require the company's sponsorship of a visa.

About The Opportunity

We are hiring Data Operations Engineers / Data Platform SREs in Singapore to support the reliability, observability, performance, and cost efficiency of our multi‑cloud data platform. You will work on core data platform components across Alibaba Cloud and AWS, covering daily operations, incident response, monitoring, automation, resource optimization, and platform reliability improvements. You will collaborate closely with data engineering, data warehouse, BI, platform teams, and cloud vendors to ensure our data pipelines and platform services run reliably at scale.

What You’ll Be Doing

  • Operate and support core data platform components, including Alibaba Cloud DataWorks, MaxCompute / ODPS, Hologres, VVP / Flink, and AWS‑based platforms such as Databricks and StarRocks.
  • Build and maintain monitoring, alerting, and SLA / SLO metrics for data platform services.
  • Respond to production incidents, troubleshoot issues, participate in post‑incident reviews, and drive long‑term fixes.
  • Analyze compute, storage, job, and cluster resource usage to improve performance and optimise cloud costs.
  • Develop or integrate automation scripts and operational tools to improve health checks, releases, scaling, and troubleshooting efficiency.
  • Collaborate with data engineering, data warehouse, BI, platform teams, and cloud vendors to continuously improve platform reliability and operational efficiency.

What We Look For In You

  • 3+ years of experience in data platform operations, big data engineering, platform engineering, SRE, or cloud infrastructure operations.
  • Experience with either AWS or Alibaba Cloud.
  • Hands‑on experience with at least one big data or cloud data platform, such as MaxCompute / ODPS, Hologres, Databricks, StarRocks, Flink, Spark, Hive, Presto, or Trino.
  • Strong SQL skills, with the ability to troubleshoot data issues, job failures, and performance problems.
  • Solid Linux fundamentals and scripting experience with Shell, Python, or similar languages.
  • Good understanding of monitoring, alerting, capacity, access control, and resource management.
  • Strong incident response and troubleshooting skills, with the ability to drive recovery under pressure.
  • Good communication and collaboration skills in a cross‑functional environment.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available