Point your AI agent at freehire and let it find you a job.

Get the CLI →

NEXBRIDGE RECRUITMENT PTE. LTD.

Open 31d reposted 4× · 4 open copies

Data Engineer - Observability / SRE

Posted Updated 5 views
Discussion

Summary

Builds and maintains observability for applications and infrastructure across cloud and hybrid environments — dashboards, monitoring, alerts, logs and traces — while supporting incident investigation and reliability improvements. Core stack includes OpenTelemetry, Grafana, Prometheus, Dynatrace, Elastic, AWS, Docker and Terraform.

Job Responsibilities

  • Design and maintain observability solutions across applications, infrastructure, cloud and hybrid environments.
  • Work with metrics, logs, events and traces to provide visibility into system health and performance.
  • Develop and maintain dashboards, monitoring solutions, alerts and service health indicators.
  • Support application and infrastructure teams with instrumentation and telemetry integration.
  • Work with technologies such as OpenTelemetry, Grafana, Prometheus, Dynatrace, Elastic or equivalent platforms.
  • Support incident investigation, troubleshooting and root-cause analysis.
  • Identify monitoring gaps and recommend improvements to system reliability and performance.
  • Work with engineering, infrastructure, network, security and platform teams on operational improvements.
  • Maintain technical documentation, operational procedures and runbooks.

Job Requirement

  • Minimum 3 years of experience in Observability, SRE, Platform Engineering, Infrastructure Engineering, DevOps or related areas.
  • Hands-on experience with production monitoring and observability environments.
  • Good understanding of metrics, logging, tracing, dashboards and alerting.
  • Experience with AWS and/or cloud environments.
  • Familiarity with Docker, CI/CD, Terraform/OpenTofu, Ansible or similar technologies.
  • Experience working with hybrid or distributed environments.
  • Strong troubleshooting and problem-solving skills.

Good to Have

  • Experience with OpenTelemetry at scale.
  • AWS or Azure certification.
  • Experience with large-scale or distributed environments.
  • Familiarity with SRE practices, SLOs, incident response and reliability engineering.

We regret that only shortlisted candidates will be notified.

Skills

Apply

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available