Data Engineer - Observability / SRE
Summary
This role designs and maintains observability for applications, infrastructure, and cloud/hybrid environments — building dashboards, alerts, and telemetry (metrics, logs, traces), supporting incident investigation and root-cause analysis, and improving reliability with teams like SRE, platform, and security. Core stack includes OpenTelemetry, Grafana, Prometheus, Dynatrace, Elastic, and AWS.
Job Responsibilities
- Design and maintain observability solutions across applications, infrastructure, cloud and hybrid environments.
- Work with metrics, logs, events and traces to provide visibility into system health and performance.
- Develop and maintain dashboards, monitoring solutions, alerts and service health indicators.
- Support application and infrastructure teams with instrumentation and telemetry integration.
- Work with technologies such as OpenTelemetry, Grafana, Prometheus, Dynatrace, Elastic or equivalent platforms.
- Support incident investigation, troubleshooting and root-cause analysis.
- Identify monitoring gaps and recommend improvements to system reliability and performance.
- Work with engineering, infrastructure, network, security and platform teams on operational improvements.
- Maintain technical documentation, operational procedures and runbooks.
Job Requirement
- Minimum 3 years of experience in Observability, SRE, Platform Engineering, Infrastructure Engineering, DevOps or related areas.
- Hands-on experience with production monitoring and observability environments.
- Good understanding of metrics, logging, tracing, dashboards and alerting.
- Experience with AWS and/or cloud environments.
- Familiarity with Docker, CI/CD, Terraform/OpenTofu, Ansible or similar technologies.
- Experience working with hybrid or distributed environments.
- Strong troubleshooting and problem-solving skills.
Good to Have
- Experience with OpenTelemetry at scale.
- AWS or Azure certification.
- Experience with large-scale or distributed environments.
- Familiarity with SRE practices, SLOs, incident response and reliability engineering.
Interested candidates who wish to apply for the above positions, please click "Apply now".
We regret that only shortlisted candidates will be notified.
Nexbridge Recruitment Pte Ltd
EA License 26C3616
Tell employers what skills you have Distributed Computing EnvironmentIncident Support
Skip Tracing
AWS
reliability monitoring
Operational Monitoring
OpenTelemetry
Analytical and Problem-Solving Skills
CI/CD
Telemetry
Engineering
Health Services
new technologies
Ansible
Hybrid Cloud
Metrics