freehire launches on Product Hunt on 26 August.

Follow →

Site Reliability Engineer Telemetry

Summary

Maintain and improve the company’s telemetry stack—metrics, logs, traces, dashboards, and alerting—using Prometheus, Grafana, and related tools to keep production systems reliable.

You will operate and improve shared telemetry platforms for metrics, logs, traces, alerting, dashboards, and profiling. You will manage telemetry services and pipelines, troubleshoot production issues, build automation, participate in incident response and on-call, write runbooks, and improve reliability based on incidents.

Responsibilities

  • Operate and improve the shared telemetry platform
  • Maintain metrics collection, storage, querying, dashboards, and alerting
  • Operate log pipelines
  • Operate distributed tracing and profiling capabilities
  • Deploy and manage telemetry services with Terraform and Terragrunt
  • Troubleshoot missing data, slow queries, broken alerts, backpressure, and capacity issues
  • Build reusable configuration and automation
  • Participate in incident response and on-call
  • Write runbooks
  • Improve the platform using incident findings

Requirements

  • 3+ years of experience as a Site Reliability Engineer, Platform Engineer, Infrastructure Engineer, Observability Engineer, or similar
  • Production systems experience at scale
  • Prometheus or a Prometheus-compatible monitoring stack
  • Distributed systems troubleshooting
  • Infrastructure as Code
  • Terraform
  • CI/CD
  • Nomad, Kubernetes, or similar platforms
  • Scripting or programming ability
  • Incident response
  • Documentation and collaboration skills
  • VictoriaMetrics, Grafana, Tempo, Loki, Vector, Splunk, Alertmanager, or OpenTelemetry experience is a plus
  • PromQL or LogQL experience is a plus

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available