SRE Monitoring Platform Software Engineer Early Career Temporary

Summary

Build components for a multi-region observability and automation platform—ingestion, query, storage, alerting, topology, remediation, and workflow code—using Go/Python/Java/Rust with OpenTelemetry, Prometheus, Grafana, and Kubernetes.

You will build well-scoped components for a multi-region observability and automation platform. You will write ingestion, query, storage, alerting, topology, remediation, and workflow code; instrument services; create dashboards and runbooks; write comprehensive tests; and participate in on-call as a shadow before taking primary responsibility.

Responsibilities

  • Build collection ingestion query and storage components
  • Implement alert correlation and SLO functionality
  • Develop topology and cluster-health integrations
  • Build remediation workflow and job-scheduling components
  • Instrument services with metrics logs and traces using OpenTelemetry
  • Create dashboards and operational runbooks
  • Write unit integration and contract tests
  • Participate in chaos soak testing and on-call operations

Requirements

  • 0–2 years of software engineering experience; strong projects or internships are welcome
  • Programming ability in Go Python Java or Rust
  • Knowledge of data structures algorithms concurrency networking and operating systems
  • Understanding of distributed systems concepts such as idempotency retries back-pressure caching and eventual consistency
  • Exposure to Prometheus Grafana Loki or similar observability tools
  • Familiarity with Linux shell and system debugging tools
  • Basic Kubernetes knowledge
  • Experience with Git and CI pipelines
  • Habit of writing unit and integration tests
  • Clear written and verbal English communication

See also

SRE jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available