SRE Monitoring Platform Software Engineer Early Career Temporary
Summary
Build components for a multi-region observability and automation platform—ingestion, query, storage, alerting, topology, remediation, and workflow code—using Go/Python/Java/Rust with OpenTelemetry, Prometheus, Grafana, and Kubernetes.
You will build well-scoped components for a multi-region observability and automation platform. You will write ingestion, query, storage, alerting, topology, remediation, and workflow code; instrument services; create dashboards and runbooks; write comprehensive tests; and participate in on-call as a shadow before taking primary responsibility.
Responsibilities
- Build collection ingestion query and storage components
- Implement alert correlation and SLO functionality
- Develop topology and cluster-health integrations
- Build remediation workflow and job-scheduling components
- Instrument services with metrics logs and traces using OpenTelemetry
- Create dashboards and operational runbooks
- Write unit integration and contract tests
- Participate in chaos soak testing and on-call operations
Requirements
- 0–2 years of software engineering experience; strong projects or internships are welcome
- Programming ability in Go Python Java or Rust
- Knowledge of data structures algorithms concurrency networking and operating systems
- Understanding of distributed systems concepts such as idempotency retries back-pressure caching and eventual consistency
- Exposure to Prometheus Grafana Loki or similar observability tools
- Familiarity with Linux shell and system debugging tools
- Basic Kubernetes knowledge
- Experience with Git and CI pipelines
- Habit of writing unit and integration tests
- Clear written and verbal English communication