Point your AI agent at freehire and let it find you a job.

Get the CLI →

Domyn

New

Senior SRE: AI-Scale Infra, Observability & Resilience

Posted Updated
Discussion

Domyn in Milan is seeking an experienced Site Reliability Engineer to help shape the future of Colosseum, our flagship AI supercomputer in development. You will design observability and control mechanisms, extract operational data, and feed it into automated systems to optimize power, cooling, and service levels.

You will guard budgets, improve system resilience, and participate in blameless post-mortems. Join a cross-functional team with Platform Engineering to ensure secure, scalable

Design and implement observability and control mechanisms for infrastructure. Extract operational data and feed it into automated systems for optimization. Guard and maintain operational budgets (power, cooling, SLOs). Contribute to blameless post-mortems and continuous improvement. Bachelor’s or Master’s degree in CS/CE/EE or related field. 6+ years as an SRE or similar role. Strong experience with Prometheus, Thanos, Grafana, OpenTelemetry. Experience with low-level instrumentation eBPF. Security monitoring tools such as Zeek or Wazuh. Kubernetes and cloud-native environments; Python automation. Learning Friday Smart Working Stock options

Skills

See also

SRE jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available