Senior SRE: AI-Scale Infra, Observability & Resilience
Posted Updated
Domyn in Milan is seeking an experienced Site Reliability Engineer to help shape the future of Colosseum, our flagship AI supercomputer in development. You will design observability and control mechanisms, extract operational data, and feed it into automated systems to optimize power, cooling, and service levels.
You will guard budgets, improve system resilience, and participate in blameless post-mortems. Join a cross-functional team with Platform Engineering to ensure secure, scalable
Design and implement observability and control mechanisms for infrastructure. Extract operational data and feed it into automated systems for optimization. Guard and maintain operational budgets (power, cooling, SLOs). Contribute to blameless post-mortems and continuous improvement. Bachelor’s or Master’s degree in CS/CE/EE or related field. 6+ years as an SRE or similar role. Strong experience with Prometheus, Thanos, Grafana, OpenTelemetry. Experience with low-level instrumentation eBPF. Security monitoring tools such as Zeek or Wazuh. Kubernetes and cloud-native environments; Python automation. Learning Friday Smart Working Stock options