Platform Infra & SRE Lead (Hybrid Role)
Posted Updated
Deltatre is seeking a Team Lead for Platform Infrastructure and SRE, based in Italy. The role combines hands-on engineering with people leadership for a 5–8 person team, guiding reliability, scalability, and operability of our infrastructure and internal tooling.
You will own technical direction, participate in design and code, and help translate infrastructure needs into a roadmap while mentoring engineers and coordinating with product and other teams.
Set technical direction for platform infrastructure and SRE initiatives, balancing reliability, cost, and delivery speed. Stay hands-on: design reviews, architecture decisions, and direct contribution to high-leverage or high-risk work. Own capacity planning, cost visibility, and infrastructure spend across our cloud providers. Define and track operational health metrics (SLOs, error budgets, MTTR) for owned services. Drive incident response process and post-incident reviews, and use them to prioritise reliability investments. Manage and grow a team of 5 to 8 engineers: 1:1s, career development, performance feedback, and hiring. Partner with product and other engineering teams to translate infrastructure needs into a prioritised roadmap. Represent the team in cross-functional planning and communicate trade-offs to stakeholders and leadership. 6+ years' experience in infrastructure, platform, or SRE engineering, including production ownership at scale. 1+ years' experience leading a team (people management or strong tech lead track record). Deep hands-on experience with Kubernetes and at least one major public cloud (AWS, GCP, or Azure). Working knowledge of a second cloud provider (multi-cloud footprint). Strong experience with infrastructure as code, ideally Terraform. Experience building or operating observability stacks (metrics, logging, tracing). Experience owning CI/CD platforms and pipelines, not just using them. Track record of leading incident response and driving reliability improvements from postmortems. Clear written and verbal communication; able to present trade-offs to non-infra stakeholders. Datadog, Prometheus, and Grafana experience Cost optimisation experience across multiple cloud providers. Prior experience in a regulated or high-availability environment.