Engineering Manager, Platform Engineering – Infra
- Leading engineering teams responsible for platforms, tooling and capabilities enabling application teams to build, deploy, observe and operate software reliably
- Leading approximately 10–12 engineers across Delivery and Observability
- Working closely with Tech Leads, Architects, application engineering teams and other Engineering Leaders to set direction, prioritise improvements and ensure platform delivers reliable and effective capabilities
- Owning roadmap and delivery including prioritisation, planning and execution across the domain
- Building the right team structure, skills and culture while owning recruitment, performance, career development and succession planning
- Partnering with Tech Leads and Architects to challenge technical decisions, manage risks and ensure platform solutions are reliable, secure and maintainable
- Owning platform performance and stability by monitoring SLAs/SLOs, service health and engineering metrics, driving continuous improvement, reducing incidents and improving reliability
- Leading operational excellence including operational readiness, observability, incident management and effective RCA
- Acting as an escalation point for major incidents and ensuring preventative actions are delivered
- Partnering across the organisation with application engineering teams, Architects and Engineering Leaders to understand needs, manage expectations, remove blockers and continuously improve platform capabilities and ways of working
- Strong technical background with ability to understand complex systems, review changes and challenge technical decisions
- Experience leading and developing engineering teams including recruitment, performance and career development
- Strong roadmap, prioritisation and delivery management skills with accountability for quality and timelines
- Experience working with Tech Leads, Architects, Product and engineering stakeholders
- Good understanding of platform engineering, cloud, CI/CD, automation and Infrastructure as Code
- Strong understanding of observability, reliability and production operations
- Experience managing SLAs/SLOs, service performance and operational metrics with focus on improving stability and reducing incidents
- Strong communication, stakeholder management and problem-solving skills with pragmatic and outcome-focused approach
- Experience managing teams responsible for CI/CD, developer platforms, engineering tooling or software delivery platforms
- Experience managing Observability, SRE or Platform Reliability teams
- Experience with AWS or other major cloud platforms
- Familiarity with technologies such as Kubernetes, GitHub Actions, Argo CD, Terraform, Ansible, Prometheus, Grafana, OpenTelemetry or similar tooling