Lead Site Reliability Engineer
Summary
Lead a small SRE team keeping an application-monitoring platform reliable and secure 24/7. Player-coach role: own reliability/security strategy, join on-call and run blameless postmortems, while staying hands-on with bare-metal infra managed by Ansible, Rust/Kafka pipelines, a Rails app, and MongoDB, ClickHouse and ElasticSearch.
Lead Site Reliability Engineer (SRE)
Ocho are working with a client to find a Lead Site Reliability Engineer (SRE) to lead the team responsible for keeping their platform running reliably and securely, 24/7.
Our client helps thousands of teams in 60+ countries monitor and improve their applications, and is remote-first, valuing impact, transparency and continuous improvement.
The role
This is a player-coach position. You'll set the technical direction and own reliability and security strategy for the platform, while staying hands-on with the systems your team runs. It's a small team with a long-standing habit of fixing root causes, not just alerts, and they're now growing it as the business scales.
Their stack
- Mostly bare-metal infrastructure, managed by Ansible
- Data ingestion and processing in Rust, running on Kafka
- A Rails app serving the customer-facing UI
- MongoDB, ClickHouse and ElasticSearch
Responsibilities
- Lead the SRE team: set priorities, mentor engineers, grow the team
- Own reliability strategy and the long-term infrastructure roadmap
- Be part of the on-call rotation, and keep improving it
- Act as incident coordinator, and lead blameless postmortems
- Guide strategic projects, including new AWS infrastructure
- Stay hands-on: tune the Rust codebase and infrastructure automation
- Handle security researcher reports, coordinate penetration tests, support ISO renewals
What you bring
- 8+ years keeping large Linux systems reliable, with experience leading an SRE, platform or infrastructure team (formally or as a technical lead).
- Competent developer across multiple languages, ideally with Rust and Ansible experience.
- Strong incident response and postmortem experience, comfortable translating business growth into infrastructure strategy.
- Bonus: AWS, Kubernetes and Docker.
What's on offer
- Competitive salary
- Remote-first culture - UK wide
- Stock options,
- Flexible PTO
- Personal development budget.
Please apply now if you are meeting the above criteria or contact Andrew Harrison directly.