Senior Data Reliability Engineer
You will drive engagement with Site Reliability across the full breadth of engineering, holding every engineer and team accountable in building highly-resilient, robust, reliable software. You'll join a cross-functional, cross-discipline team of subject matter experts and on-callers whose mission is to keep the platform highly performant 24/7/365. You'll oversee site reliability for enterprise grade applications running thousands of QPS, and you'll play a critical role in defining and building a market-leading foundation for data quality and control, powering observability, quality, data lineage, and remediation across the data and intelligence platform.
Responsibilities
- Evangelise SRE and DRE practices across engineering
- Lead the development of a data quality framework that guarantees data fidelity for customers and supports marketing and revenue functions
- Define and own the on-call process
- Establish a strong working knowledge of company systems
- Command incidents
- Run mop-ups
- Ensure follow-up actions are completed on schedule
- Evaluate and improve the existing end-to-end on-call process
- Take part in the on-call rotation, one week every 4-5 weeks with 24x7x365 coverage
- Evaluate, manage and maintain existing solutions for monitoring, alerting, paging, response and documentation
- Report on uptime, availability and performance across the product suite
- Write post-mortems for internal and external consumption
- Represent the SRE and DRE function on sales calls with tier one enterprise financial institutions
- Work with product, sales and customer service to define SLAs for different products and use cases
- Work with internal product teams to define SLOs for internal consumption and measurement
- Work with engineering teams directly to embed DRE practices
Requirements
- Proven experience at leveling up the quality and reliability of large datasets not just services and APIs
- Experience leading site reliability for a high volume SaaS product
- Experience supporting distributed systems in AWS
- Presence and empathy required to hold teams to account
- Experience defining SLAs and SLOs both internal and client facing
- Experience delivering post mortems to enterprise clients verbally and in writing
- Fluency in AI and agentic workflows
- Genuine interest in the crypto ecosystem is a bonus
- Working knowledge of Kubernetes is a bonus
Benefits
- Hybrid working with option to work from almost anywhere for up to 90 days per year
- £500 remote working budget for home office setup
- $1,000 learning and development budget
- 25 days of annual leave plus bank holidays
- Extra day off for your birthday
- Enhanced parental leave: 16 weeks fully-paid leave
- Private health insurance with Vitality
- Full access to Spill mental health support
- Life assurance covering 4 times your salary
- Cycle to Work Scheme