Site Reliability Engineer
Summary
Improve reliability of market-critical trading systems by automating infrastructure, CI/CD, and observability while leading incident response and 24/7 on-call support.
You will improve the reliability of market-critical services through automation, observability, resilience engineering and operational readiness. You will develop CI/CD and infrastructure-as-code solutions, support capacity and performance planning, manage incidents and changes, and provide rostered 24/7 on-call support.
Responsibilities
- Reduce operational toil through automation
- Improve observability across logs, metrics and traces
- Support production readiness and non-functional testing
- Drive resilience, capacity, incident learning and reliability improvement
- Design and maintain CI/CD pipelines
- Develop and maintain infrastructure as code
- Automate operational tasks, deployments and service recovery
- Conduct production readiness assessments
- Support capacity planning and performance engineering
- Validate failover and recovery and participate in resilience exercises
- Lead or contribute to post-incident reviews
- Design observability practices and actionable alerts
- Provide rostered 24/7 on-call support
- Perform weekend and after-hours installations and upgrades
- Manage incidents, problems, releases and changes
- Undertake business-as-usual team work
Requirements
- 5+ years of experience in a similar SRE role
- Experience with incident response, post-incident review and problem management
- Experience with production readiness, release readiness and operational acceptance
- Experience with capacity, performance and resilience testing
- Knowledge of observability design
- Experience supporting high-availability distributed business-critical platforms
- Scripting and automation skills using Python, PowerShell and shell scripting
- AWS experience with EC2, S3, Lambda and RDS
- Understanding of microservices and containerisation with Docker
- Kubernetes administration experience
- CI/CD pipeline experience
- Experience with CloudWatch, Grafana, Prometheus and OpenTelemetry
- Database operations experience with Oracle and/or Microsoft SQL Server
- Linux and Unix administration and troubleshooting experience
- Microsoft Windows Server 2019-2022 skills
- Troubleshooting, problem-solving and root cause analysis skills
- AWS certification at Associate level or above
- Experience with distributed transactions, high availability and performance-critical systems
- Networking troubleshooting across DNS, TLS, load balancers, firewalls and TCP/IP
Benefits
- Hybrid working
- Flexible working arrangements