hipages | Senior Site Reliability Engineer
Overview
As a Senior Software Engineer within the SRE team, you will own the delivery of key infrastructure and reliability features for your team, applying a software engineering mindset to solve complex operational challenges. You will take full responsibility for the technical implementation of features you own — from initial design through to their behaviour in production.Your focus will be on building robust, scalable, and observable platform components on our AWS and Kubernetes-based infrastructure, improving developer experience, and lifting the reliability of the services your team owns. You will set a high bar for engineering quality, mentor engineers around you, and ensure the reliability and performance of what you ship align with team and business goals.
Responsibilities
- Take full ownership of the technical implementation of key platform and reliability features — from design through to production behaviour — ensuring solutions on our AWS and Kubernetes-based infrastructure are robust, scalable, and maintainable.
- Design for operability by building in observability, graceful degradation, and recoverability so the systems your team owns are understandable and safe to run in production.
- Define and maintain SLOs for services your team owns, working with product and stakeholders to set meaningful reliability targets, and make data-driven cases for system health improvements.
- Lead incident response for your team’s systems — coordinating investigation, communicating user impact to stakeholders, and driving post-mortems through to completed corrective actions.
- Actively reduce toil and incident risk by identifying and eliminating root causes, improving runbooks, and keeping on-call practices healthy (fair rosters, current documentation, clean handovers)
- Champion a developer-centric platform with efficient CI/CD pipelines (GitHub Actions, ArgoCD) and internal tooling that removes bottlenecks and enables high team velocity.
- Use AI-assisted engineering workflows (e.g. Claude Code) to improve delivery velocity, code quality, and developer experience.
Qualifications
- 5+ years in infrastructure or platform/SRE engineering roles, with hands-on delivery in high-scale production environments.
- Deep hands-on expertise with AWS, Kubernetes (pref EKS), and Terraform, and modern CI/CD tooling (e.g. GitHub Actions, ArgoCD).
- Solid understanding of Site Reliability principles — SLOs, error budgets, alerting, on-call, and the incident lifecycle - with proven experience running incident response and post-mortems.
- Ability to own the technical implementation of features end-to-end, making sound independent decisions in the face of open-ended requirements.
- Hands-on experience using AI-assisted development tools (e.g. Claude Code) in day-to-day engineering work.
- Clear written and verbal communication - able to present technical arguments concisely and adapt the message to the audience
Nice to have
- Experience with observability/incident management platforms such as Honeycomb.io or Incident.io.
- Experience with messaging technologies such as Apache Kafka (or AWS MSK), RabbitMQ, AWS SQS etc.
- Experience with databases such as RDS, MySQL, or PostgreSQL.