Site Reliability Engineer
Summary
Build and maintain the reliability, observability and automation of Westpac’s digital banking platforms using SRE practices, cloud-native tech and AI-driven operations.
Site Reliability Engineer
Location: Sydney, NSW, Australia
Westpac One Digital is seeking a highly motivated Site Reliability Engineer to help build, operate and continuously improve the reliability, resilience, observability and performance of critical digital banking platforms. This role combines strong Site Reliability Engineering practices with modern AI-driven engineering to shape Agentic SRE and intelligent operations across Westpac One Digital, ensuring highly available, scalable and secure customer-facing services while driving automation, operational excellence and continuous improvement.
Key Responsibilities
- Improve service reliability, availability and resilience by defining and managing SLIs, SLOs, Error Budgets, disaster recovery capabilities and operational readiness activities.
- Design and enhance observability through monitoring, alerting, dashboards, synthetic monitoring, automated health checks and end-to-end visibility across applications, infrastructure and cloud environments.
- Drive automation and engineering excellence by developing self-healing solutions, operational tooling, platform automation, Infrastructure as Code and reliable CI/CD deployment practices.
- Participate in major incident management, service restoration and root cause analysis, implementing permanent solutions to reduce recurring issues, customer impact and recovery times.
- Lead the adoption of AI-driven operations, Agentic SRE capabilities, LLM-powered incident management and GitHub Copilot-enabled engineering practices to improve efficiency and reduce operational toil.
- Proven experience in Site Reliability Engineering, Production Engineering, DevOps or Platform Engineering, with a strong understanding of distributed systems, cloud-native technologies and modern application architectures.
- Hands-on experience with AWS and/or Azure, Kubernetes, container platforms, CI/CD pipelines and Infrastructure as Code (IaC) practices.
- Strong experience designing and supporting observability solutions using tools such as Splunk, Dynatrace, Grafana, Prometheus, OpenTelemetry or equivalent monitoring platforms.
- Proficiency in software development and automation using languages such as Python, Java, Go, PowerShell or similar scripting and programming technologies.
- Experience leveraging AI-assisted engineering tools including GitHub Copilot, Microsoft Copilot, Claude, ChatGPT or similar technologies, with an understanding of LLMs, AI agents, prompt engineering, AI governance and responsible AI principles.
- Demonstrated expertise in operational excellence, including incident management, problem management, change management, root cause analysis, performance engineering, capacity planning, disaster recovery testing and operational readiness.
- Desirable experience within banking or financial services, including knowledge of APRA CPS 230 Operational Resilience requirements, support of mission-critical digital banking or payments platforms, and exposure to AIOps, AI agents or autonomous operational capabilities.
- Special offers on banking products and discounts from top brands, including generous employee-only mortgage rates!
- Flexible work arrangements to help you achieve a greater work/life balance, and a variety of leave options including Culture, Lifestyle and Wellbeing leave.
- Tailored learning and development opportunities to help your grow your career within the bank.
- Lots of opportunities to ‘give back’ to the Community by getting involved in our many volunteering initiatives.