Site Reliability Engineer
Summary
The Site Reliability Engineer will manage and optimize Azure-based cloud infrastructure, CI/CD pipelines, and observability tools to ensure the reliability of AI-driven cancer diagnostic platforms. The role involves automating operational tasks, participating in an on-call rotation, and leading incident reviews.
H2R Technology is excited to be partnering with Lunit International in the search for a Site Reliability Engineer based in Wellington.
From screening to research and development, Lunit is transforming how the world advances cancer detection - boldly, intelligently, and with humanity at the core.
We build trusted partnerships with clinicians, health systems, and companies worldwide to deliver validated AI solutions that enhance clinical confidence, accelerate discovery, prioritize precision care - and bring the future of cancer diagnostics within reach.
We are currently looking for an intermediate SRE to ensure the reliability, availability, and performance of critical applications. To find out how you could contribute to our mission to conquer cancer as an SRE please read on.
About the Role
You'll work closely with product and engineering teams to improve reliability, automation, observability, and operational excellence across customer-facing platforms.
As part of the Site Reliability Engineering function, you'll investigate incidents, improve system resilience, drive automation, and enable teams to deploy and operate software more effectively.
What You'll Be Doing
From screening to research and development, Lunit is transforming how the world advances cancer detection - boldly, intelligently, and with humanity at the core.
We build trusted partnerships with clinicians, health systems, and companies worldwide to deliver validated AI solutions that enhance clinical confidence, accelerate discovery, prioritize precision care - and bring the future of cancer diagnostics within reach.
We are currently looking for an intermediate SRE to ensure the reliability, availability, and performance of critical applications. To find out how you could contribute to our mission to conquer cancer as an SRE please read on.
About the Role
You'll work closely with product and engineering teams to improve reliability, automation, observability, and operational excellence across customer-facing platforms.
As part of the Site Reliability Engineering function, you'll investigate incidents, improve system resilience, drive automation, and enable teams to deploy and operate software more effectively.
What You'll Be Doing
- Supporting and enhancing Azure-based cloud infrastructure
- Developing and maintaining CI/CD pipelines using GitHub and Azure DevOps
- Building and improving Infrastructure as Code solutions
- Creating tooling and automation to reduce operational overhead
- Improving monitoring, alerting, logging, and observability across platforms
- Working within engineering teams to improve reliability, availability, and deployment practices
- Supporting customer-facing systems and participating in an on-call roster (approximately one week per month)
- Investigating production incidents through to resolution
- Leading blameless post-incident reviews and driving problem management initiatives
- 4+ years' experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, or a similar role
- Strong Azure cloud experience, including managing and supporting production environments
- Hands-on experience designing, building, and supporting CI/CD pipelines using GitHub Workflows and/or Azure DevOps Pipelines
- Infrastructure as Code experience using Bicep, ARM, Terraform, or similar technologies
- Strong scripting and automation skills using PowerShell and/or Bash
- Experience supporting production systems and responding to operational incidents in business-critical environments
- Understanding of monitoring, alerting, logging, and observability practices
- Strong knowledge of Git and source control practices
- SQL relational database experience
- Linux administration and troubleshooting experience
- Azure Entra ID experience, including applications, managed identities, and user management
- Experience with Docker, container technologies, and package repositories
- Knowledge of Information Security frameworks and standards such as ISO/IEC 27001
- Experience with ServiceNow, ITIL practices, or formal incident/problem management processes
- Experience with Snowflake, Dagster, or data pipeline technologies
- Ability to read code and develop small tools or utilities in C#
- Join a global organisation at the forefront of AI-driven healthcare innovation
- Work on technology that directly contributes to improved cancer detection and patient outcomes
- Be part of a collaborative engineering culture focused on continuous improvement
- Work with modern Azure cloud technologies and automation tooling
- Hybrid working environment with excellent employee benefits
- Meaningful work with genuine real-world impact