freehire launches on Product Hunt on 26 August.

Follow →

Site Reliability Engineer (SRE) / Senior DevOps Engineer – Azure (4 Positions)

Site Reliability Engineer (SRE) / Senior DevOps Engineer – Azure (4 Roles)

Melbourne CBD + Hybrid


Introduction

9X5 Consulting is seeking experienced Site Reliability Engineer (SRE) / Senior DevOps Engineers to join one of our clients and support the reliability, scalability and operational performance of business-critical production platforms.

This role would suit an experienced SRE, DevOps, Platform or Cloud/Infrastructure Engineer who has worked with production systems at scale and has strong hands-on experience across Microsoft Azure, automation, observability and modern infrastructure practices.

You don't necessarily need to have held the title of "Site Reliability Engineer". We are particularly interested in candidates who can demonstrate strong production operations experience and an understanding of SRE principles, including SLIs, SLOs, error budgets, observability and continuous improvement.


About 9X5 Consulting

We are transforming the way business and technology is managed by putting real-time data into the hands of every decision maker across organisations. The insight garnered from diverse backgrounds, perspectives and lived experiences results in pioneering innovations across the organisation and better experiences for our customers. The more diverse our talent, the more impact we have on each other and on our valued clients.

Our clients range from small SME's right through to some of the largest ASX listed organisations in Australia, so experience at selling to stakeholders at this level is required.


A bit about you

You are an experienced SRE, DevOps, Platform or Infrastructure Engineer who enjoys working with complex production environments and finding ways to make systems more reliable, scalable and easier to operate.

You have strong hands-on experience with Microsoft Azure and are comfortable working across cloud infrastructure, automation, CI/CD, monitoring and observability. You understand that reliability is more than simply responding to incidents — it's about proactively identifying issues, automating repetitive tasks and continually improving the environment.

Ideally, you will bring:

  • 3+ years' experience in SRE, DevOps, Platform or Infrastructure Engineering, supporting production services at scale.

  • Strong hands-on Microsoft Azure experience across compute, networking, App Service/AKS, storage and databases.

  • Experience working with SLIs, SLOs and error budgets in production environments.

  • Experience with Azure Monitor, Log Analytics and Application Insights.

  • Infrastructure as Code experience using Terraform and/or Bicep.

  • CI/CD experience using Azure DevOps and/or GitHub Actions.

  • Automation and scripting skills using PowerShell, Python and/or Go.

  • A solid understanding of networking, including DNS, TCP/IP, load balancing and NSGs.

  • Experience supporting Linux and/or Windows Server environments.

  • Experience with incident management, root-cause analysis and blameless post-mortems.

  • Exposure to Docker, Kubernetes/AKS, Prometheus and Grafana would be highly regarded.

  • An understanding of ITIL and experience working within Scrum or Kanban environments.

Most importantly, you're someone who is comfortable taking ownership of production reliability, enjoys solving complex technical problems and works collaboratively with development, infrastructure and operational teams.


Key Responsibilities

As part of the role, you will:

  • Ensure the reliability, availability, performance and scalability of business-critical production services.

  • Define, implement and manage SLIs, SLOs and error budgets to measure and improve service reliability.

  • Design, deploy and support solutions across Microsoft Azure, including compute, networking, App Services, AKS, storage and databases.

  • Build and maintain monitoring, logging and observability solutions using Azure Monitor, Log Analytics and Application Insights.

  • Develop and maintain infrastructure using Terraform and/or Bicep and Infrastructure as Code best practices.

  • Build, maintain and improve CI/CD pipelines using Azure DevOps and/or GitHub Actions.

  • Automate operational and infrastructure processes using PowerShell, Python and/or Go.

  • Monitor production environments, troubleshoot complex issues and respond to incidents.

  • Lead or contribute to root-cause analysis and blameless post-mortems, ensuring lessons learned result in measurable improvements.

  • Work across networking and infrastructure technologies including DNS, TCP/IP, load balancing, NSGs, Linux and Windows Server.

  • Support containerised environments using Docker and Kubernetes/AKS, where required.

  • Identify opportunities to reduce manual intervention, improve resilience and increase operational efficiency through automation.

  • Collaborate closely with development, infrastructure, security and operational teams to embed reliability and operational best practices throughout the delivery lifecycle.


Key Deliverables

In this role, you will contribute to the delivery of:

  • Reliable, scalable and highly available production services and infrastructure.

  • Clearly defined and measurable SLIs and SLOs, supported by appropriate error budgets.

  • Effective monitoring, alerting and observability across critical applications and infrastructure.

  • Automated and repeatable infrastructure deployments using Terraform and/or Bicep.

  • Reliable and efficient CI/CD pipelines supporting application and infrastructure deployments.

  • Increased automation of operational processes, reducing manual effort and improving consistency.

  • Improved production stability through proactive monitoring, performance analysis and remediation.

  • Effective incident response, root-cause analysis and blameless post-mortems, with identified improvements followed through to completion.

  • Improved platform resilience, performance and scalability across the Azure environment.

  • Clear operational documentation, procedures and knowledge transfer to support ongoing service management.

  • Continuous improvements that strengthen reliability, operational efficiency and overall service performance.


Job Requirements

Essential

To be successful in this role, you will have:

  • Minimum 3 years' experience in SRE, DevOps, Platform Engineering, Infrastructure Engineering or a similar role supporting production environments.

  • Strong hands-on experience with Microsoft Azure, including compute, networking, App Service, storage and databases.

  • Demonstrated experience supporting and operating production services at scale.

  • Experience defining, implementing or working with SLIs, SLOs and error budgets.

  • Strong experience with Azure monitoring and observability tools, including Azure Monitor, Log Analytics and Application Insights.

  • Experience with Infrastructure as Code (IaC) using Terraform and/or Bicep.

  • Experience building and maintaining CI/CD pipelines using Azure DevOps and/or GitHub Actions.

  • Strong scripting and automation capability using PowerShell and Python, or a similar language such as Go.

  • Solid understanding of networking concepts including DNS, TCP/IP, load balancing and Network Security Groups (NSGs).

  • Experience administering or supporting Linux and/or Windows Server environments.

  • Practical experience with incident management, troubleshooting, root-cause analysis and post-incident reviews.

  • Strong problem-solving, communication and stakeholder engagement skills.

Desirable

The following skills and experience will be highly regarded:

  • Hands-on experience with Docker and containerised applications.

  • Experience with Kubernetes and Azure Kubernetes Service (AKS).

  • Experience with Prometheus and Grafana for monitoring and observability.

  • Understanding of distributed systems and highly available architectures.

  • Experience implementing or improving formal Site Reliability Engineering (SRE) practices.

  • Knowledge of ITIL principles and service management practices.

  • Experience working within Scrum and/or Kanban delivery environments.

  • Relevant Microsoft Azure, DevOps, Kubernetes or cloud certifications would be advantageous.


Personal attributes

To be successful in this role you will demonstrate:

  • A genuine passion for technology, automation, reliability and continuous improvement.

  • A proactive, solutions-focused approach to identifying and resolving complex technical issues.

  • Strong analytical and troubleshooting skills with the ability to remain methodical when responding to production incidents.

  • A strong sense of ownership and accountability for the reliability and performance of the environments you support.

  • The ability to work calmly and effectively when managing critical production issues and incidents.

  • A mindset focused on automation, with a desire to eliminate repetitive manual processes wherever possible.

  • Strong attention to detail and a commitment to building reliable, secure and scalable solutions.

  • The ability to work independently while contributing positively to collaborative engineering and operational teams.

  • Excellent communication skills with the ability to explain complex infrastructure, cloud and reliability concepts to both technical and non-technical stakeholders.

  • A continuous improvement mindset, with a willingness to challenge existing processes and identify better ways of working.

  • The ability to adapt to changing priorities and work effectively within dynamic production environments.

  • A professional, reliable and accountable approach to your work.

  • A collaborative and blameless approach to incident management, focused on learning, improvement and prevention rather than assigning fault.

  • A willingness to share knowledge, support colleagues and contribute to the ongoing development of the broader team.

  • A commitment to applying SRE, DevOps, security and cloud engineering best practices to improve platform reliability and operational performance.

  • A professional, reliable and accountable approach to your work.

  • A positive attitude and willingness to share knowledge, mentor others and contribute to team success.

  • A commitment to delivering solutions that are secure, scalable and aligned with best practice engineering principles.


Deal-breakers (stuff we don’t want in a new team member)

  • A “that’ll do” approach to our clients’ projects;

  • A closed mind to change;

  • Big egos. We’re really, really not keen on big egos here.


If this opportunity aligns with your experience, please apply via Seek with your updated CV.


*Only applicants with citizenship, permanent visa will be considered.

No agencies please!

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available