Senior DevOps & Site Reliability Engineer

Summary

The Senior DevOps & Site Reliability Engineer will design, build, and maintain highly available, scalable, and automated enterprise platforms on Microsoft Azure. The role focuses on implementing CI/CD pipelines, managing Kubernetes environments, and driving reliability through observability and infrastructure-as-code practices.

We are looking for an experienced Senior DevOps & Site Reliability Engineer (SRE) to design, build and operate highly available, secure, scalable and automated enterprise technology platforms.

This is a senior hands-on engineering role spanning DevOps, Site Reliability Engineering, Azure Cloud, Platform Engineering, Kubernetes, Infrastructure as Code, CI/CD, Observability and DevSecOps.

The successful candidate will work across engineering and delivery teams to improve platform reliability, deployment velocity, resilience, automation, operational efficiency and production performance, while supporting mission-critical enterprise applications.

Key Responsibilities

DevOps & Platform Engineering

  • Design, build and maintain cloud-native infrastructure and platform services.

  • Develop and maintain Infrastructure as Code (IaC) solutions.

  • Automate infrastructure provisioning, configuration and operational processes.

  • Build reusable engineering tools, deployment templates and platform components.

  • Establish and standardise platform engineering practices across multiple delivery teams.

  • Identify opportunities to reduce manual intervention and increase engineering automation.

CI/CD & Release Automation

  • Design, implement and maintain enterprise-grade CI/CD pipelines for application and infrastructure deployments.

  • Implement automated testing, security scanning, code-quality controls and release automation.

  • Enable automated deployments, rollback and recovery processes.

  • Improve deployment frequency while reducing change and deployment risk.

  • Continuously optimise software delivery and release-management processes.

Site Reliability Engineering

  • Implement and mature Site Reliability Engineering practices across production environments.

  • Define, monitor and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs) and Service Level Agreements (SLAs).

  • Improve application and platform availability, scalability, resilience and performance.

  • Lead production incident response, troubleshooting, problem management and Root Cause Analysis (RCA).

  • Drive proactive reliability improvements and reduction of technical debt.

  • Improve Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).

Azure Cloud Engineering

  • Design, implement and operate enterprise Microsoft Azure environments.

  • Work extensively with technologies such as:

    • Azure Kubernetes Service (AKS)

    • Azure App Services

    • Azure Networking

    • Azure Monitor

    • Azure Storage

    • Azure Identity Services

  • Design highly available and disaster-recovery-capable environments.

  • Optimise cloud environments for performance, resilience, security and cost.

  • Support hybrid-cloud and multi-cloud environments where required.

Containers & Kubernetes

  • Build, deploy and support containerised applications using Docker and Kubernetes.

  • Manage Kubernetes environments, particularly Azure Kubernetes Service (AKS).

  • Develop and maintain deployment configurations using Helm.

  • Support container-platform reliability, scalability and operational performance.

  • OpenShift experience would be advantageous.

Infrastructure as Code & Automation

Hands-on experience with technologies such as:

  • Terraform

  • Bicep

  • ARM Templates

  • Ansible

Candidates should be comfortable using Infrastructure as Code to build repeatable, scalable and governed enterprise infrastructure.

Monitoring & Observability

  • Implement comprehensive logging, monitoring, metrics, tracing and alerting.

  • Build operational dashboards and platform insights.

  • Establish enterprise observability standards.

  • Implement proactive and predictive monitoring.

  • Use observability information to improve application and infrastructure reliability.

Relevant technologies may include:

  • Dynatrace

  • Grafana

  • Prometheus

  • Elastic Stack / ELK

See also

DevOps jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available