Senior devops & site reliability engineer

We are looking for an experiencedSenior Dev Ops & Site Reliability Engineer (SRE) to design, build and operate highly available, secure, scalable and automated enterprise technology platforms.

This is a senior hands-on engineering role spanningDev Ops, Site Reliability Engineering, Azure Cloud, Platform Engineering, Kubernetes, Infrastructure as Code, CI/CD, Observability and Dev Sec Ops .

The successful candidate will work across engineering and delivery teams to improveplatform reliability, deployment velocity, resilience, automation, operational efficiency and production performance , while supporting mission-critical enterprise applications. Key Responsibilities Dev Ops & Platform Engineering

Design, build and maintaincloud-native infrastructure and platform services .

Develop and maintainInfrastructure as Code (Ia C) solutions.

Automate infrastructure provisioning, configuration and operational processes.

Build reusable engineering tools, deployment templates and platform components.

Establish and standardise platform engineering practices across multiple delivery teams.

Identify opportunities to reduce manual intervention and increase engineering automation.

CI/CD & Release Automation

Design, implement and maintain enterprise-gradeCI/CD pipelines for application and infrastructure deployments.

Implement automated testing, security scanning, code-quality controls and release automation.

Enable automated deployments, rollback and recovery processes.

Improve deployment frequency while reducing change and deployment risk.

Continuously optimise software delivery and release-management processes.

Site Reliability Engineering

Implement and matureSite Reliability Engineering practices across production environments.

Define, monitor and manageService Level Indicators (SLIs), Service Level Objectives (SLOs) and Service Level Agreements (SLAs) .

Improve application and platformavailability, scalability, resilience and performance .

Lead production incident response, troubleshooting, problem management andRoot Cause Analysis (RCA) .

Drive proactive reliability improvements and reduction of technical debt.

ImproveMean Time to Detect (MTTD) andMean Time to Recover (MTTR) . Azure Cloud Engineering

Design, implement and operate enterpriseMicrosoft Azure environments.

Work extensively with technologies such as:

Azure Kubernetes Service (AKS )

Azure App Services

Azure Networking

Azure Monitor

Azure Storage

Azure Identity Services

Design highly available and disaster-recovery-capable environments.

Optimise cloud environments forperformance, resilience, security and cost .

Support hybrid-cloud and multi-cloud environments where required.

Containers & Kubernetes

Build, deploy and support containerised applications usingDocker andKubernetes .

Manage Kubernetes environments, particularlyAzure Kubernetes Service (AKS) .

Develop and maintain deployment configurations usingHelm .

Support container-platform reliability, scalability and operational performance.

Open Shift experience would be advantageous.

Infrastructure as Code & Automation

Hands-on experience with technologies such as:

Terraform

Bicep

ARM Templates

Ansible

Candidates should be comfortable using Infrastructure as Code to build repeatable, scalable and governed enterprise infrastructure.

Monitoring & Observability

Implement comprehensivelogging, monitoring, metrics, tracing and alerting .

Build operational dashboards and platform insights.

Establish enterprise observability standards.

Implement proactive and predictive monitoring.

Use observability information to improve application and infrastructure reliability.

Relevant technologies may include:

Dynatrace

Grafana

Prometheus

Elastic Stack / ELK

Splunk

Azure Monitor

Open Telemetry Dev Sec Ops & Security

EmbedDev Sec Ops practices throughout the software-delivery lifecycle.

Integrate security scanning and controls into CI/CD pipelines.

Support vulnerability identification, remediation and risk reduction.

Ensure cloud and platform environments comply with enterprise security and regulatory requirements.

Work closely with information-security teams to continuously improve platform security.

Technical Leadership

Provide technical leadership across Dev Ops, Cloud, Platform and SRE teams.

Mentor and coach junior and intermediate engineers.

Contribute to architecture decisions and technology roadmaps.

Promote engineering standards and operational best practice.

Lead cross-functional initiatives aimed at improving engineering productivity and reliability.

Minimum Experience

8+ years' experience across software engineering, infrastructure engineering, cloud engineering, Dev Ops or platform engineering.

5+ years' hands-on Dev Ops engineering experience.

3+ years' Site Reliability Engineering or production-operations experience.

Proven experience supportingmission-critical production systems .

Experience operatinglarge-scale enterprise technology platforms .

Strong exposure to highly available and business-critical environments.

Essential Technical Skills Cloud

Microsoft Azure

Azure Kubernetes Service (AKS)

Azure Networking

Azure App Services

Azure Monitor

Azure Storage

Azure Identity

Dev Ops / CI/CD

Azure Dev Ops

Git Hub / Git

Jenkins

Sonar Qube

Artifactory and/or Nexus

Infrastructure Automation

Terraform

Bicep

ARM Templates

Ansible

Containers

Kubernetes

Docker

Helm Observability

Dynatrace

Grafana

Prometheus

Elastic Stack

Splunk

Azure Monitor

Open Telemetry

Scripting / Development

Strong scripting or programming ability using technologies such as:

Python

Power Shell

Bash

C#

Java

Go experience would be advantageous.

Core Technical Competencies

Dev Ops Engineering

Site Reliability Engineering

Azure Cloud Engineering

Platform Engineering

Infrastructure Automation

Kubernetes / Container Orchestration

CI/CD

Infrastructure Automation

Dev Sec Ops

Cloud Architecture

Observability

Continuous Delivery

Systems Integration

Capacity Planning

Performance Optimisation

Incident & Problem Management

Root Cause Analysis

Behavioural Competencies

Strong technical problem-solving ability

Strategic thinking

Strong decision-making skills

Collaboration across engineering disciplines

Stakeholder management

Continuous-improvement mindset

Coaching and mentoring capability

Accountability and ownership

Customer-centric approach

Qualifications

A Bachelor's Degree or equivalent technical qualification in one of the following areas is preferred:

Computer Science

Information Technology

Software Engineering

Information Systems

Preferred Certifications

Relevant certifications would be advantageous, including:

Microsoft Certified:Azure Dev Ops Engineer Expert

Microsoft Certified:Azure Solutions Architect Expert

Certified Kubernetes Administrator (CKA)

Certified Kubernetes Application Developer (CKAD)

Hashi Corp Terraform Associate

AWS Certified Dev Ops Engineer

ITIL Foundation

SRE Foundation Certification

Ideal Candidate

The ideal candidate is a senior, hands-on engineer who can bridgesoftware development, cloud infrastructure, Dev Ops, platform engineering and production operations .

They should have deep experience building and running highly available enterprise environments and possess a strongautomation-first and reliability-focused mindset .

This person should be equally comfortable troubleshooting a critical production issue, building Terraform infrastructure, improving a Kubernetes platform, designing a CI/CD pipeline, implementing observability, defining SLOs and mentoring other engineers. #J-18808-Ljbffr

See also

DevOps jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available