Senior devops & site reliability engineer at datonomy solutions
We are looking for an experienced Senior Dev Ops & Site Reliability Engineer (SRE) to design, build and operate highly available, secure, scalable and automated enterprise technology platforms.
This is a senior hands‑on engineering role spanning Dev Ops, Site Reliability Engineering, Azure Cloud, Platform Engineering, Kubernetes, Infrastructure as Code, CI/CD, Observability and Dev Sec Ops.
The successful candidate will work across engineering and delivery teams to improve platform reliability, deployment velocity, resilience, automation, operational efficiency and production performance, while supporting mission‑critical enterprise applications.
Key Responsibilities Dev Ops & Platform Engineering Design, build and maintain cloud-native infrastructure and platform services. Develop and maintain Infrastructure as Code (Ia C) solutions. Automate infrastructure provisioning, configuration and operational processes. Build reusable engineering tools, deployment templates and platform components. Establish and standardise platform engineering practices across multiple delivery teams. Identify opportunities to reduce manual intervention and increase engineering automation. CI/CD & Release Automation Design, implement and maintain enterprise- grade CI/CD pipelines for application and infrastructure deployments. Implement automated testing, security scanning, code-quality controls and release automation. Enable automated deployments, rollback and recovery processes. Improve deployment frequency while reducing change and deployment risk. Continuously optimise software delivery and release-management processes. Site Reliability Engineering Implement and mature Site Reliability Engineering practices across production environments. Define, monitor and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs) and Service Level Agreements (SLAs). Improve application and platform availability, scalability, resilience and performance. Lead production incident response, troubleshooting, problem management and Root Cause Analysis (RCA). Drive proactive reliability improvements and reduction of technical debt. Improve Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR). Azure Cloud EngineeringDesign, implement and operate enterprise Microsoft Azure environments.
Work extensively with technologies such as:
Azure Kubernetes Service (AKS) Azure App Services o Azure Networking o Azure Monitor o Azure Storage o Azure Identity Services Design highly available and disaster-recovery-capable environments. Optimise cloud environments for performance, resilience, security and cost. Support hybrid-cloud and multi-cloud environments where required. Containers & Kubernetes Build, deploy and support containerised applications using Docker and Kubernetes. Manage Kubernetes environments, particularly Azure Kubernetes Service (AKS). Develop and maintain deployment configurations using Helm. Support container-platform reliability, scalability and operational performance. Open Shift experience would be advantageous. Infrastructure as Code & Automation Hands‑on Experience With Technologies Such As Terraform Bicep ARM Templates Ansible Monitoring & Observability Implement comprehensive logging, monitoring, metrics, tracing and alerting. Build operational dashboards and platform insights. Establish enterprise observability standards. Implement proactive and predictive monitoring. Use observability information to improve application and infrastructure reliability. Relevant Technologies May Include Dynatrace Grafana Prometheus Elastic Stack / ELK Splunk Azure Monitor Open Telemetry Dev Sec Ops & Security Embed Dev Sec Ops practices throughout the software-delivery lifecycle. Integrate security scanning and controls into CI/CD pipelines. Support vulnerability identification, remediation and risk reduction. Ensure cloud and platform environments comply with enterprise security and regulatory requirements. Work closely with information-security teams to continuously improve platform security. Technical Leadership Provide technical leadership across Dev Ops, Cloud, Platform and SRE teams. Mentor and coach junior and intermediate engineers. Contribute to architecture decisions and technology roadmaps. Promote engineering standards and operational best practice. Lead cross-functional initiatives aimed at improving engineering productivity and reliability. Minimum Experience 8+ years' experience across software engineering, infrastructure engineering, cloud engineering, Dev Ops or platform engineering. 5+ years' hands-on Dev Ops engineering experience. 3+ years' Site Reliability Engineering or production-operations experience. Proven experience supporting mission-critical production systems. Experience operating large-scale enterprise technology platforms. Strong exposure to highly available and business-critical environments. Essential Technical Skills Cloud Microsoft Azure Azure Kubernetes Service (AKS) Azure Networking Azure App Services Azure Monitor Azure Storage Azure Identity Dev Ops / CI/CD Azure Dev Ops Git Hub / Git Jenkins Sonar Qube Artifactory and/or Nexus Infrastructure Automation Terraform Bicep ARM Templates Ansible Containers Kubernetes Docker Helm Observability Dynatrace Grafana Prometheus Elastic Stack Splunk Azure Monitor Open Telemetry Scripting / Development Strong Scripting Or Programming Ability Using Technologies Such As Python Power Shell Bash C# JavaGo experience would be advantageous.
Core Technical Competencies Dev Ops Engineering Site Reliability Engineering Azure Cloud Engineering Platform Engineering Infrastructure as Code Kubernetes / Container Orchestration CI/CD Infrastructure Automation Dev Sec Ops Cloud Architecture Observability Continuous Delivery Systems Integration Capacity Planning Performance Optimisation Incident & Problem Management Root Cause Analysis Behavioural Competencies Strong technical problem-solving ability Strategic thinking Strong decision-making skills Collaboration across engineering disciplines Stakeholder management Continuous-improvement mindset Coaching and mentoring capability Accountability and ownership Customer-centric approach QualificationsA Bachelor's Degree or equivalent technical qualification in one of the following areas is preferred:
Computer Science Information Technology Software Engineering Information Systems Preferred CertificationsRelevant certifications would be advantageous, including:
Microsoft Certified: Azure Dev Ops Engineer Expert Microsoft Certified: Azure Solutions Architect Expert Certified Kubernetes Administrator (CKA) Certified Kubernetes Application Developer (CKAD) Hashi Corp Terraform Associate AWS Certified Dev Ops Engineer ITIL Foundation SRE Foundation Certification Ideal CandidateThe ideal candidate is a senior, hands‑on engineer who can bridge software development, cloud infrastructure, Dev Ops, platform engineering and production operations.
They should have deep experience building and running highly available enterprise environments and possess a strong automation‑first and reliability‑focused mindset.
This person should be equally comfortable troubleshooting a critical production issue, building Terraform infrastructure, improving a Kubernetes platform, designing a CI/CD pipeline, implementing observability, defining SLOs and mentoring other engineers.
Desired Skills Senior Dev Ops Engineer Site Reliability Engineer SRE