Software & Applications Manager (Technical Lead/Supervisor)
Role Overview
The Observability Engineer leads the enterprise observability and service reliability strategy across infrastructure and application domains. This role drives proactive monitoring, automation, and resilience initiatives to enhance system availability, ensure regulatory compliance, and reduce mean time to resolution (MTTR).
Key Responsibilities
- Observability Strategy & Governance: Define enterprise observability architecture aligned with operational resilience standards. Deploy and optimise full-stack observability platforms (metrics, logs, traces) and integrate them with ITSM and AIOps systems for predictive alerting.
- Reliability Engineering & Automation: Implement SRE frameworks, define error budget policies, and codify operational reliability. Automate runbooks, self-healing workflows, and auto-remediation actions using Python, Ansible, and Terraform.
- Cloud & Platform Observability: Architect and manage telemetry solutions for cloud-native workloads across AWS and Azure, embedding observability into landing zones and CI/CD pipelines via Infrastructure-as-Code (IaC).
- Operational Excellence: Partner with cross-functional teams to conduct resilience testing, chaos engineering, and capacity validation. Maintain executive dashboards for operational risk indicators and act as a technical advisor during major incidents and audits.
Competencies and Qualifications
- Technical Background: Qualification in Computer Science, Information Technology, or equivalent practical experience.
- Domain Experience: Track record in Infrastructure, Cloud, or Site Reliability Engineering (SRE), with experience operating as an SRE subject matter expert, ideally within financial services or regulated environments.
- Observability Platforms: Hands-on proficiency with tools such as Datadog, Dynatrace, Splunk, or ELK Stack.
- Automation & Cloud: Expertise in Infrastructure-as-Code and scripting (Terraform, Ansible, Python) alongside cloud observability tooling (AWS CloudWatch/X-Ray, Azure Monitor/Log Analytics).
- SRE & Compliance: Deep understanding of SRE principles (SLOs/SLAs, error budgets), financial sector operational resilience frameworks (e.g., MAS TRM, DORA), and automated remediation.
- Certifications: Relevant certifications in observability platforms, IaC, cloud engineering (AWS/Azure), SRE, or ITIL are advantageous.