Head of Site Reliability Engineering (SRE) & Information Security
Summary
Leads cloud infrastructure, DevOps, and security for a government-backed platform, designing AWS/Kubernetes environments with IaC, CI/CD, and observability while enforcing compliance and automation.
About the Role
We are looking for an experienced Head of Site Reliability Engineering (SRE) & Information Security to lead our cloud infrastructure, DevOps, security, and reliability initiatives. The ideal candidate will have extensive experience designing and managing secure, highly available, cloud-native platforms with expertise in Kubernetes, AWS, Infrastructure as Code, CI/CD automation, and cloud security.
Key Responsibilities
· Design, implement, and manage highly available AWS cloud infrastructure.
· Lead Kubernetes (EKS) platform engineering and container orchestration.
· Build and maintain CI/CD pipelines using Jenkins, GitLab, ArgoCD, and Git workflows.
· Implement Infrastructure as Code (Terraform) for provisioning and disaster recovery.
· Drive Site Reliability Engineering practices, including monitoring, alerting, incident management, and capacity planning.
· Strengthen cloud security through IAM, WAF, encryption, vulnerability management, and security monitoring.
· Develop disaster recovery and business continuity solutions with defined RTO/RPO objectives.
· Drive observability using Prometheus, Grafana, CloudWatch, and SigNoz.
· Optimise cloud costs through FinOps best practices.
· Collaborate with engineering teams to improve deployment automation and operational excellence.
· Lead internal/external security audits and compliance initiatives.
· Evaluate and implement modern DevSecOps technologies.
Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- 9 - 11 years of experience in DevOps, Cloud Infrastructure, Site Reliability Engineering (SRE), or Platform Engineering.
- AWS and/or Kubernetes certifications are preferred.
Skills & Competencies
- Strong experience with AWS (EC2, EKS, VPC, Route 53, RDS, Aurora, S3, EFS, ALB/NLB) and cloud infrastructure design.
- Hands-on expertise in Kubernetes, Docker, Helm, Kustomize, Terraform, and CI/CD tools such as Jenkins, GitLab CI/CD, and ArgoCD.
- Experience with monitoring, observability, and reliability engineering using Prometheus, Grafana, CloudWatch, and SigNoz.
- Solid understanding of cloud security practices, including GuardDuty, Security Hub, Cloudflare WAF, IAM, and security compliance.
- Proficiency in Python and Shell scripting for automation and infrastructure management.
- Experience with PostgreSQL, Aurora PostgreSQL, Redis, and Solr; exposure to Kafka, AWS AI Services, DevSecOps, FinOps, and security audits is an advantage.
- Strong problem-solving, leadership, stakeholder management, and communication skills with the ability to lead large-scale production environments.