Site Reliability Engineer (SRE) – DevOps & Cloud Engineering | 5+ Years | Abu Dhabi, UAE
Summary
An experienced Site Reliability Engineer in Abu Dhabi owns and supports customer-facing production systems on AWS and Kubernetes (EKS), building Python automation, Terraform/CDK infrastructure-as-code, and CI/CD pipelines. Day-to-day work spans observability (CloudWatch, X-Ray), incident management and on-call, Kubernetes scaling/upgrades, and production troubleshooting.
Location
Abu Dhabi, United Arab Emirates — Local Candidates Only
Job Category
- Information Technology (IT) & Software
- Engineering & Technical
- Telecommunications
- Remote / Work-from-Home Opportunities
Job Overview
An opportunity is available for an experienced Site Reliability Engineer (SRE) / Production Engineer in Abu Dhabi, UAE. The role focuses on owning and supporting customer-facing production systems with strong hands-on expertise in AWS, Amazon EKS, Kubernetes, Python, automation, and cloud infrastructure.
The successful candidate will be responsible for production reliability, cloud infrastructure, observability, incident management, automation, and continuous improvement. Strong experience with Kubernetes scaling, AWS services, Infrastructure as Code, CI/CD, and production troubleshooting is required. The position offers opportunities for professional development, cloud and DevOps training, relevant certification, and long-term career growth in site reliability and cloud engineering.
Key Responsibilities
- Own and support customer-facing production systems.
- Develop Python-based automation and engineering solutions.
- Manage AWS and Amazon EKS/Kubernetes production environments.
- Configure and optimize HPA, Karpenter/Cluster Autoscaler, networking, ingress, storage, and RBAC.
- Perform Kubernetes cluster upgrades and production maintenance.
- Manage AWS Lambda and serverless architectures.
- Administer IAM/IRSA, VPC, S3, ECR, API Gateway, SQS, SNS, EventBridge, and managed databases.
- Manage Helm and Kustomize deployments.
- Implement GitOps and Infrastructure as Code using Terraform, CDK, or CloudFormation.
- Support CI/CD pipelines and deployment automation.
- Monitor production systems using CloudWatch, X-Ray, and CloudTrail.
- Participate in incident management, on-call support, postmortems, and production troubleshooting.
- Define and monitor SLOs and reliability objectives.
- Troubleshoot Linux, containers, networking, and distributed systems.
- Work with stakeholders and technical teams to improve system reliability and performance.
Requirements & Qualifications
- 5+ years of experience in SRE, DevOps, Production Engineering, or Software Engineering.
- Strong Python development and automation experience.
- Deep hands-on AWS and Amazon EKS/Kubernetes production experience.
- Strong knowledge of HPA, Karpenter/Cluster Autoscaler, networking, ingress, storage, RBAC, and Kubernetes upgrades.
- Production experience with AWS Lambda and serverless architectures.
- Strong knowledge of IAM/IRSA, VPC, S3, ECR, API Gateway, SQS/SNS/EventBridge, and managed databases.
- Experience with Helm, Kustomize, GitOps, Terraform, CDK, CloudFormation, and CI/CD.
- Strong AWS observability experience with CloudWatch, X-Ray, and CloudTrail.
- Experience with incident management, on-call operations, postmortems, SLOs, and production troubleshooting.
- Strong Linux, container, networking, and distributed-systems fundamentals.
- Excellent communication and stakeholder-management skills.
- Bachelor’s degree in Computer Science, Engineering, or a related field.
- Candidates must be currently based locally in Abu Dhabi/UAE.
Salary, Benefits & Career Growth
Salary and compensation details were not provided by the employer and therefore are not listed.
The role offers opportunities for:
- Professional development in SRE, DevOps, and cloud engineering.
- Advanced AWS and Kubernetes technical exposure.
- Cloud infrastructure and automation training.
- Relevant AWS, Kubernetes, DevOps, and cloud certification opportunities.
- Experience managing large-scale production environments.
- Career progression in Site Reliability Engineering, DevOps, and Cloud Engineering.