DevOps Engineer
Summary
A DevOps Engineer in Abu Dhabi who designs CI/CD pipelines, builds and maintains scalable cloud/on-prem infrastructure for AI model deployment, automates provisioning, and runs monitoring, security compliance, and on-call support. Core stack: Docker, Kubernetes, AWS/Azure/GCP, Terraform/Ansible, Prometheus/Grafana, Python/Bash/Go. Role is via recruitment firm GCS.
The Role
We are seeking a highly motivated and skilled DevOps Engineer to join our team. You'll play a
crucial role in building, deploying, and maintaining scalable and reliable systems and
infrastructure, working closely with development teams to ensure operational efficiency and
smooth deployment pipelines.
What You’ll Do
- Design, implement, and maintain CI/CD pipelines to streamline development workflows.
- Build and manage scalable infrastructure for AI model deployment and lifecycle management.
- Automate infrastructure provisioning and management using tools like Terraform,
Ansible, or CloudFormation. - Optimize cloud-based and on-premises resources for scalability and cost efficiency.
- Manage and fine-tune queuing systems and real-time streaming architectures.
- Monitor and troubleshoot production systems to ensure uptime and performance.
- Implement logging, monitoring, and alerting solutions using tools such as Prometheus,
Grafana, ELK stack, etc. - Set up comprehensive monitoring for both system metrics and ML model performance.
- Conduct root cause analyses and post-mortems to improve system reliability.
- Collaborate with development and QA teams to deploy new features into production seamlessly.
- Promote best practices in system architecture, security, and performance.
- Participate in a rotating on-call schedule for production system support.
- Ensure infrastructure complies with security and compliance standards (e.g., SOC2,
ISO27001). - Securely manage secrets and credentials using tools like Vault or AWS Secrets Manager.
What You’ll Bring
- Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience).
- Proficiency in at least one scripting language: Python, Bash, or Go.
- Hands‑on experience with cloud platforms like AWS, Azure, or Google Cloud.
- Skilled in containerization and orchestration with Docker and Kubernetes.
- Experience using CI/CD tools such as Azure DevOps, Jenkins, GitLab CI/CD, or CircleCI.
- Knowledge of monitoring and observability tools like Prometheus, Datadog, New Relic,
Grafana, or PagerDuty. - Understanding of networking fundamentals including DNS, load balancing, and firewalls.
- Familiarity with real-time streaming architectures for AI and data applications.
Great Pluses / Preferred Experience
- Experience with Infrastructure as Code (IaC) tools like Terraform or Pulumi.
- Understanding of service mesh technologies like Istio or Linkerd.
- Familiarity with database scaling and administration, including VectorDBs, SQL, and
NoSQL systems. - Previous experience in a high-traffic production environment
GCS is acting as an Employment Business in relation to this vacancy.
Skills
- AI
- Ansible
- AWS
- Azure
- Azure DevOps
- Bash
- CI/CD
- CircleCI
- Cloud
- CloudFormation
- Containerization
- Datadog
- DevOps
- DNS
- Docker
- ELK
- Firewall
- GCP
- GitLab
- Grafana
- Infrastructure as Code
- ISO 27001
- Istio
- Jenkins
- Kubernetes
- Machine Learning
- Model Deployment
- Networking
- New Relic
- NoSQL
- Observability
- PagerDuty
- Prometheus
- Pulumi
- Python
- SOC 2
- SQL
- Terraform
- Vault