DevOps Engineer
Summary
DevOps engineer (5-10 yrs) building and running infrastructure for an early-stage Sovereign AI platform in Kuala Lumpur: infrastructure-as-code with Terraform, Kubernetes and Docker, CI/CD pipelines, monitoring/observability and security hardening across AWS/GCP and on-prem, including GPU and AI/ML workloads.
- Design, build and maintain infrastructure-as-code for our Sovereign AI platform using Terraform, Kubernetes and Docker.
- Define and establish processes and standards for infrastructure provisioning, deployment and security — as an early team member, you'll be shaping how things get done, not just following an existing playbook.
- Build and maintain CI/CD pipelines to support fast, reliable deployment of infrastructure and platform changes.
- Operate and improve the reliability, scalability and performance of our cloud and on-prem infrastructure.
- Implement monitoring, logging and alerting to support platform observability and operational readiness.
- Apply security best practices across the infrastructure lifecycle — from image/container hardening and vulnerability scanning to access control and secrets management.
- Support incident response and troubleshooting across infrastructure and deployment issues.
- Work closely with the Head of Sovereign AI Infra and the rest of the team to continuously improve automation, tooling and platform maturity.
- Contribute to securing AI/ML and high-performance computing (HPC) workloads as the platform evolves, including GPU infrastructure and model-serving environments.
- 5–10 years of experience in a DevOps, Platform Engineering, SRE or DevSecOps role.
- Strong, hands-on Linux skills — comfortable operating, troubleshooting and optimising at the OS and systems level.
- Strong command of the open-source infrastructure ecosystem, with hands‑on experience in Terraform, Kubernetes and Docker.
- Hands‑on experience with AWS or GCP in a production environment.
- Experience building and maintaining CI/CD pipelines.
- Practical understanding of security fundamentals — vulnerability management, container/image hardening, access control, secrets management.
- Comfortable working in a small, fast‑moving team where you're building and operating, not just designing.
- Demonstrated ability to define and introduce processes, standards or best practices from scratch, ideally in a startup or early‑stage environment.
Highly Regarded
- Experience securing or deploying generative AI, AI/ML or high‑performance computing (HPC) workloads.
- Experience with GPU infrastructure and related tooling.
- Contribution to open‑source infrastructure or security tooling projects.
- Experience with observability stacks (e.g. Prometheus, Grafana), open source SIEM (Wazuh) and security scanning tools (e.g. Trivy, Snyk).
- Prior experience in an infrastructure‑as‑a‑service or platform‑as‑a‑service environment.