DevOps / SRE Engineer AI Cloud
Summary
Operates and supports an AI cloud/GPU infrastructure platform day to day: keeping Linux and Kubernetes environments healthy, maintaining monitoring and observability (Prometheus, Grafana, OpenSearch), handling security tasks like CVE remediation and access audits, and automating operations with Python and Shell.
We are hiring for a fast-growing global technology company headquartered in Singapore, currently expanding its AI Cloud and GPU infrastructure platform.
What you'll do:
- Support day-to-day operations of Linux and Kubernetes environments
- Maintain CMDB / asset management processes and perform routine platform health checks
- Operate internal platforms such as OpenBao / Vault, Sonatype and audit systems
- Maintain monitoring and observability platforms including Prometheus, Grafana, OpenSearch and Alert manager
- Support CVE remediation, patching, access control, account audits and security compliance activities
- Support K8s-based AI workloads including monitoring, RBAC, backup and resource management
- Automate operational tasks using Python / Shell
What We're Looking For
- 3+ years of experience in DevOps, SRE, Cloud Operations, or System Engineering
- Hands-on experience with Linux and Kubernetes, including deployment, scaling, monitoring, and basic troubleshooting
- Experience with Prometheus / Grafana or similar monitoring and observability tools
- Good scripting skills in Python and/or Shell
- Basic understanding of CI/CD, networking, RBAC, and infrastructure security
- Strong troubleshooting skills and a good sense of ownership and operational discipline
- Experience with OpenBao / Vault, CMDB, CVE remediation, or SOC2 is a plus