Kubernetes & Cloud Operations Specialist
We are looking for a Kubernetes & Cloud Operations Specialist with strong experience in Kubernetes, monitoring, automation and production operations to join our client, a leading Systems Integrator (SI) in Singapore supporting large-scale technology projects. Candidates with relevant experience supporting public sector projects will be highly preferred.
Job Description
- Gather and analyze metrics from operating systems as well as applications to assist in performance tuning and fault finding
- Partner with various teams to improve services through testing, workflow and best practices
- Participate in system design consulting, platform management, and capacity planning, knowledge in Broadcom VMware is a plus
- Create sustainable systems and services through automation and uplifts
- Balance feature development speed and reliability with well-defined service-level objectives
- Proactive approach to identifying problems, performance bottlenecks, and areas for improvement through automation, adding of workflow etc.
- Collaborate with delivery teams, to delivery service and project on time and to requirement while focusing on operation work
- Understand SDLC and have some development knowledge
Requirements
- Bachelor’s Degree in Computer Science, IT, Engineering or a related field.
- 6–8+ years of experience in IT infrastructure, Cloud Operations, DevOps, SRE or related fields.
- Strong hands-on experience with Kubernetes and Docker in production environments.
- Experience with Linux/Windows server administration and infrastructure operations.
- Good understanding of CI/CD, monitoring, observability and automation.
- Programming/scripting experience in Python, Bash or similar languages.
- Strong troubleshooting, analytical and problem-solving skills.
- Good communication skills and ability to work independently and collaboratively.
- Experience supporting enterprise or public sector projects is an advantage.
Technical Skills
- OS: RHEL, CentOS, Ubuntu, Windows Server
- Containerisation: Docker, Kubernetes
- Monitoring & Observability: Prometheus, Grafana, ELK/OpenSearch or similar
- DevOps & Automation: CI/CD, scripting, Infrastructure as Code
- Infrastructure: VMware/Broadcom and/or AWS, Azure or GCP