DevOps / Site Reliability Engineer
Summary
DevOps/Site Reliability Engineer in Hong Kong who keeps production infrastructure and services stable, performant and reliable — managing Linux servers, virtualization and Kubernetes/Docker workloads, building HA, DR and auto-scaling, and maintaining CI/CD pipelines, monitoring (Prometheus/Grafana) and cloud infrastructure.
Ensure the stability, performance and reliability of production infrastructure and services. Manage Linux servers, virtualization platforms, Kubernetes clusters and containerized applications.
Key responsibilities
Manage Linux servers, virtualization platforms, Kubernetes clusters and containerized applications
Design and improve high availability, disaster recovery, auto-scaling and resource management solutions
Perform system monitoring, performance tuning, troubleshooting and capacity planning
Develop automation tools and improve deployment, monitoring, alerting and incident response processes
Maintain and optimize CI/CD pipelines and cloud infrastructure
About you
Strong Linux administration and troubleshooting skills
Proficient in Bash, Python or Go
Hands-on experience with Kubernetes and Docker in production environments
Familiar with KVM/VMware, Ansible/Terraform and CI/CD tools such as Jenkins, GitLab CI or ArgoCD
Experience with Prometheus, Grafana and monitoring systems
Familiar with Nginx, HAProxy, Kafka or RabbitMQ
Experience with AWS, Alibaba Cloud or Tencent Cloud is a plus
Strong problem-solving and automation mindset