Devops Engineer or Software Engineer or System Analyst
Summary
Operate and maintain enterprise-scale Kubernetes-based systems in Hong Kong, implementing SRE practices, managing incidents, reviewing vendor deliverables, and automating operational tasks using Bash, Python and Ansible.
Responsibilities
- Own the day-to-day operation, maintenance and support of designated enterprisescale systems, including standalone and Kubernetes-based environments.
- Administer and support Kubernetes environments, including application deployment, configuration, monitoring, troubleshooting and operational maintenance.
- Review vendors’ deliverables, including technical documents, source code, test evidence and deployment materials, to ensure compliance with contractual requirements, internal security standards and guidelines.
- Coordinate with business users, vendors, platform and infrastructure teams to resolve incidents, deliver system changes and support production releases.
- Implement and enhance Site Reliability Engineering (SRE) practices, including monitoring, alerting, logging and incident management, to ensure system availability, reliability and performance.
- Investigate production incidents, perform root-cause analysis, and drive corrective and preventive actions.
- Identify repetitive manual operational tasks and improve efficiency, consistency and reliability through scripting, standardisation and workflow automation.
- Maintain system operational documentation, including runbooks, support procedures, incident reports and knowledge-base materials.
- Degree in Computer Science, Information Technology or a related discipline.
- At least 3 years of IT experience, including 2 years of relevant experience in system development, application implementation and/or enterprise-scale system support.
- Demonstrates hands-on experience in Kubernetes (K8s) operations, including application deployment, configuration management, troubleshooting, monitoring, and incident resolution. Administrative experience with Kubernetes is advantageous.
- Familiarity with container technologies, such as Docker, and Kubernetes deployment/configuration artifacts.
- Proven ability to coordinate cross-functional teams to drive delivery and issue resolution.
- Familiar with Agile/Scrum methodologies and collaboration tools such as Jira and Confluence.
- Solid understanding of modern application concepts, including but not limited to RESTful APIs, Single Sign-On (SSO) and microservices architecture.
- Good scripting skills in Bash and Python.
- Strong automation mindset, with the ability to identify and transform manual operational processes into standardised, automated workflows.
- Hands-on experience with Ansible Playbooks for operational, deployment or configuration automation is highly preferred.
- Experience with monitoring, alerting and observability tools is preferred.
- Strong analytical, problem-solving and incident-management skills.
- Excellent communication skills, with proficiency in written and spoken English, Cantonese and/or Mandarin