Senior Platform Engineer (SRE)
Team Segment : Music
KKBOX is one of the leading music streaming platforms in Asia, serving millions of users across multiple regions. The platform handles large-scale traffic and concurrent streaming workloads.
Our infrastructure spans both on-premise data centers and cloud environments, leveraging global CDN networks and multi-region architecture to deliver highly available and high-performance services.
We are looking for an experienced Senior Platform Engineer (SRE focus) to build and maintain reliable, scalable, and highly available platforms. You will work closely with Engineering, Product, Data, and Infrastructure teams to improve development efficiency, system reliability, and operational excellence.
Responsibilities:
Cloud Infrastructure & Platform Management
Design, build, and maintain cloud infrastructure on AWS.
Manage Kubernetes clusters and containerized environments.
Build highly available and scalable platform architectures.
Support hybrid cloud infrastructure planning and operations.
Automation & CI/CD
Design and maintain CI/CD pipelines.
Implement Infrastructure as Code (IaC) practices.
Automate deployment, testing, and operational workflows.
Continuously improve development and delivery efficiency.
Reliability & Observability
Build and maintain monitoring, alerting, and observability platforms.
Define and manage SLI/SLO metrics.
Participate in on-call rotations and respond to production incidents in real time.
Define and manage Error Budgets, and assess release risk and cadence based on Error Budget policy.
Participate in incident response and root cause analysis.
Perform capacity planning and performance optimization.
Database Operations
Maintain and optimize MySQL and PostgreSQL databases.
Implement backup, restore, and disaster recovery solutions.
Perform database performance tuning and troubleshooting.
Support database upgrades and migration projects.
Security & Compliance
Manage IAM, access control, and credential management.
Implement secret management solutions.
Enhance infrastructure and container security.
Support compliance and audit requirements.
Requirements:
Technical Skills
5+ years of experience in SRE or Platform Engineering.
Strong knowledge of AWS cloud services.
Hands-on experience with Docker and Kubernetes.
Solid Linux administration and troubleshooting skills.
Experience with Infrastructure as Code (Terraform preferred).
Experience with production monitoring, alerting, and on-call incident response.
Soft Skills
Strong analytical and problem-solving skills.
Excellent communication and collaboration abilities.
Ability to independently drive technical initiatives.
Nice to Have:
Technical Skills
GitOps (ArgoCD)
Redis
OpenSearch / Elasticsearch
Prometheus / Grafana
Experience with large-scale streaming services
Multi-cloud environments
FinOps experience
Experience designing and maintaining CI/CD pipelines.
Experience operating MySQL or PostgreSQL databases.
Experience with physical data center operations
Certifications
AWS Certified Solutions Architect
Kubernetes Certifications (CKA / CKAD)