Site Reliability Engineer
Summary
Maintain and automate cloud infrastructure for a global restaurant-tech platform, using AWS, Python/Bash/Go, and observability tools like Datadog and Prometheus.
Site Reliability Engineer II
Location: Ho Chi Minh, Dong Nam Bo, Viet Nam; Hybrid
2+ years of experience in SRE, DevOps, production support, or infrastructure engineering roles
Hands-on experience with monitoring and observability tooling (e.g., Datadog, Prometheus, Grafana, CloudWatch, or similar)
Working knowledge of at least one major cloud provider (AWS preferred)
Proficiency in at least one scripting or programming language (e.g., Python, Bash, Go) for automation, with demonstrated examples of automating away manual operational work
Experience participating in incident response and on-call or shift-based operations
Understanding of SLI/SLO concepts and reliability engineering fundamentals
Ability to work follow-the-sun shift rotations, including structured handoffs with teams in other regions
Strong written and verbal English communication skills for cross-region collaboration
• 2+ years of experience in site reliability engineering, DevOps, infrastructure, or production operations roles. • Hands-on experience with monitoring and observability tooling (e.g., Datadog, Prometheus, Grafana, CloudWatch, or similar) • Working knowledge of at least one major cloud provider (AWS preferred) • Proficiency in at least one scripting or programming language (e.g., Python, Bash, Go) for automation, with demonstrated examples of automating away manual operational work • Experience participating in incident response and on-call or shift-based operations • Understanding of SLI/SLO concepts and reliability engineering fundamentals • Ability to work follow-the-sun shift rotations, including structured handoffs with teams in other regions • Strong written and verbal English communication skills for cross-region collaboration