Senior Site Reliability Engineer (SRE/DevOps)

Summary

Senior Site Reliability Engineer designs and operates secure, scalable cloud infrastructure for AI systems, focusing on reliability, automation, and observability using Kubernetes, Terraform, and major cloud providers.

Senior Site Reliability Engineer (SRE/DevOps)

Location: Vietnam

Workplace Type: Remote


About the Role

We are hiring on behalf of our client, a growing technology company, for a Senior Site Reliability Engineer to help shape their infrastructure strategy, design resilient cloud architectures, and ensure their platforms are secure, scalable, and high-performing. This role is critical to bringing AI systems into production with reliable delivery, strong observability, and operational excellence.


Our client is specifically looking for someone with a genuine SRE mindset - not a traditional DevOps background. Candidates should have hands-on experience optimizing systems for reliability and building zero-downtime production systems, not just automating deployments.

Responsibilities

  • Design and operate secure, scalable, high-quality infrastructure supporting modern applications and AI workloads.
  • Build and maintain robust automation across CI/CD pipelines, infrastructure provisioning, and operational processes to improve reliability and minimize manual effort.
  • Integrate AI-driven solutions into operational workflows to enhance efficiency, detect anomalies, and accelerate delivery.
  • Apply strong systems engineering practices: monitoring, incident management, performance optimization, and capacity planning.
  • Establish and uphold SRE/DevOps best practices, including reproducibility, testing, documentation, and operational excellence.
  • Communicate technical decisions clearly and collaborate cross-functionally to support predictable delivery.
  • Provide mentorship and technical leadership, raising the bar on platform engineering and DevOps maturity across the organization.

Requirements

Must-Have

  • 6+ years of progressive experience in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering (SRE background strongly preferred over traditional DevOps).
  • Proven track record designing, building, and operating highly reliable, zero-downtime production systems, using patterns such as Blue/Green, Canary, progressive delivery, or Preview Environments.
  • Deep, hands-on expertise in Kubernetes (or equivalent container orchestration) running in production.
  • Strong experience with Infrastructure as Code (Terraform, Pulumi, or CloudFormation).
  • Solid experience with at least one major cloud provider (AWS, GCP, or Azure), including networking, compute, storage, and security.
  • Experience building CI/CD pipelines from the ground up (not just using pre-built templates).
  • Practical experience with modern observability stacks (e.g., Prometheus, Grafana, Datadog).
  • Some hands-on experience supporting or deploying AI/ML workloads (model inference, vector databases, or GPU workloads).
  • Strong background in platform security: secrets management, IAM, and runtime/security hardening.
  • Excellent communication skills, with a proven ability to explain complex infrastructure decisions and mentor other engineers.

Nice-to-Have

  • Experience with GitOps practices (ArgoCD, Flux).
  • True multi-cloud experience across AWS/GCP/Azure.
  • Experience with Multi-Cloud API Gateways and Edge Routing.
  • Experience building Self-Service Developer Platforms.
  • Familiarity with Node.js, NestJS, or Python for extending DevOps tooling.
  • Experience collaborating with QA/IT/ISRM teams on vulnerability remediation and incident investigation.

Benefits

  • Attractive salary range and open to negotiate for strong fits.
  • Hybrid/Remote-friendly culture. Work where you grow best!
  • Flexible hours, async teamwork. Focus time is respected.
  • Work equipment support.
  • Allowance for certification & skill development.
  • Year-end bonus & performance-based rewards.
  • 22 paid leaves from your 5th year. Take a full month off.
  • Career growth with personal coaching sessions.
  • Open, collaborative team culture. No micromanagement, only trust.
  • Tools & AI-powered workflows that make remote work easier.