Manager, Cloud SysOps & SRE Automation VN

Key Responsibilities
  • Design, build, and operate scalable, secure, and highly available application platforms across GCP and on-premises environments.

  • Site Reliability Engineering initiatives through Infrastructure as Code (IaC), automation, cloud-native engineering, and reliability best practices.

  • Design and develop automation frameworks and engineering platforms for infrastructure provisioning, deployment orchestration, scaling, observability, and operational reliability using tools such as Terraform, Ansible, Helm.

  • Optimize Kubernetes platforms, cloud architecture, deployment workflows, and operational processes to improve scalability, reliability, performance, operational efficiency, and cost optimization.

  • Define and manage SLIs/SLOs, observability standards, reliability metrics, telemetry, alerting strategies, and operational governance to ensure production stability and service reliability.

  • Design and enhance enterprise observability capabilities including metrics, logging, tracing, monitoring dashboards, and incident visibility using Prometheus, Grafana, and cloud-native observability solutions.

  • Lead incident response, root cause analysis (RCA), post-incident reviews, resilience validation, disaster recovery exercises, and continuous operational improvement initiatives.

  • Perform capacity planning, performance optimization, and reliability engineering to ensure high availability and minimal downtime for mission-critical banking services.

  • Collaborate with development, security, architecture, and engineering teams to strengthen platform reliability, deployment consistency, security compliance, and operational excellence.

  • Drive a culture of automation, reliability, accountability, continuous improvement, and engineering excellence across the organization.

Requirements:

  • Minimum with at least 4 years of experience as Cloud Architect (GCP, AWS) , and at least 2 years of site reliability engineering experience
  • Strong hands-on experience in Kubernetes platforms, Infrastructure as Code (IaC), and cloud-native technologies.
  • Practical experience in Site Reliability Engineering (SRE), observability, incident management, and production operations.
  • Minimum 2 years of experience in developer roles.
  • Experience in banking or financial services environments is highly preferred.
  • Relevant certifications such as Google Professional Cloud Architect, AWS Solutions Architect, Certified Kubernetes Administrator (CKA), or Terraform Associate are an advantage.