Site Reliability, Sr Staff - 17536
Staff Site Reliability Engineer
Location: HCM/Da Nang/Hanoi
Job Description
Synopsys is expanding Synopsys.ai alongside traditional EDA engineering compute for semiconductor customers in Vietnam and across the APAC region. We are looking for an experienced Unix/Linux infrastructure professional to keep these platforms reliable, secure, scalable, and operable in Synopsys sites and major customer environments, including customer-hosted, air-gapped, high-security, and multi-tenant on-premises deployments.
You will join the regional EDA Engineering Compute / platform operations team working with global IT, AE engineering, product, security, network, and infrastructure teams. Your work directly enables R&D and customer engineering teams to run EDA and GenAI workloads, from HPC batch grids through containerized AI gateways, vector data services, and agent/MCP-based tool execution on engineering compute platforms.
The successful candidate will support Synopsys Vietnam engineering compute and data center services and contribute to regional infrastructure initiatives. The role will also support deployment and day-2 operations of secure multi-tenant infrastructure, OpenStack-based private cloud environments, and container platforms used by internal teams and semiconductor customers.
How This Role Contributes
- Improves availability, operability, performance, and utilization of business-critical EDA and Synopsys.ai infrastructure in Vietnam.
- Supports secure multi-tenant platforms that isolate compute, network, storage, identity, and access between customers, projects, or business environments.
- Reduces operational risk for customer-hosted and air-gapped deployments through strong runbooks, monitoring, automation, incident response, and lifecycle management.
- Connects HPC schedulers, shared storage, GPU compute, private cloud infrastructure, container platforms, and platform gateways so Synopsys services run predictably at scale.
- Improves deployment speed and configuration consistency through repeatable OpenStack infrastructure automation and configuration management.
- Extends APAC shared-support coverage and technical capability across Vietnam, Japan, Korea, Taiwan, China, and other regional locations.
Responsibilities
Engineering Compute and Data Center
- Maintain Synopsys Vietnam engineering compute and data center environments in accordance with corporate IT, data center, security, and operational standards.
- Support server hardware lifecycle activities, including installation, provisioning, maintenance, upgrade planning, retirement, and decommissioning.
- Perform capacity planning, performance troubleshooting, hardware coordination, and vendor or customer coordination at Synopsys and customer sites.
- Operate and troubleshoot Linux-based EDA compute environments, including virtualization, engineering resource management, job schedulers, license connectivity, shared storage, and remote access services.
- Run HPC operations, including cluster health, scheduler integration using LSF, Slurm, or similar platforms, InfiniBand where deployed, and performance-related incident handling.
- Support storage and file services used by engineering compute environments, including NFS-based services and customer-controlled storage infrastructure.
- Collaborate with corporate network and information security teams on data center connectivity, switching, firewalls, routing, circuits, remote access, and infrastructure security.
Multi-Tenant and OpenStack Platform Operations
- Support deployment, administration, monitoring, troubleshooting, upgrade, and day-2 operations of OpenStack-based private cloud environments.
- Support secure multi-tenant infrastructure for different customers, projects, or business units, with appropriate separation of compute, network, storage, identity, access, and operational data.
- Assist with configuration and troubleshooting of OpenStack services related to compute, networking, storage, images, identity, orchestration, and the management interface.
- Support tenant or project provisioning, quotas, images, flavors, virtual networks, security controls, volumes, and access policies according to approved architecture and security standards.
- Work with network, storage, and security teams to integrate OpenStack with enterprise network segmentation, shared or tenant-specific storage, authentication services, monitoring, and security controls.
- Participate in platform capacity planning, patching, upgrades, backup and recovery planning, incident response, and lifecycle management.
- Develop and maintain operational runbooks, validation procedures, troubleshooting guides, and standardized deployment patterns for multi-tenant environments.
- Support technical assessment and implementation of regional private cloud and customer isolation requirements, including Japan-related and future APAC initiatives.
Container and Platform Operations
- Support Docker-based application and platform environments, including container images, registries, runtime configuration, networking, storage, security, and troubleshooting.
- Support deployment and day-2 operations of Kubernetes workloads and clusters in customer-controlled, on-premises, private cloud, and air-gapped environments.
- Troubleshoot Kubernetes workloads and platform services, including pods, deployments, services, ingress, configuration, secrets, persistent storage, resource allocation, logging, and connectivity.
- Apply Kubernetes namespaces, role-based access control, network policies, storage policies, and other platform controls to support workload and tenant isolation.
- Support platform lifecycle activities, including installation, configuration, monitoring, patching, upgrades, backup, recovery, and production incident handling.
- Support enterprise Kubernetes distributions or management platforms such as Red Hat OpenShift and Rancher where deployed.
- Work with application, product, and engineering teams to distinguish infrastructure, platform, network, storage, and application issues and coordinate effective resolution.
Synopsys.ai Platform Operations
- Support deployment and day-2 operations for Synopsys EDA and AI reference patterns in Vietnam and other supported APAC environments.
- Support customer-hosted and air-gapped container environments, API/AI gateways, LLM gateway integration, on-premises inference endpoints, and customer-controlled vector database or GPU compute platforms where applicable.
- Partner on installation, upgrade, and configuration of platform components that connect user or agent workflows to the engineering grid, including job services, unified MCP server patterns, and tool execution on LSF or Slurm-backed clusters.
- Implement and maintain observability, alerting, authentication and authorization integration, and operational documentation for distributed platform services.
- Support customer identity integration using applicable enterprise authentication patterns, including IdP, OIDC, SSO, or OAuth-based integration.
- Support agentic AI platform rollouts where components run locally, centrally, in containers, or on the engineering grid.
- Escalate architecture and product design questions to global architecture, product, and engineering teams while maintaining ownership of operational implementation and follow-up.
Infrastructure Automation and Operational Improvement
- Build and maintain repeatable infrastructure automation for OpenStack resource provisioning, tenant or project creation, network and storage configuration, image management, platform validation, and environment lifecycle operations.
- Use OpenStack-native orchestration, configuration management, scripting, APIs, or equivalent automation approaches to reduce manual deployment effort and configuration inconsistency.
- Create, review, and maintain reusable templates, playbooks, scripts, configuration definitions, and operational workflows using technologies such as OpenStack Heat, Ansible, Python, shell scripting, or similar tools.
- Store automation and configuration definitions in version control and follow appropriate review, documentation, testing, and change management practices.
- Automate post-deployment validation, infrastructure health checks, configuration compliance checks, and repeatable operational tasks.
- Support integration of infrastructure provisioning with Linux configuration, storage, network, Kubernetes, monitoring, and customer-specific settings.
- Own or co-own runbooks, monitoring and alerting policies, and operational automation to reduce repetitive work across the APAC team.
- Apply AI-assisted workflows where appropriate to improve troubleshooting, knowledge management, documentation, and routine operations.
Reliability, Incident Response, and Customer Engagement
- Participate in a shared on-call rotation covering weekday nights, weekends, and holidays.
- Lead or support incident bridges, technical troubleshooting, stakeholder communication, service restoration, and post-incident follow-up.
- Produce clear incident reports, investigation notes, runbooks, and operational documentation in English.
- Support planned upgrades, break-fix activities, platform deployments, infrastructure validation, and customer training at air-gapped or high-security sites.
- Coordinate with customer IT, security, data center, network, storage, and application teams during platform deployment and incident resolution.
- Travel within Vietnam and APAC as required, typically 10 to 20 percent of the time.
- Support cross-regional handoffs, documentation sharing, technical reviews, and mentoring of junior team members where applicable.
Required Skills and Experience
Technical
- Five or more years of experience in Linux system administration, infrastructure engineering, private cloud operations, or production network and storage support.
- Strong hands-on Linux administration and troubleshooting experience in production environments.
- Enterprise data center experience involving servers, storage, networking, monitoring, security controls, and operational processes.
- Experience supporting HPC or EDA engineering compute environments, including Linux clusters, workload schedulers such as LSF or Slurm, and production incident support.
- Understanding of secure multi-tenant infrastructure and the isolation of resources, networks, storage, identity, access, and operational data between customers or projects.
- Practical experience with virtualization, private cloud, or infrastructure platform operations.
- Practical experience supporting Docker-based application or platform environments.
- Working knowledge of Kubernetes architecture and operations, including workload deployment, service exposure, persistent storage, access control, monitoring, and troubleshooting.
- Ability to support GPU-enabled compute and containerized platform services in customer-controlled environments.
- Understanding of infrastructure automation and Infrastructure as Code principles, including repeatable provisioning, configuration consistency, version-controlled definitions, reusable automation, and environment validation.
- Ability to use infrastructure automation tools, scripts, APIs, configuration management, or orchestration frameworks to provision and maintain infrastructure.
- Familiarity with observability platforms and operational telemetry, including logs, metrics, tracing, alerting, and infrastructure health monitoring.
- Understanding of secure multi-tier deployments, gateways, authentication and authorization, service-to-service access, and distributed platform troubleshooting.
- Ability to troubleshoot issues across Linux, compute, virtualization, network, storage, container, and platform layers.