Sustaining Software Engineer – Virtualization and Cloud Platform
Salary: $100,000 – $150,000 per year
About the Role
We are looking for a Sustaining Software Engineer with experience in virtualization, cloud infrastructure, and software development using Python or Go.
You will be responsible for maintaining and improving our virtualization and cloud platform products. The role focuses on troubleshooting production issues, resolving customer-reported defects, developing software fixes, and improving the reliability and serviceability of the platform.
This is not a traditional technical support role. You will work directly with source code, system logs, APIs, operating systems, virtualization components, and distributed cloud services. You will collaborate with support engineers, field teams, QA, and core development teams to diagnose complex issues and deliver production-quality solutions.
Key Responsibilities
Maintain and improve virtualization and cloud platform products.
Investigate customer-reported issues involving virtual machines, hosts, clusters, networking, storage, orchestration, and management services.
Reproduce production issues and perform detailed root-cause analysis.
Develop, test, review, and deliver bug fixes using Python or Go.
Troubleshoot problems across application, service, operating system, virtualization, networking, and storage layers.
Analyze logs, traces, metrics, API requests, service states, database records, and system configurations.
Diagnose issues involving virtual machine lifecycle management, scheduling, high availability, live migration, resource management, and cluster operations.
Work closely with technical support and field engineering teams to resolve complex customer escalations.
Provide technical guidance, workarounds, diagnostic procedures, and corrective actions for production issues.
Collaborate with core engineering teams on architectural defects and long-term product improvements.
Develop scripts, diagnostic tools, automated tests, and troubleshooting utilities.
Improve product monitoring, logging, observability, upgradeability, and operational reliability.
Participate in release validation and safely backport fixes to supported product versions.
Write clear technical documentation, root-cause analysis reports, and internal knowledge-base articles.
Required Qualifications
Bachelor’s degree or above in Computer Science, Computer Engineering, Software Engineering, or a related field.
Professional software development experience using Python or Go.
Strong Linux administration and troubleshooting skills.
Solid understanding of operating systems, processes, threads, networking, storage, and distributed services.
Experience with virtualization technologies or cloud infrastructure.
Familiarity with one or more of the following:
KVM and QEMU
libvirt
VMware vSphere or ESXi
OpenStack
Container runtimes
Cloud management platforms
Hyper-converged infrastructure
Experience troubleshooting REST APIs, background services, or distributed control-plane components.
Strong debugging, problem-solving, and root-cause analysis skills.
Ability to understand and modify an existing production codebase.
Ability to communicate technical findings clearly in written and spoken English and Chinese.
Willingness to work with customer-facing teams on complex production incidents.
Must be a Singapore citizen, permanent resident, or hold a valid Singapore work permit.
Preferred Qualifications
Experience developing or maintaining virtualization, cloud management, or infrastructure software.
Understanding of virtual machine lifecycle management, CPU and memory virtualization, virtual networking, and virtual storage.
Familiarity with high availability, live migration, cluster scheduling, resource management, and failure recovery.
Familiarity with Linux networking, Open vSwitch, bridges, VLANs, VXLAN, routing, and software-defined networking.
Experience with storage protocols or systems used by virtualization platforms.
Experience with diagnostic and observability tools such as GDB, strace, tcpdump, perf, Prometheus, Grafana, or distributed tracing systems.
Experience writing automated tests, integration tests, and troubleshooting tools.
Previous experience in sustaining engineering, escalation engineering, product support engineering, site reliability engineering, or cloud operations.
What We Value
Strong ownership of customer-impacting technical problems.
The ability to troubleshoot systematically across multiple layers.
A software engineering approach to product maintenance and customer support.
Attention to backward compatibility, upgrade safety, and production reliability.
Clear communication during complex or high-priority incidents.
The ability to balance short-term customer recovery with long-term engineering improvements.