Sustaining Software Engineer – Distributed Storage
Salary: $100,000 – $150,000 per year
About the Role
We are looking for a Sustaining Software Engineer with a strong background in distributed systems, storage technologies, and C/C++ development.
In this role, you will be responsible for maintaining and improving our production-grade storage products. You will investigate complex software defects, troubleshoot customer-reported issues, develop reliable fixes, and work closely with support, QA, and engineering teams to ensure product stability.
This is a hands-on engineering position. It combines software development, production troubleshooting, root-cause analysis, and customer issue resolution. The ideal candidate enjoys working on mature and complex systems, understands how software behaves in real-world production environments, and is comfortable debugging problems across multiple layers of the system.
Key Responsibilities
Maintain and improve distributed storage and data infrastructure products.
Investigate and resolve customer-reported product issues, including performance degradation, service interruptions, data-path failures, and unexpected system behavior.
Reproduce production issues in laboratory or test environments and identify their root causes.
Develop, review, test, and deliver high-quality bug fixes using C or C++.
Analyze system logs, core dumps, stack traces, performance metrics, and storage-related diagnostic information.
Troubleshoot issues involving distributed systems, storage engines, networking, operating systems, concurrency, and hardware interactions.
Work closely with customer support and field engineering teams to collect technical information and provide troubleshooting guidance.
Collaborate with product development teams on complex defects, architectural improvements, and product reliability.
Create diagnostic tools, scripts, test cases, and internal documentation to improve troubleshooting efficiency.
Participate in product release validation and ensure fixes are safely backported to supported product versions.
Contribute to improving product observability, serviceability, stability, and maintainability.
Required Qualifications
Bachelor’s degree or above in Computer Science, Computer Engineering, Software Engineering, or a related field.
Strong programming experience in C or C++.
Solid understanding of data structures, algorithms, multithreading, memory management, and network programming.
Practical experience with Linux systems and Linux debugging tools.
Good understanding of distributed systems concepts, such as replication, consistency, consensus, fault tolerance, distributed state management, and failure recovery.
Experience with one or more storage technologies, such as:
Distributed storage systems
Block, file, or object storage
Storage virtualization
RAID, snapshots, replication, or data protection
Local file systems or storage engines
NVMe, SCSI, iSCSI, Fibre Channel, or NVMe over Fabrics
Strong troubleshooting and root-cause analysis skills.
Ability to read and understand a large and mature codebase.
Ability to communicate technical findings clearly in written and spoken English and Chinese.
Willingness to work directly with customer-facing teams on complex production issues.
Must be a Singapore citizen, permanent resident, or hold a valid Singapore work permit.
Preferred Qualifications
Experience developing or maintaining distributed storage products.
Experience with Ceph, SPDK, RocksDB, LevelDB, distributed databases, or similar infrastructure software.
Familiarity with Linux kernel storage, device drivers, file systems, or networking subsystems.
Experience analyzing core dumps with GDB and diagnosing memory corruption, deadlocks, race conditions, and performance bottlenecks.
Experience with observability and diagnostic tools such as perf, eBPF, Valgrind, AddressSanitizer, strace, SystemTap, or similar tools.
Familiarity with Python, Bash, or other scripting languages for testing and automation.
Experience in enterprise infrastructure, cloud platforms, hyper-converged infrastructure, or data centre products.
Previous experience in sustaining engineering, escalation engineering, product support engineering, or site reliability engineering.
What We Value
A strong sense of ownership and responsibility for product quality.
Patience and persistence when investigating difficult technical problems.
The ability to distinguish symptoms from root causes.
A practical engineering mindset focused on reliable and maintainable solutions.
The ability to work effectively across engineering, QA, support, and customer-facing teams.
A willingness to understand both the product source code and the customer’s production environment.