Sr. System Engineer/GPU Platforms
NewJob Req ID: 30083
About Supermicro:
Supermicro® is a Top Tier provider of advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing company among the Silicon Valley Top 50 technology firms. Our unprecedented global expansion has provided us with the opportunity to offer a large number of new positions to the technology community. We seek talented, passionate, and committed engineers, technologists, and business leaders to join us.
Job Summary:
Supermicro is seeking an experienced Senior Systems Engineer / GPU Platforms to support the bring-up, qualification, enablement, and customer deployment of advanced GPU computing platforms.
This role focuses on multi-GPU server systems used for AI, HPC, enterprise computing, and accelerated workloads. The successful candidate will work across the product lifecycle, from initial system bring-up and qualification through product release, customer POC/EVAL support, debugging, and post-launch technical enablement.
The ideal candidate combines strong server hardware knowledge with hands-on Linux and GPU software experience and can independently troubleshoot complex issues across hardware, firmware, operating systems, networking, and GPU software environments.
Essential Duties and Responsibilities:
Key Responsibilities
- Support system bring-up, configuration, integration, validation, and troubleshooting of advanced GPU server platforms.
- Execute and support GPU platform qualification activities, including NVIDIA NVQUAL or equivalent validation processes.
- Install, configure, and troubleshoot Linux, GPU drivers, CUDA environments, firmware, libraries, and related software components.
- Diagnose complex system issues using logs, telemetry, diagnostics, and vendor tools, and drive issues to resolution or appropriate engineering escalation.
- Support multi-GPU server platforms throughout qualification, product launch, and post-release engineering activities.
- Participate in customer-facing POC/EVAL engagements, including system preparation, technical calls, debugging, and issue resolution.
- Collaborate with internal Architecture, Systems, Software, Validation, Product Management, and other engineering teams, as well as external technology partners.
- Develop technical documentation, troubleshooting guides, and best practices.
- Deliver technical presentations, training sessions, and internal knowledge-sharing activities.
- Serve as a technical resource and mentor for other engineers when appropriate.
Key Competencies
- Strong technical depth and systems-level troubleshooting ability.
- Ownership and accountability for complex technical issues.
- Ability to quickly learn new server, GPU, and software technologies.
- Effective cross-functional collaboration.
- Clear technical communication and documentation.
- Willingness to share knowledge and support team development
5–15 years of relevant experience
Candidates should demonstrate the ability to independently support complex GPU platforms and technical customer environments. More senior candidates should additionally bring broad system-level expertise, technical leadership, mentoring experience, and ownership of complex platform or customer-facing initiatives.
Qualifications:
Required Qualifications
- Bachelor’s degree in Computer Engineering, Electrical Engineering, Computer Science, Information Technology, or a related discipline, or equivalent practical experience.
- 5–15 years of relevant industry experience in systems engineering, server engineering, platform engineering, validation, technical enablement, HPC, AI infrastructure, or a related field.
- Strong knowledge of enterprise server hardware and system architecture.
- Hands-on experience with Linux server environments.
- Experience installing, configuring, validating, and troubleshooting server hardware and software.
- Strong system-level troubleshooting and root-cause-analysis skills.
- Working knowledge of PCIe architectures and high-performance I/O.
- Experience with GPU computing, accelerators, or comparable high-performance computing technologies.
- Ability to independently manage complex technical assignments and drive issues toward resolution.
- Strong written and verbal communication skills.
- Ability to work effectively with cross-functional and geographically distributed engineering teams.
- Comfortable participating in customer-facing technical discussions.
Preferred Qualifications
- Hands-on experience with NVIDIA data center or professional GPU platforms.
- Experience with CUDA and NVIDIA GPU software environments.
- Experience with NVIDIA NVQUAL or similar platform qualification processes.
- Experience with 4-GPU or 8-GPU server platforms.
- Familiarity with NVIDIA Blackwell, B200, Rubin, or comparable accelerator architectures.
- Knowledge of PCIe topology, NUMA, DMA, IOMMU, and GPU-to-NIC communication.
- Experience with GPUDirect RDMA, InfiniBand, RoCE, or high-speed Ethernet.
- Familiarity with NCCL, NVML, DCGM, Fabric Manager, or similar GPU diagnostic and management tools.
- Experience with Docker, containers, Kubernetes, or related orchestration technologies.
- Experience supporting AI, machine learning, HPC, or accelerated computing environments.
- Experience with customer POCs, technical evaluations, or engineering escalations.
- Experience delivering technical training or knowledge-sharing sessions.
- Bash, Python, or other scripting experience is a plus.
Salary Range
$137,000 - $156,000
The salary offered will depend on several factors, including your location, level, education, training, specific skills, years of experience, and comparison to other employees already in this role. In addition to a comprehensive benefits package, candidates may be eligible for other forms of compensation, such as participation in bonus and equity award programs.
EEO Statement
Supermicro is an Equal Opportunity Employer and embraces diversity in our employee population. It is the policy of Supermicro to provide equal opportunity to all qualified applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, protected veteran status or special disabled veteran, marital status, pregnancy, genetic information, or any other legally protected status.