Software Engineer Kernels CUDA C++
You will design, build, and optimize GPU clusters for training and inference. You will develop CUDA kernels, profile and remove bottlenecks across GPU memory, networking, filesystems, and multi-GPU operations, and deliver production-grade performance and scalability with research teams.
Responsibilities
- Design, build, and optimize GPU clusters for training and inference workloads
- Develop and tune low-level CUDA kernels
- Profile, debug, and eliminate bottlenecks across GPU memory, networking, filesystems, and multi-GPU operations
- Collaborate with AI research teams to deliver production-grade performance and scalability
Requirements
- Low-level systems programming knowledge in C, C++, PTX, or SASS
- Experience with large-scale GPU clusters or distributed compute infrastructure
- Experience with GPU kernel optimization using CUTLASS, custom kernels, or Nsight profiling
- Experience building or running high-performance infrastructure for AI training or inference workloads
- Ability to optimize memory-bound and compute-bound scenarios
Benefits
- Equity
- Medical coverage
- Vision coverage
- Dental coverage
- 401(k) retirement plan
- Short-term disability insurance
- Long-term disability insurance
- Life insurance
- Discounts and perks