AMD GPU Performance Engineer — Inference Backends & Kernels
Summary
Optimize AMD GPU performance for vLLM by tuning inference backends, kernels, and benchmarking tools, leveraging deep ML expertise.
Inferact is seeking an AMD GPU performance engineer in San Francisco, California, to enhance vLLM's capabilities in the AMD accelerator ecosystem. The role focuses on optimizing backends, kernels, and benchmarking infrastructure.
Qualified candidates will have experience with AMD GPU tools and a strong foundation in machine learning. The competitive salary range is $200,000 - $400,000 annually, with additional equity and generous benefits offered.
a16z