Research Engineer - LLM/VLM Inference Optimization (Seed Infra)
Summary
Design and build high-performance inference engines for large language and vision-language models, optimizing latency, throughput, and cost while collaborating with research teams.
- Build model inference engines using performance optimization
- Collaborate with research teams to improve model toolchains
- Conduct performance analysis and identify bottlenecks
- Design high performance inference systems for LLMs and VLMs
- Develop end-to-end deployment pipelines
- Optimize large model serving latency throughput and cost