Software Engineer, Inference Runtime
Summary
Builds and optimizes AI inference engines, integrating engines, scheduling models, and diagnosing performance on CPU/GPU for cloud and edge devices.
- Benchmark and diagnose correctness and performance issues
- Bring up model architectures and multimodal models
- Build model loading batching scheduling caching and distributed execution
- Contribute upstream to open source inference projects
- Integrate inference engines and runtime capabilities
- Maintain inference stack on device and in cloud
- Optimize model execution for CPU and GPU targets
Perks/Benefits:
- Dental insurance
- Flexible PTO
- Flexible WFH
- Medical insurance
- Team meals
- Vision insurance