Machine Leaning Performance Engineer (Inference)
- Apply quantization pruning and distillation
- Benchmark inference platforms
- Collaborate with ML and hardware engineers
- Design low latency inference strategies
- Develop optimized GPU kernels
- Ensure numerical stability and low latency inference
- Identify and resolve memory and interconnect bottlenecks
- Integrate performance libraries
- Optimize inference execution pipelines
- Profile and optimize inference performance
Perks/Benefits:
- Fitness Events
- Free meals
- Hybrid work options
- Paid time off
- Volunteer opportunities
- Wellness reimbursement
- Workshops and continuous learning