freehire launches on Product Hunt on 26 August.

Follow →

Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization

Summary

Designs and optimizes large-language-model inference systems, quantizing models, deploying LLMs, and scaling GPU-based pipelines for low-latency production use.

- Apply model quantization - Architect LLM deployment environments - Build system stability frameworks - Collaborate with perception prediction teams - Enable rapid model iteration - Establish profiling and telemetry frameworks - Integrate optimization toolchains - Lead root cause analysis - Monitor GPU utilization - Optimize hardware execution - Optimize inference performance - Own model deployment - Scale service deployment pipelines - Schedule compute resources

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available