freehire launches on Product Hunt on 26 August.

Follow →

Senior On-Premise LLM Inference & GPU Systems Consultant

Summary

Senior engineer building and maintaining on-premise LLM inference systems on NVIDIA H200 clusters and OpenShift AI, deploying open-source LLMs like Llama for private GenAI environments.

Job Description: Senior On-Premise LLM Inference & GPU Systems Engineer Location: Charlotte, NC Role Overview We are seeking an AI Infrastructure Runtime Engineer to build and maintain large-scale on-prem LLM infrastructure. This is an enterprise private GenAI environment running on NVIDIA H200 GPU clusters and an OpenShift AI deployment ecosystem. You will manage production inference internally, including self-hosting open-source LLMs like Llama. We are focused exclusively on inferencing; this…

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available