freehire launches on Product Hunt on 26 August.

Follow →

LLM Interface GPU Consultant

Summary

Engineer and maintain on-prem LLM inference systems using NVIDIA H200 GPUs and OpenShift AI to host open-source models like Llama.

Job Title: LLM Inference & GPU Systems Consultant Location: Charlotte, NC (Hybrid) Role Overview: We are seeking an AI Infrastructure Runtime Engineer to build and maintain large-scale on-prem LLM infrastructure. This is an enterprise private GenAI environment running on NVIDIA H200 GPU clusters and an OpenShift AI deployment ecosystem. You will manage production inference internally, including self-hosting open-source LLMs like Llama. We are focused exclusively on inferencing; this role involv…

See also