MLOps engineer to run self-hosted LLM inference (vLLM/SGLang) on NVIDIA GPUs behind a LiteLLM gateway in a closed, air-gapped corporate environment — plus Prometheus/Grafana monitoring, model evaluation/licensing, client tooling setup, and a RAG retrieval layer over internal code, docs, and tasks. Remote position based in Moscow.
Sign in to see your match