Senior AI engineer building and running Firmus's self-hosted AI model inference platform on its GPU cloud — onboarding models, provisioning secure scalable endpoints, optimizing LLM serving performance (quantization, batching, parallelism), and benchmarking runtimes like TensorRT-LLM, vLLM, and Triton on NVIDIA CUDA/Kubernetes infrastructure.
Sign in to see your match