Builds and operates self-hosted AI inference infrastructure: onboarding models and running them as secure, scalable, high-performance GPU-backed endpoints for internal products, customers, and future Inference-as-a-Service. Day-to-day work centers on model serving frameworks (TensorRT-LLM, vLLM, SGLang, Triton), CUDA/NVIDIA stack optimization, benchmarking, and Kubernetes-based deployment.
Sign in to see your match