Integrates open-source AI models (LLMs, embeddings, rerankers, speech) into business-critical on-prem production services at a Swiss telecom, owning solutions end to end: design, deployment, RAG pipelines, self-hosted inference stacks (vLLM, Ollama), GPU tuning, monitoring, and a 24/7 on-call rotation, using Linux, Kubernetes, and SQL.
Sign in to see your match