Build and own the LLM/agentic evaluation infrastructure—datasets, scorers, harnesses, and CI gates—that rigorously measures grounding, faithfulness, and hallucination before AI-generated financial commentary ships. Core tech: Python, LLM eval frameworks (promptfoo, DeepEval, Ragas, LangSmith), managed LLMs (Bedrock/Vertex/Azure OpenAI), LangGraph, GitHub Actions.
Sign in to see your match