Senior LLMOps Engineer, Development (51577)
Description
Citrin Cooperman offers a dynamic work environment, fostering professional growth and collaboration. We’re continuously seeking talented individuals who bring a problem-solving mindset, fresh perspectives, and sharp technical expertise. We know you have choices, so our team of collaborative, innovative professionals are ready to support your professional development. At Citrin Cooperman, we offer competitive compensation and benefits and most importantly, the flexibility to manage your personal and professional life to focus on what matters most to you!
We are seeking a Senior – MLOps/LLMOps Engineer, Development, to join our Development team within the Information Technology department. The AI Solutions team is the vanguard of our enterprise AI competency, bridging the gap between rapid generative AI pilots and our enterprise operations. As we industrialize these advanced applications, you’ll build the operational backbone for our non-deterministic systems.
In this critical deployment and observability role, you’ll define how generative AI and agentic workflows are shipped to production. Working with frontier models (Anthropic, Google, OpenAI) and custom frameworks (LangGraph), you’ll transition pilot code into robust solutions, automated with CI/CD pipelines. You’ll own the infrastructure for prompt versioning, while establishing automated evaluation gates (e.g., LLM-as-a-judge), and implementing the deep telemetry required to monitor token costs, latency, and hallucination rates. The ideal candidate has a strong DevOps foundation but has successfully pivoted into the unique challenges of machine learning and generative AI operations, as well as views observability as the ultimate defense against model drift.
Responsibilities are, but not limited to
- LLMOps CI/CD Pipelines: Design and build automated deployment pipelines specifically for generative AI applications. Ensure that updates to prompts, LangGraph state machines, or RAG retrieval logic can be safely promoted across environments (Dev, Test, Prod).
- Evaluation Infrastructure: Deploy and manage the infrastructure required for continuous AI evaluation (e.g., LangSmith, Braintrust, or custom evaluation harnesses). Embed precision, recall, and toxicity checks directly into the deployment gates.
- Telemetry & Observability: Instrument the AI applications to capture deep operational metrics. Build dashboards to monitor token consumption, end-to-end latency, reasoning traces, and API failure rates across multiple LLM providers.
- Prompt & Model Registry Management: Implement version control for prompts and model configurations, ensuring the enterprise has a strict, auditable history of what instructions are running in production at any given time.
- Guardrails & Content Filtering: Integrate input/output guardrails (e.g., Azure AI Content Safety, NeMo Guardrails) into the application flow to automatically block prompt injection attacks, PII leakage, or off-topic responses.
- Cost Management (FinOps for AI): Actively monitor the financial footprint of our AI solutions. Set up alerting for token usage spikes and work with AI Engineers to optimize embedding and retrieval strategies for cost efficiency.