AI Engineer (Agentic Systems / LLM)
Summary
Build and maintain production-grade LLM-based agentic systems for a large North American energy company, focusing on multi-agent workflows, RAG pipelines, reliability, and performance optimization.
Full-time
6months+
EU Only
About the Client
A major North American energy company — one of the largest players in energy infrastructure and distribution on the continent. The organization is investing in AI to modernize how it operates at scale, building LLM-based agentic systems to support decision-making and automation across a complex, data-rich environment. You'll join a team applying cutting-edge AI engineering to real production problems in a large, established enterprise.
Overview
You'll work on an AI/LLM platform built around agentic systems — designing, orchestrating, and hardening multi-agent workflows in production. This is a deeply hands-on engineering role focused on making LLM-based systems reliable, fast, and measurable, not just wiring up prompts.
Responsibilities
Agentic systems: Build and operate multi-agent architectures — agent roles, triggers, handoffs, orchestration.
RAG pipelines: Set up, tune, and fine-tune RAG systems end-to-end — chunking strategies, retrieval quality, iteration.
Reliability engineering: Implement self-healing behavior — infinite-loop mitigation, graceful failures, clean handoffs between agents.
Memory systems: Design and implement shared-memory concepts, with emphasis on fast memory.
Evaluation & KPIs: Build evaluators of LLM performance, define KPIs, and automate evaluation (golden images, sample sizing).
Performance & cost: Apply caching and token-optimization strategies to keep systems fast and economical.
Who You Are
A hands-on builder who's actually shipped agentic/LLM systems, not just experimented with them
Rigorous about reliability and measurement — you treat evals and failure modes as first-class work
Comfortable owning ambiguous, fast-moving AI infrastructure end-to-end
Tech Stack You'll Work With
Core: Python, LLM tooling (llama / ollama), CUDA
AI systems: Multi-agent orchestration frameworks, RAG pipelines, LLM evaluation tooling
Concepts: Shared/fast memory, self-healing patterns, caching & token optimization, chunking strategies
Cloud: Azure (strong advantage)
Qualifications
Must-Have
Strong hands-on experience with agentic systems (multi-agent orchestration, agent roles, triggers, handoffs)
Strong hands-on experience building, tuning, and finetuning RAG pipelines (incl. chunking strategies)
Experience with LLM performance evaluation — building evaluators, defining KPIs, automating them
Self-healing / resilience implementation (loop mitigation, graceful failures, handoffs)
Shared memory concepts and implementation, especially fast memory
Caching and token-optimization strategies
Python; llama/ollama; CUDA
Nice-to-Have
Azure (products relevant to the above) — strong advantage
Mistral — strong advantage
Golden images and sample-size methodology for evaluation