Point your AI agent at freehire and let it find you a job.

Get the CLI →

Mapgenesys

NewBe an early applicant

Sr AI Engineer

Posted Updated
Discussion

Summary

Senior AI Engineer owning the AI layer of an autonomous IT-operations platform that uses multi-agent LLM systems to detect, diagnose, and resolve infrastructure incidents. Day to day: building agent reasoning and planning pipelines, advanced RAG retrieval, and continuous learning in Python with frameworks like LangChain/LangGraph.

Compensation: $80k – $150k

Senior AI Engineer Location: Chicago, US Type: Full-time Department: AI & Automation — IT Operations ──────────────────────────────────────────────────────────── About the Role We're building an AI-powered autonomous operations platform that uses multi-agent AI to detect, diagnose, and resolve IT infrastructure incidents — with minimal human intervention. Our platform coordinates specialized AI agents that collaborate across the incident lifecycle, from initial alert through resolution. We have a working MVP processing real incidents end-to-end. We need someone to take the AI layer from functional to intelligent — improving agent reasoning, building smarter retrieval, and making the system genuinely learn from every incident it handles. You'll own the AI layer: agent reasoning, multi-step planning, tool use, knowledge retrieval, and continuous learning. ──────────────────────────────────────────────────────────── What You'll Do • Advance agent reasoning — Design and refine multi-step reasoning pipelines where AI agents dynamically plan, execute, and re-plan based on results. Push toward 80%+ autonomous incident resolution • Build intelligent knowledge retrieval — Implement advanced RAG patterns including hybrid search, re-ranking, chunking strategies, and adaptive retrieval across enterprise knowledge bases • Develop continuous learning capabilities — Build pattern mining that discovers recurring incident patterns, constructs causal models, and auto-generates operational runbooks from historical resolutions • Design multi-agent collaboration — Architect how agents share context, negotiate priorities, and dynamically route work based on incident complexity and confidence levels • Implement safety and guardrails — Extend the safety layer with risk scoring, blast radius estimation, and progressive autonomy models (shadow → supervised → autonomous) • Optimize cost and latency — Implement intelligent model routing (powerful models for complex reasoning, lightweight models for classification), response caching, and token optimization strategies ──────────────────────────────────────────────────────────── What You Bring Must Have • 5+ Python — async/await, FastAPI or similar frameworks, strong software engineering fundamentals • Hands-on LLM application development — You've built production systems where LLMs reason, plan, and use tools. Prompt engineering is second nature • Experience with agentic AI frameworks — LangChain, LangGraph, CrewAI, AutoGen, or similar. You understand state management, checkpointing, and conditional routing in agent workflows • RAG implementation experience — Vector databases, embedding models, hybrid search, retrieval evaluation. You know why naive RAG fails and how to fix it • Comfortable with ambiguity — You'll design, build, test, and iterate in a fast-moving environment. You figure out the right approach, not wait for detailed specs • Observability & production readiness — Tracing (Langfuse/OpenTelemetry), structured logging, prompt/version tracking, monitoring token usage, model drift detection • Cost-performance optimization — Model selection tradeoffs (latency vs reasoning depth), caching strategies, batching, tool-call efficiency, and context window optimization • Security & governance awareness — Prompt injection mitigation, data isolation, secrets handling, PII redaction, and secure tool access patterns Nice to Have • Experience with IT operations, NOC, or ITSM platforms (ServiceNow, PagerDuty, etc.) • Knowledge of causal inference or pattern mining from operational data • Experience with MCP (Model Context Protocol) or similar tool-use frameworks • Familiarity with vector databases (pgvector, Pinecone, Weaviate, Qdrant) • Prior work on AI safety and guardrails for autonomous systems • Experience deploying LLM systems in cloud-native environments (AWS/GCP/Azure, Kubernetes) ──────────────────────────────────────────────────────────── Why Join • Greenfield AI architecture — You're building autonomous agents that reason, learn, and collaborate — not maintaining legacy ML pipelines • Real production impact — Every improvement directly reduces mean time to resolution and operational toil • Cutting-edge applied AI — Multi-agent coordination, agentic RAG, continuous learning from operations data • Ownership — Small team, high autonomy. You'll shape the AI architecture, not just implement tickets

Skills

See also

AI Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available