Senior AI Platform Engineer — Multi-Agent Systems
Summary
Senior engineer who designs, ships, and operates production multi-agent AI infrastructure — agent orchestration, tool calling, state/memory, model routing, evaluation, and observability — for a large client's AI-native products. Requires 5+ years backend engineering (Python/Go/Java/TypeScript), distributed systems, and cloud (AWS/GCP/Azure) experience.
Compensation: $200k – $250k • No equity
About the Opportunity Our client is a large, high-growth software company investing heavily in AI-native product experiences used by a massive customer base. We are hiring a Senior AI Platform Engineer focused specifically on production agentic and multi-agent systems. This is not a prompt-engineering position. This is not a role for someone whose primary AI experience is building simple chatbots, adding an LLM API to an application, or creating basic RAG pipelines. We want an exceptional software engineer who has actually architected, built, deployed, and operated sophisticated agentic systems in production. What You'll Build Production infrastructure for running intelligent agents at scale. Multi-agent systems where specialized agents coordinate, delegate work, share state, use tools, and execute complex workflows. Agent orchestration and execution frameworks. Tool-calling architectures connecting agents with APIs, applications, search systems, and internal services. Planning and task-decomposition systems. Long-running and asynchronous AI workflows. Context management, state management, and memory systems. Model-routing and provider-fallback infrastructure. Agent evaluation and testing frameworks. Human-in-the-loop systems for uncertain or sensitive actions. Retrieval and enterprise knowledge systems supporting agent reasoning. Monitoring, tracing, feedback, and observability for production AI. Backend services capable of supporting agentic systems under significant production load. Multi-Agent Experience Is the Highest Priority Strong candidates should have personally built production systems involving multiple areas such as: Multi-agent architectures Agent-to-agent communication Orchestrator and worker-agent patterns Tool selection and tool calling Planning and task decomposition Persistent agent state Short-term or long-term memory Context management Autonomous execution Long-running workflows Human-in-the-loop controls Agent evaluation Agent tracing and observability Retry and recovery mechanisms Model routing and fallback You should be able to clearly explain: Why multiple agents were used, what each agent was responsible for, how they communicated, how state was maintained, how failures were handled, and how you evaluated whether the system actually worked. Software Engineering Requirements Candidates must be strong software engineers first. 5+ years of professional software engineering experience. Strong backend engineering fundamentals. Experience building and operating distributed production systems. Strong system-design skills. Experience designing production APIs and services. Experience with asynchronous or event-driven architectures. Understanding of queues, retries, concurrency, idempotency, failure handling, and scaling. Strong programming ability in Python, Go, Java, TypeScript/Node.js, or comparable backend languages. Cloud production experience across AWS, GCP, or Azure. Experience with Docker, Kubernetes, CI/CD, monitoring, and observability. Highly Relevant Technologies Relevant experience may include: LangGraph LangChain LlamaIndex AutoGen CrewAI Semantic Kernel Temporal Ray OpenAI Anthropic Google Gemini Vector databases Semantic and hybrid search RAG LLM evaluation systems AI observability and tracing Listing these technologies on a résumé is not enough. We care about what you personally designed and shipped with them. Production AI Experience The strongest candidates have: Shipped customer-facing AI products. Operated agentic systems under real production load. Built multi-agent workflows used by real customers. Debugged unpredictable agent or model behavior. Built automated evaluation systems. Improved latency, cost, throughput, or reliability. Designed model-routing and fallback mechanisms. Implemented safeguards around sensitive information and permissions. Owned systems after launch rather than stopping at the prototype stage. Ideal Background We are intentionally looking for candidates from strong engineering environments. Relevant backgrounds may include: Leading AI-native software companies High-growth SaaS companies Developer infrastructure companies Enterprise search and knowledge platforms Top-tier consumer or B2B technology companies Strong AI startups Teams building production copilots, coding agents, autonomous workflows, or complex AI platforms Experience at companies similar to OpenAI, Anthropic, Google, Microsoft, Meta, Stripe, Datadog, Snowflake, Scale AI, Glean, Notion, Figma, Ramp, Rippling, or other high-bar engineering organizations can be a strong signal. But pedigree alone will not carry someone through the process. We care much more about whether you can deeply explain a sophisticated production system that you personally helped architect and build. This Role Is Not a Fit If Your experience is primarily prompt engineering. Your primary work has been creating prompts or prompt templates. You have mostly built simple chatbots. Your experience is primarily single-step LLM API integrations. You have only built basic RAG applications. Your work has mainly been prototypes, hackathons, demos, or notebooks. Your background is primarily ML research or data science. You have used LangChain or LangGraph but cannot explain the underlying architecture. You cannot explain how agents communicate, maintain state, use tools, recover from failures, or get evaluated. You have never supported an AI system after it entered production. Why Consider It This role sits well beyond basic generative AI. You will work on problems involving: multi-agent coordination, autonomous execution, tool use, planning, state, memory, retrieval, model routing, evaluation, observability, reliability, privacy, and distributed systems. If you want to build the infrastructure that allows AI systems to reason and act, rather than simply generate text, this is the type of role we want you working on.Skills
- Agentic AI
- AI
- Anthropic
- API
- AutoGen
- AWS
- Azure
- CI/CD
- Cloud
- CrewAI
- Data Science
- Datadog
- Distributed Systems
- Docker
- Event Driven Architecture
- Figma
- GCP
- Gemini
- Generative AI
- Java
- Kubernetes
- LangChain
- LangGraph
- LlamaIndex
- LLM
- Machine Learning
- Node.js
- Notion
- Observability
- OpenAI
- Prompt Engineering
- Python
- SaaS
- Semantic Kernel
- Snowflake
- TypeScript
- Vector Databases