Senior AI Developer
About the role
We build production AI agent systems for enterprise clients: a multi-agent system licensed to an external client, reasoning across several distinct data sources (survey, ad-platform, psychographic, live web search) and routing between knowledge bases, SQL-based RAG, and MCP tools; and a conversational analytics agent used by hundreds of people across multiple markets, recently migrated onto a new Bedrock/LangGraph architecture.
This role is as hands‑on as it is senior: expect real day-to-day coding, not just review and direction. You'll also help architect new projects as they come in, and shape the team's standards while it's still being built, rather than join one where that's already fixed.
What you’ll do
- Own the architecture and engineering of a production multi-agent system, including reliability, cost, and security
- Design for heterogeneous tool use across knowledge bases, SQL-based retrieval, and MCP tools
- Review and harden access‑control and data‑boundary logic on systems handling sensitive client data
- Build evaluation and observability practices, including monitoring, logging, and error handling, so quality and reliability are measured, not assumed
- Design data ingestion and preprocessing workflows for structured and unstructured data, with a clear sense of when a stateful vs. stateless approach fits
- Collaborate with data scientists, backend engineers, and DevOps to bring systems into production
- Take on new projects as they come in, applying the same architectural approach each time
Must-have
- Proven track record building and operating agent‑based or multi‑agent architectures in production (LangGraph or comparable), not just prototypes or demos
- Real experience combining multiple retrieval or tool‑use approaches in one system (knowledge bases, SQL-based retrieval, vector search), and reasoning about when to use which
- Hands‑on experience across a real AWS environment: Bedrock, Lambda, Step Functions, DynamoDB, EventBridge, IAM, EKS, Cognito
- Hands‑on prompt engineering experience, designing effective interactions with LLMs
- Comfort owning security‑sensitive review: access control, identity/permission logic, data‑boundary design, and secure handling of confidential client data
- Strong backend engineering fundamentals (Python; FastAPI or similar), with solid Git/GitHub workflow practice
- Real, hands‑on experience building or running evaluation and observability practices for LLM/agent systems, plus familiarity with broader MLOps concepts (deployment, monitoring, versioning). This is a specific area we want to strengthen, so depth here matters more than most other line items
- Experience with infrastructure‑as‑code (Terraform or CloudFormation) and CI/CD pipelines
- Deep understanding of event‑driven and serverless architectures, and the judgment to weigh them against container‑based approaches
Nice‑to‑have
- Enough grounding in statistical or ML concepts (anomaly detection, forecasting, basic analysis methods) to guide an agent that performs this kind of analysis
- Experience in an agency, consultancy, or client‑services environment
- Familiarity with prompt‑injection defense and LLM security practices
What success looks like in the first 90 days
- You've shipped real code to a production system, not just reviewed or advised on others' work
- You're a genuine second reviewer on at least one existing production system
- You've contributed to defining a piece of the team's shared standards
- You're proposing architectural or technical improvements independently, not just executing assigned tickets