Principal AI Engineer (LATAM Remote)
Summary
Lead the design and deployment of production-grade AI agents that take real-world action, focusing on retrieval engineering, tooling, and guardrails for a mobility-focused venture studio.
Overview
Technical Challenge
Responsibilities
- Own and grow the agentic function end-to-end: set architecture across internal (data-engineering & data-science agents) and external (customer/partner-facing agents, MCP servers) surfaces, make build-vs-buy calls, and grow a team under you as we scale.
- Design, build, and deploy production LLM agents that take consequential action (write-back, execute changes) with human-in-the-loop controls — including tool interfaces (MCP, function calling), tiered tool access, approval gates, and rollback.
- Build and maintain eval harnesses for agentic systems — offline and online evaluation, regression testing for non-determinism, guardrails, and observability/tracing.
- Own context and retrieval engineering: go beyond naive chunking to build high-quality retrieval, context assembly, and grounding in proprietary domain data to reduce hallucination.
- Implement graph-based knowledge and retrieval systems (GraphRAG, property graphs, ontology/semantic layers) to ground agents in a complex, structured domain.
- Partner with applied science and the broader data/ML function on pipelines, embeddings, and vector stores where relevant, without owning ML model training.
- Communicate agent architecture, trade-offs, and roadmap to execs and investors; act as a player-coach who writes production code today while owning the function's direction and hiring as the team grows.
Required Skills
- 5+ years shipping production software; strong full-stack/backend engineering (Python core, TS/Node or Go a plus), including production APIs, data models, testing, and CI.
- Proven track record building and shipping LLM agents to production with real users — multi-step, tool-calling, stateful, with orchestration (LangGraph or equivalent) — and the ability to explain the control loop, not just the framework used.
- Experience building eval harnesses for agentic systems (e.g. MLflow, LangSmith, or custom) with fluency in determinism, drift, and guardrails.
- Experience owning a retrieval system in production, including chunking vs. structured retrieval trade-offs and evaluating retrieval quality.
- Experience shipping agents that take consequential, real-world action, with a clear point of view on approval/guardrail/rollback architecture (bonus: an incident where the agent did the wrong thing, and how it was handled).
- Track record leading an agentic initiative or team end-to-end — from architecture to production — including communicating agent systems to both execs and investors.
Preferred Skills
- Experience shipping graph-backed retrieval and explaining why graphs outperform flat vectors for structured domains (GraphRAG, knowledge graphs, ontologies, property graphs, triple stores).
- Experience fine-tuning or distilling an open-source model (LoRA/QLoRA) with measured gains, or strong context/prompt optimization as a substitute.
- Experience serving/operating open-source models (Llama, Qwen, Mistral) via Databricks, vLLM, or similar, with a point of view on self-host vs. hosted-API cost/latency trade-offs.