Software Engineer, Agent Platform
Summary
Senior engineer on Retool's Agent Platform owning AI/LLM behavior in production: designing prompting, retrieval, routing and tool-use strategies, building evaluation and guardrail systems, and preventing regressions in non-deterministic systems. Full-stack work in TypeScript, Node.js and React.
- Own the behavior of AI-powered features across multiple product surfaces, including quality, safety, variance, and failure modes
- Design and evolve prompting, retrieval, routing, and tool-use strategies that embrace non-determinism while bounding its downside
- Build and maintain evaluation systems that measure model performance using statistical signals, distributions, and trends—not just pass/fail tests
- Detect, diagnose, and resolve non-deterministic failures such as hallucinations, partial correctness, instruction drift, or sensitivity to context changes
- Define and implement guardrails, fallbacks, and degradation paths that keep systems useful even when models behave unexpectedly
- Partner with product and infra teams to decide when probabilistic behavior is “good enough” to ship—and when it isn’t
- Influence model selection, model behavior, and tool design to balance quality, cost, latency, and robustness for real user workflows
- Accountable for AI behavior, not just system correctness
- Grounded in evaluation, iteration, and regression prevention under non-determinism
- Comfortable designing systems where outputs vary, confidence is probabilistic, and correctness is contextual
- Focused on shipping dependable products on top of imperfect components
- Adding LLM calls to existing features and moving on
- Treating models as black boxes with undefined behavior
- Shipping AI features without owning their long-term reliability, drift, or user trust
- 6+ years of professional engineering experience, with ownership over complex systems in production
- Demonstrated experience owning AI/LLM behavior beyond basic integration, including mitigation of variance and failure modes
- Comfort reasoning about probabilistic systems and tradeoffs (quality vs. cost, recall vs. precision, speed vs. robustness)
- Experience designing or maintaining evaluation frameworks, golden datasets, regression detection, or human-in-the-loop feedback loops
- Strong product intuition—you care deeply about what “good” looks like even when outputs are non-deterministic
- Ability to operate independently in ambiguous problem spaces and set quality standards others rely on
- Strong opinions, weakly held—you iterate quickly and adjust based on evidence and observed runtime behavior
- Experience with RAG, agentic systems, or tool-using models in production
- Familiarity with vector databases, embeddings, or retrieval pipelines
- Exposure to fine-tuning, model routing, or post-training techniques
- Experience building shared AI infrastructure used by multiple teams
- History of mentoring engineers on designing for non-determinism and evaluation-driven development
