Staff /Principal Engineer – Core Team
Summary
Builds and scales the core infrastructure layer that routes, orchestrates, and governs production AI systems for enterprises, using cloud-native and distributed systems.
About TrueFoundry
Every production AI system, whether it's powering customer support, writing code, analyzing financial data, or diagnosing medical conditions, needs the same foundational infrastructure.A way to route between models. A way to manage tools and integrate them securely. A way to orchestrate agents and enforce governance. A unified compute layer to run it all.
That infrastructure layer is being built right now.
We are looking for a Staff/Principal Engineer – Core Team.
The Problem We're Solving
Companies are moving beyond simple chatbots to production agentic systems. These systems route between OpenAI, Anthropic, Google, and self-hosted models. They integrate dozens of tools via protocols like MCP. They orchestrate multi-agent workflows where agents coordinate with other agents.
The infrastructure to support this doesn't exist yet. You can't just duct-tape together a few API calls and call it production-ready.
You need a control plane that handles:
- Intelligent routing with observability, cost policies, and fallback logic
- Centralized tool and MCP server management with security and lifecycle controls
- Agent orchestration with governance and guardrails
- A unified compute layer to run self-hosted models, custom tools, and agents
AI Gateway is the control plane: five composable components (Prompts, LLM Gateway, MCP Gateway, Guardrails, Agent Gateway) that handle routing, orchestration, and governance.
We're Series A, backed by Intel Capital and Sequoia. Companies like CVS, Mastercard, Siemens, Paytm, Synopsys, and Zscaler run production AI workloads on our platform.
The Role:
- Solve some of the most complex Engineering problems and drive it alongside a team of engineers & ML researchers.
- Build a deep, holistic understanding of the TrueFoundry platform across all components and shape the product vision and implementation.
- Act as the technical face of engineering for customer-related discussions and escalations
- Guide and unblock engineers across projects in the US region
- Partner closely with our CTO and India-based engineering team to drive system design, architecture, and implementation of complex products
- Lead technical design, critical customer problem-solving, and platform scalability initiatives end-to-end
This is a high-ownership, high-impact role designed for an engineer who loves combining world-class systems thinking with real-world execution.
What You’ll Do:
- Build and scale TrueFoundry's MCP Gateway and Agentic Gateway, enabling secure, reliable, and scalable AI agent interactions.
- Design and develop cloud-native, distributed systems that power AI agents, tool integrations, and enterprise AI workloads.
- Build core capabilities such as authentication, authorization, routing, observability, security, and multi-tenant infrastructure for the gateway platform.
- Collaborate closely with the CTO to shape the technical architecture, product roadmap, and long-term platform vision.
- Drive architectural decisions, participate in design and code reviews, and ensure high engineering standards across the platform.
- Work closely with enterprise customers to understand their AI infrastructure needs and translate feedback into scalable platform capabilities.
- Mentor engineers across teams, helping them build high-quality, reliable, and maintainable systems.
- Continuously improve platform performance, reliability, scalability, and developer experience while reducing technical debt.
- Stay at the forefront of emerging AI infrastructure, Model Context Protocol (MCP), and agentic systems, bringing new ideas into the product
Who You Are:
- 8+ years of strong backend/systems engineering experience at top technology companies or startups
- Deep expertise in distributed systems, cloud-native architectures, and scalable system design
- Strong working knowledge of Kubernetes, containerized workloads, and infrastructure engineering
- Practical experience building or deploying ML/GenAI applications (or closely working with ML/DS teams)
- Skilled in programming languages such as Python, Golang, Node.js, and TypeScript
- Solid understanding of system observability, resiliency design, and SRE practices
- Strong technical leadership and communication skills — able to work with both customers and engineering teams
- Ability to think strategically while also executing hands-on when required
Traits we are looking for: Ownership, ability to execute, hustle and think out of the box, data-driven decision making, be comfortable with more unknowns than knowns.
Perks of Working at TrueFoundry
- Join a fast-growing Series A, Bay Area-based startup building cutting-edge AI infrastructure.
- Comprehensive health insurance for you and your family, including medical, dental, and vision coverage.
- 401(k) retirement plan.
- Flexible hybrid work, 2 days a week in the office (Tuesday & Wednesday), with flexibility around the schedule.
- Lunch and snacks are on us whenever you come into the office.
- Work from another TrueFoundry office for up to 1 month each year, whether that's our London or India office, if you'd like a change of scenery.