Staff Software Engineer, Agent Infrastructure
About Sapiom
Sapiom is the end-to-end platform that removes barriers to ship and scale agentic products.
We unify what an agent needs to act in the world — compute and sandboxes, memory, identity, spend controls, storage and queues, monitoring — provisioned together as one thing, not handed over as a framework to assemble yourself.
We have assembled a world-class team with deep infrastructure and payments DNA to build the operating system for machines. Our founder ran payments engineering at Shopify for five years and built an autonomous consumer agent company before that. We raised a $35M Series A led by Dragonfly in August, bringing us to $50M total. Accel led our seed; Menlo Ventures and Anthropic are also behind us.
The models are the brain. We're the spine.
About the role
Agents can think now. They still can't act — not economically, not reliably, not with anyone in control. Teams build a great demo in days, then find that running it in production is brutally hard: it breaks when nobody's watching, the bill arrives before the explanation, and nobody can reconstruct what happened. Closing that gap is the whole company.
At thirty people, a Staff engineer here isn't a level above the work, they’re the person who decides what the next version of a system looks like and then builds enough of it that the rest follows. You'd own one of the hard problem areas outright. The architecture is still soft enough that your judgment sets it, and hard enough that getting it wrong is expensive.
We hire at two levels: engineers one to three years in, and Staff. There's no layer between, which means your designs get built by people who will ask why, and your answers become how the team thinks.
What we're working on
Routing and capacity. Every model call has to be placed against latency, cost, and quality targets in real time, and the scheduling, reservations, and forecasting underneath have to keep those decisions honest as demand shifts. Getting this right is what makes running agents economical at all.
Metering and billing correctness. Two charging layers, hold and capture semantics, and 270M+ consumption events that all have to reconcile — in a ledger and in a usage number a customer is reading right now. Metering is the product here, not a feature, which means the correctness bar is unusually high.
Reliability at sustained scale. We've taken a 10–20x increase that kept compounding week over week for months. The open question is what reliability should mean here, and what observability, incident practice, and architecture get us ahead of the growth rather than reacting to it.
The control plane. Credential-scoped permissions, workflow state limits, gateway hardening. These are the decisions that determine how much authority a customer can safely hand an agent — largely unsettled, and consequential.
Agent execution. Sandboxed runs, state, memory, tools, and the external integrations agents depend on, at tens of thousands of runs a day.
You may be a fit if
You've designed and operated large-scale distributed systems in production and been accountable for them through real failures.
You've independently scoped and delivered ambiguous multi-month projects from a blank page to a production system other engineers build on.
You set technical direction across a team rather than executing within it, and you can bring people with you without authority.
Your architectural calls have held up over years — and you can name the ones that didn't and what you learned.
You've built the operational foundation too: incident response, postmortems, on-call that's sustainable.
You mentor deliberately and make other engineers better without becoming a bottleneck.
You use AI tools in your own engineering workflow — not instead of judgment, but as a multiplier.
Nice to have: experience with orchestration or workflow engines, queues, storage, or observability infrastructure; systems that talk to a lot of third-party APIs; metering, billing, or usage-based pricing at scale; LLM inference or agent architectures; a stint as a tech lead; early-stage companies where you set the architecture rather than inherited it.
If this reads like a stretch, apply anyway. At this stage we care more about how fast you close gaps than which ones you've already closed.
Applying
A recruiter screen, then a technical screen. If those go well, a three-part loop: an architecture deep dive, a hands-on AI project, and a conversation with our founder.
Everyone on our engineering team carries the title Member of Technical Staff internally. We post by level so the scope is clear.
Skills
As published by ashby · 3 questions
Basics
Name, Email, Location, Resume
Short answers (1)
- LinkedIn optional
Pick from a list (2)
- Are you able to be our San Francisco office 5 days a week?
- Do you currently reside in the Bay Area? optional
