Member of Technical Staff, Model Capabilities
Summary
Build data-tooling interfaces, evaluation frameworks, and datasets to advance frontier AI models. Work end-to-end on UI, data, and evaluation systems using Python, TypeScript, Next.js, and Kubernetes in a fast-moving in-person San Francisco team.
About TypeSafe
TypeSafe is an AI lab building intelligence beyond chat: AI designed to work inside software and power real-world automation. Our mission is to make intelligence composable—dependable enough for builders to invoke, inspect, constrain, and layer into production systems.
Today’s models are remarkably capable, but that intelligence is still hard to build on. We’re rethinking the AI stack from first principles so software can branch on meaning, judgment, and intent while code continues to handle exact computation.
We’re a small, fast-moving team from OpenAI, Google Brain, and Meta/FAIR. We care about technical rigor, real-world impact, and building systems people can trust. Join us to help turn today’s intelligence into the foundation for a new generation of software.
About the role
We're looking for scrappy engineers who ship data-tooling interfaces fast. You will progress model intelligence through the interfaces, datasets, and evaluation tooling you build, combining product sense, data science, and engineering. You'll be part of the real "secret sauce" of TypeSafe: figuring out how our models can provide production-ready reliability in the real world.
Our tech stack is primarily Python. We also use TypeScript, Next.js, and Tailwind CSS for frontend, with Kubernetes for orchestration. We empower developers to use any tooling they find helpful for getting their job done, including Claude Code and Cursor.
What you'll do
Create high-leverage datasets, products, and user interfaces that unlock new capabilities and use cases on top of our model
Build internal data-tooling and evaluation interfaces — fast — that make model development clearer (evaluation, debugging, data inspection)
Develop and own evaluation frameworks to measure quality, reliability, and emergent characteristics across model iterations
Run rigorous analyses and experiments to understand how data, training choices, and targeted interventions impact model behavior
Turn raw model outputs and data into clear, inspectable UIs — owning a domain end-to-end, with craft and judgment over simple optimization
Partner closely with product managers and customers to translate real-world needs into concrete model and system improvements
Continuously iterate, learn, and co-discover new techniques for building exceptional, trustworthy AI products
We are looking for people who
Are generalists (~2–7 yrs) with solid fundamentals, strong coding ability, and product intuition
Maintain high attention to detail and a strong bar for quality
Enjoy thinking beyond the code to how systems are used in the real world
Are scrappy and hacky, and thrive in ambiguous problem spaces — taking satisfaction in finding creative solutions
Don't trust LLMs blindly — they look at the reality of model generations
Have hands-on experience implementing LLMs and understand their capabilities and limitations
Life at TypeSafe
We’re a small, flat, close-knit team working to make intelligence dependable enough to become part of everyday software. We work fully in person from our San Francisco office near Embarcadero station. We love what we do and care deeply about the work.
We strive for excellence and craftsmanship and won’t stop until we get there. When the team wins, we all win, and we enjoy collaborating and inspiring each other to grow—as a team and as individuals.
We value emotional honesty, kindness, and bringing your whole self to work. We build machines; we don’t try to be machines.
We want TypeSafe to be the place where you do the most impactful work of your career and help define our future as a company.
We provide
Base salary of $150k–250k plus equity, based on leveling
100% covered health insurance
Daily lunch and dinner
Visa sponsorships
401K plans
Skills
As published by ashby · 9 questions · 2 written answers
Basics
Name, Email, Location, Resume/CV
Short answers (4)
- LinkedIn Profile optional
- Github Profile or Equivalent optional
- Website optional
- Why TypeSafe?
Pick from a list (3)
- Preferred Pronouns
- Work Authorization Status
- This role requires working in person five days per week at TypeSafe’s downtown San Francisco office. Are you willing to meet this requirement?
Written answers (2)
- What career accomplishments are you proud to share?
- What have you tried to get working technically with an LLM, and what limitations did you discover? optional
