freehire launches on Product Hunt on 26 August.

Follow →

Technical Program Manager

Summary

Own end-to-end delivery of RL environments that stress-test frontier AI agents; manage scope changes, QA, and customer handoffs while staying hands-on with code and data.

About Patronus AI

Patronus AI is a frontier lab developing simulation research and infrastructure to accelerate progress toward human-aligned AGI. We are on a mission to simulate all of the world’s intelligence.

We are the team behind some of the earliest and most influential research in AI evaluation like FinanceBench, Lynx, SimpleSafetyTests, CopyrightCatcher, Humanity’s Last Exam, and more. We are formerly AI researchers and engineers from companies like Meta AI, Amazon AGI, and Google. Our customers include foundation model labs and Fortune 500 enterprises like Adobe. We are backed by top-tier investors like Lightspeed Venture Partners, Notable Capital, Stanford University, Noam Brown, Gokul Rajaram, and more.

Responsibilities

As a Technical Program Manager at Patronus AI, you will run the production line that turns frontier-lab demand into delivered RL environments. Our customers commission simulations of real applications and workflows, and you own the path from signed work order through SME sourcing, environment build, task generation, QA, and delivery.

Like all cutting edge fields, the state of the industry and the work changes shape constantly. App lists get swapped after a work order is signed, difficulty definitions get renegotiated a week before delivery etc. Your job is to keep the plan honest through all of it: who owns what, what is due when, and whether a delivery is actually ready to ship. When scope moves, you move the plan with it.

This is not a coordination-only role. We believe you should not manage work you don't understand. You will stay in the details, reading QA feedback rows, poking at environments, and sanity-checking task quality, and you are expected to pitch in directly where it unblocks the team. You will own identifying opportunities for automation and build them yourself.

Your work will help frontier labs stress-test and improve the next generation of AI agents, advancing progress toward safe, human-aligned general intelligence.

In this role, you will:

  • Own delivery programs end-to-end: track work orders, deliverables, due dates, and owners across engineering, QA, and SME teams, and maintain the consolidated view the rest of the company plans against.
  • Manage scope changes mid-program. A previously committed app list doubles in size, or a customer redefines task difficulty on an active work order, and you update the plan, the pricing inputs, and our commitments without losing the thread.
  • Run the handoffs between GTM, SME sourcing, environment engineering, task generation, and QA. This is where work gets stranded today. Build the process that stops that: clear entry and exit criteria, and visibility into who is staffed on what.
  • Define what "ready to ship" means and hold the line on it. QA tickets are green before anything is marked complete, and tasks pass in the customer's harness, not just locally. We deliver work that already passes rather than delivering and triaging after.
  • Work directly with customers alongside account leads. Turn their asks into plans with clear outcomes, owners, and dates, and keep them current on progress, risks, and timeline confidence. Respond quickly; customers notice when we don't.
  • Get teams to time-bound their work and make estimates explicit ("expect this to take X hours, tell me if it takes longer"), then hold everyone, including yourself, to them.
  • Identify opportunities for automation and process improvements across the delivery pipeline and build them yourself.
  • Stay hands-on. Run agents in environments, review QA feedback and trajectories, and build small tools (trackers, scripts, Claude skills) that scale your own function.

Qualifications

"The number one qualification to succeed in this machine learning course is gumption" - John Lafferty, CS Professor at Yale

Above all, we look for a proactive mindset, willingness to learn, unlimited energy, and relentless optimism. You are a great fit if you have a background in the following:

  • BS, MS, or equivalent experience in Computer Science, Engineering, or another technical / quantitative field, with 3+ years as a technical program manager, delivery lead, engineering program manager, or in a similar role on complex, multi-team technical programs.
  • A track record of shipping when requirements move after kickoff, without letting quality or the customer relationship slip.
  • Strong technical fluency, including comfort using AI tools, reading or reviewing code, and analyzing data or model outputs. You should be able to go deep enough into the work to have credibility with engineers.
  • Excellent organization and execution skills, with the ability to manage tasks, timelines, quality reviews, customer requirements, and cross-functional stakeholders. You are the person who always knows the current state.
  • Clear written and verbal communication skills, including the ability to translate customer asks into concrete plans with owners and dates.
  • Strong eye for quality and detail, with a bias toward catching edge cases, inconsistencies, and subtle failure modes before the customer does.

Nice to have:

  • Experience with reinforcement learning, agent evaluation, RL environments, or human-data / SME-sourced data pipelines.
  • Experience delivering to frontier labs or other research-driven customers with short turnaround expectations.
  • Experience standing up QA or review processes for software or data deliverables.

To support close collaboration, this role is based in our San Francisco headquarters and requires in-office attendance 5 days a week.

The expected base salary range for this role is $125,000 - $250,000 USD. In addition to base salary, we offer equity and benefits. Actual compensation will be determined based on experience, qualifications, skills, and location.

Benefits

  • Competitive salary and equity packages
  • 15 days of paid vacation per annum
  • Parental & sick leave
  • Health, dental, and vision insurance plans
  • 401(k) plan + matching
  • In-office lunch & dinner
  • Sponsored personal tax accounting
  • Whoop band
  • Monthly meal stipend
  • Monthly health and wellness stipend
  • Equinox membership
  • Fun global offsites!

Patronus AI is an equal opportunity employer. We celebrate diversity in our workplace, and all qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or other legally protected characteristics.

By clicking ‘Apply’, you agree to Greenhouse's Terms of Service and Privacy Policy.

By clicking 'Apply', you agree to Patronus AI, Inc. Privacy Policy.

What this application asks

greenhouse

First Name, Last Name, Email, Phone, Resume/CV, Cover Letter, Location

  • Preferred First Name optional
  • LinkedIn Profile
  • Twitter / X account
  • Are you currently authorized to work in the United States without requiring employer sponsorship, now or in the future? choose one
  • This role is based in our San Francisco office and requires working onsite five days per week. Are you able to meet this requirement? choose one
  • If you would need to relocate to the San Francisco Bay Area, when would you be available to relocate? choose any
  • Our interview process typically takes about two weeks from initial screen to final decision. Does this timeline align with your availability? choose one
  • If selected, when would you be available to start?
  • Website optional
  • I consent to the processing of my personal data by Patronus AI, Inc for this application and for the purpose of being considered for future career opportunities. I understand that my data will be retained in the talent pool for 24 months, and I can request deletion at any time by contacting [email protected]. choose one
  • How did you find out about Patronus AI? choose any
  • Describe a program where the requirements changed significantly after kickoff. What did you change, and what did the customer see? (150 words)
  • What is something repetitive in your last role that you automated yourself? (100 words)

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available