Member of Technical Staff, Embedded Assessments
About METR
We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation and misalignment.
We believe it is robustly good for policymakers and civil society to have a clear understanding of risks from AI systems, and we are extremely excited to build a team of ambitious, excellent people to tackle one of the most important challenges of our time.
Overview
METR has started embedding researchers inside frontier labs to investigate incidents, stress-test labs’ internal agent monitoring systems, and assess loss-of-control risks from internal deployment. As agent capabilities increase, we expect this to be one of the most important sources of independent information the world has about catastrophic risks from advanced AI.
As the source of information becomes more important, we'll need many more talented researchers and engineers who can conduct embedded exercises. We expect these assessors to have deep access, and for their work to be a large part of METR's impact in the next year. We want to build on the momentum from previous exercises to further develop our risk assessments.
What the job looks like
You'll be embedded in a frontier AI lab for up to several weeks at a time, likely alongside 1-4 other METR staff. Between exercises, you'll practice, develop the general methodology, talk to other researchers, build tooling to make future exercises go better, help us hire and scale, write up results, and plan/coordinate future exercises.
We are looking for embedded researchers and engineers across multiple current and potential future exercises: AI R&D acceleration assessment, compute allocation, monitorability red-teaming, incident investigations, and more. We expect this role to evolve significantly as we develop and prototype this new form of risk assessment.
Ideal candidate
We are looking for candidates with at least two of the following skills, although the ideal candidate will have most or all of them:
-
You’re good at prompting LLMs, are familiar with research relevant to the fields above (or can get up to speed quickly).
-
You have some security experience.
-
You can get spun up on large codebases quickly.
-
You’re scrappy, creative, independent, and self-directed (because during the exercises you’ll only have a few other METR employees you can talk to). The work is novel and you’ll need to figure a lot of stuff out on the fly largely by yourself.
-
You are excellent at loss-of-control threat modeling and breaking down safety cases.
-
You’re good at verbal communication, writing, and stakeholder management (e.g., navigating complex relationships with frontier labs).
-
You’re trustworthy and have a good reputation.
As published by lever
Resume/CV, Full name, Email, Phone, Current location, Current company, Linkedin URL, GitHub URL, Personal website URL
- Are you currently authorized to work in the United States? (We are a cap-exempt nonprofit and can sponsor visas.) choose one
- Where are you currently based? written answer
- Are you happy to relocate to, or already in, the Bay Area and excited to work in-person at our office in Berkeley? choose one
- When is the earliest you would want to start working with us? optional
- How did you hear about this role? choose one · optional
- Do you have an accelerated timeline by which you need a decision (e.g., because of a competing offer)? If so, describe your preferred timeline here. If not, leave this blank. written answer · optional
- What information about your application would you like us to share with similar organizations? choose one · optional
- What are your 2–5 most impressive achievements? written answer · optional
- What's the best evidence that you'd potentially be great in this position? written answer
- Why METR? written answer