AI Red Teamer (LLM Generalist)
AI Red Teamer (LLM Generalist)
Location: Seattle, WA (candidates must reside in the Seattle metro area or be willing to relocate prior to start)
Type: Contract, 40 hours per week
About the Role
As an AI Red Teamer, you will stress-test large language models by intentionally trying to break them. Rather than checking whether an answer is correct, you will design creative, adversarial prompts that expose vulnerabilities: unsafe content, bias, broken guardrails, hallucinations, prompt injection weaknesses, and unexpected behaviors. Your work directly supports AI safety and model robustness for leading research labs.
This is a generalist red teaming role. You will probe models across the full spectrum of risk categories, including content safety, CBRN (chemical, biological, radiological, nuclear), cybersecurity, persuasion and influence operations, child safety, self-harm, over-companionship, and regulatory compliance. Red teaming may span text, image, voice, and agentic model capabilities depending on project needs.
This role requires creativity, curiosity, and an ability to think like an adversary while operating with strong ethical judgment.
Day-to-Day Responsibilities
Craft creative prompts and multi-turn scenarios to stress-test AI guardrails across diverse risk categories
Discover ways around safety filters, restrictions, and defenses using jailbreak, evasion, and prompt injection techniques
Explore edge cases to provoke disallowed, harmful, or incorrect outputs
Evaluate and score model responses against structured harm taxonomies and severity rubrics
Document experiments clearly, including what you tried, why you tried it, and what it revealed
Review and refine adversarial prompts generated by other team members
Contribute to harm taxonomy development, calibration exercises, and inter-rater reliability work
Collaborate with engineers, data scientists, and researchers to share findings and strengthen defenses
Work with potentially disturbing content on a regular basis (see Content Warning below)
Stay current on jailbreaks, attack methods, and evolving model behaviors
Desired Capabilities
Core
Strong hands-on experience using multiple LLMs (ChatGPT, Claude, Gemini, open-source models, etc.)
Intuition for crafting adversarial prompts; familiarity with jailbreak or evasion techniques is a strong plus
Creative, adversarial problem-solving skills
Clear and thoughtful written communication
Strong ethical judgment and the ability to separate adversarial thinking from personal values
Self-directed, collaborative, and comfortable in feedback-heavy environments
Curiosity, persistence, and comfort with frequent failure in experimentation
Nice to Have
Familiarity with Python or other scripting languages
Experience working with LLM APIs or evaluation tooling
Comfort with structured data annotation and rubric-based scoring
Prior work in trust and safety, content moderation, QA, or security research
Subject matter expertise in any high-risk domain (cybersecurity, chemistry, biology, medicine, law, finance, etc.)
You Will Thrive Here If
You treat every model response as a hypothesis to challenge
You can switch between creative free-association and rigorous documentation in the same session
You go deep into unusual interests (fandoms, niche internet cultures, gaming exploits, Wikipedia rabbit holes, etc.)
You come from a creative background: writing, visual art, improv, puzzle design, or similar
You are energized by finding the thing nobody else thought to try
You are genuinely passionate about AI and follow the space closely
Content Warning
This role involves regular and deliberate exposure to harmful content. You will encounter and intentionally generate content involving violence, self-harm, hate speech, sexually explicit material, child safety scenarios, and other categories of harmful output as part of structured adversarial testing. Candidates must be able to engage with this material professionally and sustainably. Support resources are available.
About Handshake AI
Handshake AI partners with leading AI research labs to make models safer and more robust. Our red teaming operations help identify vulnerabilities before they reach users, contributing directly to the responsible development of frontier AI systems.
Skills
As published by ashby · 13 questions · 4 written answers
Basics
Name, Email, Resume, Location
Short answers (4)
- Name Pronunciation optional
- Phone number
- LinkedIn Profile optional
- When could you start?
Pick from a list (5)
- Can you work onsite in Seattle, Monday through Friday, 40 hours per week?
- Have you reviewed the Content Warning and are you willing to perform the work described?
- Are you authorized to work lawfully in the United States for Handshake?
- Will you now, or in the future, require sponsorship for employment visa status?
- Which relevant experience do you have? Select all that apply.
Written answers (4)
- Share any relevant work samples, projects, or writeups. Please include links or a brief description; do not include confidential material. optional
- Which AI models have you used regularly, and what have you used them for? (100 words maximum)
- Describe a time an AI model behaved unexpectedly or failed. What were you trying to do, what happened, and how did you investigate it? Personal experimentation counts. (150 words maximum)
- What is an unusual topic, hobby, or community you know deeply? How could that knowledge help you spot failures other testers might miss? (150 words maximum)