Member of Technical Staff Safety
Summary
Owns red-teaming and adversarial evaluation pipelines for Reflection's AI models — probing for security, misuse, and alignment failures, implementing jailbreaks and defenses, building automated safety benchmarks, and validating model releases against risk thresholds. Core focus: LLM safety and ML evaluation systems.
You will own red-teaming and adversarial evaluation pipelines for AI models. You will identify security, misuse, and alignment failures; develop automated safety benchmarks; implement jailbreaking techniques and defenses; translate findings into guardrails; and validate releases against risk thresholds.
Responsibilities
- Own red-teaming and adversarial evaluation pipelines
- Probe models for security misuse and alignment failure modes
- Translate safety findings into concrete guardrails
- Validate releases against safety risk thresholds
- Develop scalable automated safety benchmarks
- Research and implement jailbreaking techniques and defenses
Requirements
- Graduate degree in Computer Science Machine Learning or a related discipline or equivalent AI safety experience
- Knowledge of LLM safety adversarial attacks red-teaming methodologies and interpretability
- Software engineering experience building automated evaluation pipelines or large-scale ML systems
- Ability to make high-stakes model release and safety-threshold decisions
Benefits
- Stock options
- Medical dental vision and life insurance
- Annual wellness allowance
- Daily office lunch and dinner
- 22 weeks of paid parental leave
- Unlimited paid time off in the United States
- 30 days of vacation in the United Kingdom
- Visa sponsorship support
- Regular off-sites happy hours and team celebrations
