AI Research Scientist, Reinforcement Learning (LLM) and Post-Training
Summary
The AI Research Scientist will develop reinforcement learning methods for post-training large language and code models, including designing reward models and conducting training experiments. The role involves analyzing model failure modes and collaborating on scalable RL infrastructure.
- Analyze failure modes reward hacking and instability
- Collaborate to scale training with RL infrastructure
- Define interfaces for rollout generation and logging
- Design reward models and training curricula
- Develop reinforcement learning methods for post training large language models and code models
- Publish research at top academic venues
- Run off policy and on policy training experiments