Research and build continuous learning systems for self-improving AI agents on the OpenPipe team. You'll investigate RLHF, reward modeling, and on-policy distillation to solve production bottlenecks in agent training. Core stack uses PyTorch/JAX for model training with Kubernetes and Megatron for distributed GPU infrastructure.
Sign in to see your match