freehire launches on Product Hunt on 26 August.

Follow →

Student Researcher (LLM Post Training – Agent & Reinforcement Learning) - 2026 Start (PhD)

Summary

Researcher optimizing large language models via post-training techniques like RLHF/DPO to improve reasoning, coding, and math while exploring future use cases.

- Conduct preference alignment - Construct training data - Explore large-scale models - Improve coding capability - Improve math capability - Improve reasoning capability - Optimize multimodal large models - Optimize post training systems - Perform instruction tuning - Research future model use cases Perks/Benefits: - Health insurance - Housing allowance - Life insurance - Paid Holidays - Paid sick time - Wellbeing benefits

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available