freehire launches on Product Hunt on 26 August.

Follow →

Multi-Agent RL Research Engineer

Open 30d

We are seeking a highly skilled and experienced AI Engineer specializing in Agent Behavior and Reinforcement Learning. This position demands deep expertise in autonomous agent architectures, a strong grasp of reinforcement learning principles and techniques, and proven ability to integrate complex AI systems like decision-making frameworks, behavior planning algorithms, multi-agent coordination, and optimized training pipelines. The ideal candidate is a proactive problem-solver and a dedicated builder, passionate about creating intelligent, adaptive, and believable AI agents for our XR simulation platforms.

About the client:


Our client is focused on building next-generation training systems powered by AI and immersive technologies. They develop solutions combining artificial intelligence, computer vision, extended reality, reinforcement learning, and embedded systems to support military training.

The team is small, about 15 people, with a flat structure and a culture that values ownership, autonomy, and long-term impact.

The company is currently working on multi-year projects related to U.S. Air Force contracts and U.S. Marine Corps programs. Candidates should be supportive of the defense mission and comfortable in a fast-paced startup environment.

What you will work on in this role:

  • Large-scale multi-agent reinforcement learning: cooperative team behavior and competitive self-play.
  • A GPU-accelerated parallel-simulation training pipeline (NVIDIA Isaac Lab/Gym, or JAX-based sims such as Brax / MuJoCo MJX), with trained policies deployed into the Unity environment at behavioral parity.
  • Grounding NPC behavior in realistic tactics via imitation learning / learning-from-demonstration on real combat footage.
  • Real-time opponent adaptation (opponent modeling / meta-RL), bounded to human-plausible movement.

Must have skills:

  • Multi-agent RL, both cooperative and competitive / self-play (e.g., MAPPO, QMIX, PPO self-play).
  • Large-scale or GPU-parallel RL training experience (Isaac Lab/Gym, JAX sims, or comparable).
  • Strong PyTorch and/or JAX.
  • Imitation learning / behavior cloning / learning from demonstration.

Strong plus:

  • Population-based training or self-play league management (AlphaStar / OpenAI Five lineage).
  • Meta-RL, fast online adaptation, or opponent modeling.
  • Policy export / sim-to-deployment into a game engine (Unity ML-Agents or custom).

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available