freehire launches on Product Hunt on 26 August.

Follow →

Research Intern — Human Pose Understanding and Vision-Language Models

Open 36d

Summary

Research intern in computer vision and vision-language models, focusing on hand/body pose tracking for Apple devices like Vision Pro, using deep learning frameworks.

We are looking for a research intern to join us for a research project aimed at publication at a top-tier venue. The intern will design and develop novel systems that explore the interaction between human pose understanding and vision-language models (VLMs), advancing how these modalities can be combined to reason about human motion, activity, and embodied behavior across images and video.

Our group develops hand and body pose tracking algorithms for various apple devices and applications. One such example includes the hand tracking input for the Vision Pro.

Minimum Qualifications

  • Currently enrolled in a graduate program (M.Sc. or Ph.D.) in Computer Science, Electrical Engineering, or a related field
  • Publications at top-tier venues (e.g., NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, ACL, EMNLP or similar)
  • Strong programming skills in Python and experience with deep learning frameworks (e.g., PyTorch)
  • Solid foundation in computer vision, natural language processing, or multimodal learning

Preferred Qualifications

  • Demonstrated expertise working with Vision-Language Models (VLMs) and/or Large Language Models (LLMs)
  • Experience with human pose estimation, motion modeling, or related body-tracking tasks
  • Familiarity with video understanding tasks and temporal modeling
  • Familiarity with multimodal learning and benchmarks that combine language with visual or spatial data
  • Experience with prompt engineering and optimization techniques

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available