freehire launches on Product Hunt on 26 August.

Follow →

Machine Learning Engineer (Egocentric 3D Human Pose)

Summary

Machine Learning Engineer developing 3D human body/hand pose estimation models from egocentric video using PyTorch/TensorFlow, camera calibration pipelines, and deep learning to generate robot training data.

About the Role

We are looking for a Machine Learning Engineer to join our core research and development team focused on recovering accurate 3D human body and hand motion from egocentric (first-person) video.

Human demonstration data is the foundation of robot learning, and its quality depends on accurately reconstructing human motion. In this role, you will develop models and production pipelines that transform head-mounted and body-mounted camera streams—including wide-FOV, stereo, motion-blurred, and heavily self-occluded video—into metrically accurate, temporally consistent 3D pose representations for robot policy training and human-to-robot motion retargeting.

You will work across the entire perception stack, including camera calibration, data annotation, model training, evaluation, and large-scale deployment. This role is ideal for engineers with strong expertise in both computer vision and deep learning who enjoy solving challenging real-world perception problems.

Responsibilities

  • Develop state-of-the-art 3D body and hand pose estimation models for egocentric video using monocular and stereo camera systems.
  • Build models for 2D/3D keypoint estimation, SMPL/SMPL-X, MANO, and full-body motion reconstruction.
  • Address challenging egocentric vision problems, including severe self-occlusion, motion blur, rolling shutter artifacts, truncated limbs, extreme viewpoints, and hand-object interaction.
  • Design and maintain camera geometry and calibration pipelines, including fisheye and wide-FOV camera models, stereo calibration, triangulation, and coordinate frame alignment.
  • Improve temporal consistency and physical plausibility using filtering, kinematic constraints, multi-view fusion, and multi-modal sensor integration.
  • Build scalable annotation, evaluation, and quality assurance pipelines for large-scale human motion datasets.
  • Optimize large-scale model training and high-throughput inference pipelines for production environments.
  • Collaborate closely with robotics engineers to convert reconstructed human motion into high-quality robot training data.
  • Contribute to system architecture, engineering best practices, and the long-term evolution of the perception platform.

Minimum Qualifications

  • Bachelor's, Master's, or PhD in Computer Science, Machine Learning, Computer Vision, Robotics, or a related field.
  • 3+ years of experience building and deploying machine learning systems.
  • Hands-on experience with 3D human pose estimation, hand pose estimation, or human motion tracking from video.
  • Strong understanding of multi-view geometry, camera calibration, triangulation, coordinate transformations, and projection models.
  • Strong Python programming skills and proficiency with PyTorch or TensorFlow.
  • Solid knowledge of modern deep learning techniques, model training, evaluation, and production ML workflows.
  • Strong analytical and problem-solving skills with the ability to thrive in a fast-paced collaborative environment.

Preferred Qualifications

  • Experience with egocentric perception systems, AR/VR headsets, smart glasses, or wearable capture rigs.
  • Expertise with SMPL, SMPL-X, MANO, inverse kinematics, markerless motion capture, or hand-object pose estimation.
  • Familiarity with egocentric vision datasets and benchmarks.
  • Experience with modern video and 3D learning architectures, including Video Transformers, diffusion-based motion models, or 3D CNNs.
  • Experience building multi-camera capture systems, synchronization, and calibration infrastructure.
  • Experience with human-to-robot motion retargeting, teleoperation, imitation learning, or dexterous manipulation.
  • Experience developing annotation tools, active learning pipelines, or large-scale data quality systems.
  • Publications at leading conferences such as CVPR, ICCV, ECCV, NeurIPS, SIGGRAPH, or 3DV, open-source contributions, or demonstrated impact in applied AI systems.

What We Offer

  • Competitive salary and equity package
  • Comprehensive medical, dental, and vision insurance
  • 401(k) retirement plan
  • Generous paid time off and company holidays
  • Paid sick leave
  • Opportunity to work alongside leading researchers and engineers in robotics, computer vision, and AI
  • High-impact role building cutting-edge perception systems for next-generation robotics

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available