Point your AI agent at freehire and let it find you a job.

Get the CLI →

persona.ai

NewBe an early applicant

AI Engineer – Robotics Data Preprocessing

Posted Updated
Discussion

Summary

An AI Engineer who builds the data preprocessing pipeline for humanoid robots: converting raw multimodal recordings (egocentric video, force-torque, IMU) into training-grade datasets via cross-modal validation, kinematic retargeting, and augmentation. Core stack is Python and PyTorch with video processing tools like OpenCV and FFmpeg.

Job Title: AI Engineer – Robotics Data Preprocessing

Department: Software

Reports To: Teleoperations Lead

Employment Type: Full-Time

Location: Houston, TX or Pensacola FL

Who We Are

Persona AI is building humanoid robots for the most demanding environments in heavy industry — shipyards, steel mills, fabrication facilities, and offshore platforms — performing welding, grinding, maintenance, inspection, and material-handling work that is dangerous, physically demanding, and increasingly difficult to staff.

We are backed by leading strategic and financial investors and engaged with global industrial leaders across Korea, Japan, the United States, and Singapore. Korea is the center of gravity for our early commercial strategy, anchored by relationships with the world’s leading shipbuilders and steelmakers. Our work spans both the robot platform itself and the systems, partners, and playbooks required to deploy it at scale.

Why Join Persona AI?

  • We offer competitive compensation, a performance-based bonus, 99% employer covered medical benefits, early-stage equity, competitive PTO, and a company-wide paid winter break between December 24th and January 2nd.

  • You’ll shape technology that’s redefining the possibilities of robotics and human interaction.

  • Work alongside passionate teammates who value creativity, and continuous learning.

  • Enjoy full access to advanced tools,

About the Role

At Persona we require an unprecedented volume of high-quality, multimodal data. We are moving beyond basic teleoperation to leverage massive datasets of in-the-wild egocentric video combined with dense sensor streams (IMU, haptics, kinematics, and high-fidelity force profiles). We are seeking a highly skilled AI Engineer to architect the systems that turn this raw, unstructured multimodal data into high-fidelity training assets for our robots.

Models are only as good as the data they learn from. In humanoid robotics, that's not a platitude, it's the bottleneck. There's no Internet-scale corpus of robots manipulating the physical world. We have to create it. That's this role.

As an AI Data Engineer, you sit at the most leveraged point in our entire training pipeline: every model we ship is downstream of the data you build. If this role succeeds, our foundation models learn dexterity faster than anyone else's. If it fails, nothing else matters.

You will architect and scale the infrastructure that turns raw, messy reality into training-grade data, extracting, augmenting, and aligning human dexterous manipulation data from massive multi-sensor and egocentric video datasets. You'll build advanced pre-processing algorithms that recover what sensors can't directly see: quantifying grasp dynamics from force-torque signals, estimating contact forces from visual cues alone, reconstructing heavily occluded hand poses, and lifting 3D geometry out of 2D frames.

And because every minute of teleoperation data is expensive, you'll make each one count: using spatial, temporal, and cross-modal augmentation to multiply the value of everything our collection team captures. Your work directly determines how fast our models learn and how far they can go.

What You Will Be Doing

  • Force Analysis & Hidden State Inference: Design cross-modal validation systems that verify video, proprioception, force/haptic signals, and language annotations agree with each other, e.g., reprojecting robot state into the image plane to confirm video–state consistency, and VLM-assisted checks that instructions match observed behavior.

  • Kinematic Retargeting & Alignment: orchestrating hand-tracking, segmentation, depth estimation, 3D reconstruction, and pose-tracking modules; retargeting human demonstrations into robot trajectories; and running simulation-in-the-loop validation (kinematic feasibility, physics replay, motion-consistency filtering) so synthesized data is physically grounded, not just visually plausible.

  • Advanced Data Augmentation: Implement robust data augmentation strategies (spatial transformations, temporal scaling, synthetic viewpoints, and sensor noise injection) to expand expert trajectories and improve the robustness of our learning models.

  • Teleoperation Synchronization: unified state–action representations across differing embodiments, coordinate frames, rotation conventions, gripper/hand parameterizations, and sampling rates, with per-dimension validity masking and per-source normalization so that adding a new robot or sensor is a configuration change, not a rewrite.

  • Close the loop with data consumers: build the tooling that lets researchers query, visualize, and audit datasets (clip browsers, trajectory viewers, annotation review UIs), and turn model-failure analyses into new curation rules and targeted re-collection requests.

  • Multimodal Data Pipelines: Architect end-to-end ingestion pipelines that take raw, unstructured recordings (egocentric video, teleoperation sessions, third-party open datasets) and produce indexed, queryable, training-ready datasets. This includes temporal segmentation of long recordings into action clips, metadata and scene-graph extraction, embedding-based retrieval, and language annotation workflows.

What We Are Looking For

  • Education: M.S., or Ph.D. in Computer Science, Data Engineering, Machine Learning, Robotics, or a related field.

  • Programming & ML Frameworks: Deep expertise in Python and extensive experience with PyTorch, specifically in handling custom dataloaders for multimodal datasets.

  • Force & Time-Series Data Processing: Experience analyzing and processing complex time-series data from force-torque (F/T) sensors, load cells, or tactile arrays, ensuring pristine alignment with visual frames.

  • Video Processing Expertise: Mastery of video processing pipelines and libraries (OpenCV, FFmpeg, Decord) and managing the I/O bottlenecks of terabyte-scale video datasets.

  • Solid working knowledge of 3D geometry and robotics data: coordinate frames and transforms, rotation representations, camera intrinsics/extrinsics, forward/inverse kinematics, URDF.

  • Data Augmentation: Proven ability to implement programmatic and generative data augmentation techniques for computer vision and time-series data.

Bonus Skills

  • Experience with NVIDIA’s robotic software stack (Open X-Embodiment, DROID, AgiBot World, EgoDex, or similar).

  • Familiarity with the modern perception toolbox as a user: segmentation (SAM-family), monocular depth, hand/body pose estimation (MANO/SMPL), 6-DoF object pose tracking, point tracking—you don't need to train these models, but you should be comfortable composing and evaluating them in a pipeline

  • Familiarity with distributed data processing systems (Ray, Apache Spark) for cluster computing.

  • Background in generating or utilizing synthetic robotic data via simulation (Omniverse, MuJoCo).

  • Experience integrating spatial awareness or tactile data representations (e.g., Fourier encoding) into visual pipelines.

Persona AI is an Equal Opportunity Employer.

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, age, disability, veteran status, or any other characteristic protected by applicable federal, state, or local law.

Skills

See also

AI Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available