Robot retargeting · August 7, 2026
What 81 robot mappings taught us about better robot data.
We took one measured human manipulation trajectory, mapped it to a five-joint SO-101 robot in 81 different ways, and inspected every frame. The encouraging result was not simply that the robot could reach the path. It was that the experiment told us exactly which frames fit the digital embodiment, which frames needed work, and what to improve before using the motion in a learning or hardware loop.
Human experience is promising. A robot still needs the right translation.
Human demonstrations can cover more objects, environments, and strategies than a small number of robots can collect alone. But the same motion does not automatically fit a new embodiment. A five-joint arm has different reach, orientation freedom, singularities, and continuity constraints than the human or capture device that produced the source trajectory.
We treated that gap as a measurable experiment. The input was one 400-frame episode from the public Universal Manipulation Interface cup dataset. The target was a pinned SO-101 model. Before running the full review, we fixed 81 candidate combinations of robot placement, motion scale, and orientation.
Every candidate was evaluated against the same checks. We retained failed frames instead of smoothing them away or changing thresholds after seeing the result.
How to read the result
- Each mapping evaluates the complete episode.
- The sweep tests all 400 source frames under each of 81 robot setups. A setup changes the arm's placement, starting pose, motion scale, and orientation relative to the source trajectory. Evaluating every setup across every frame produces 32,400 mapping-frame evaluations and exposes whether a result depends on one favorable arrangement.
- Measured cadence determines the implied motion.
- The source cadence is 59.94 Hz, so the 400-frame episode spans 6.67 seconds. Joint speed and continuity checks depend on that timing: using an assumed frame rate would change the implied motion and could make an infeasible trajectory appear acceptable.
- Feasibility depends strongly on robot setup.
- Thirteen of 81 setups preserved both position and tool direction across the entire sequence. The remaining 68 failed at least one of those checks somewhere in the episode. Placement, scale, and orientation therefore determine how much of the same measured motion transfers to the target embodiment.
- Rejected frames are diagnostic data.
- In the selected mapping, 293 of 400 frames passed every configured check. The other 107 remain in the result with their failure causes. That distinction shows whether the next intervention should change task framing, robot placement, trajectory continuity, or the source demonstration itself.
- Retargeting qualifies data; it does not produce a policy.
- This experiment converts a measured human path into candidate robot joint trajectories and checks their digital feasibility. It does not include visual feedback, adaptation to a moved object, learned decision-making, or physical execution. Those require a separate robot-learning and hardware evaluation.
Inspect the selected digital robot mapping.
Play, pause, or scrub all 400 measured frames. Green frames passed every configured digital check; purple frames were retained for review.
Data-driven 3D playback. Not physical execution or a learned policy.
Video transcript and frame-status sequence
The silent visualization uses a fixed virtual camera. On the left, a seven-point digital SO-101 link model moves across a floor grid while its tool path accumulates behind it. On the right, the current frame, aggregate check totals, and the four per-frame checks update as the video advances. Green means the frame passed every configured digital check; purple means one or more checks were retained for review.
Qualified frame ranges, using the one-based frame numbers shown in the video: 1–43, 99–143, 152–284, 296–341, and 375–400. Flagged ranges: 44–98, 144–151, 285–295, and 342–374.
The selected mapping solves position for 400/400 frames, tool direction for 351/400, continuity for 343/399 applicable transitions, conditioning for 346/400, and all configured checks together for 293/400 frames.
The source motion contained a strong transferable core.
The selected mapping solved position for all 400 frames. It preserved the configured tool direction for 351 frames, continuity for 343 of 399 applicable transitions, and task conditioning for 346 frames. When all configured checks were combined, 293 frames—73.25% of the episode—remained qualified.
Across the full sweep, 13 mappings preserved both position and tool direction for the entire sequence. That means the source was not broadly unusable. It contained meaningful robot-compatible structure, but different choices exposed different weaknesses.
No mapping was promoted to physical replay. The remaining digital failures were preserved, and physical calibration, collision, controller, payload, and safety checks were still outside this experiment. Those preserved failures make the result more useful by locating what requires adjustment before hardware.
The target path stayed within the selected arm placement's positional reach.
Most frames preserved the intended approach axis, with exact gaps identified.
Most neighboring solutions stayed coherent; abrupt changes remain visible.
The audit isolated frames where the task became poorly conditioned.
Better robot learning starts before training—with data that knows its embodiment.
Qualification can increase usable-data yield
Instead of treating every demonstration as equally trainable, the pipeline can identify compatible segments, flag repairable gaps, and avoid spending robot time on obviously mismatched motion.
Failure labels can guide the next collection
Direction, continuity, and conditioning fail for different reasons. Those labels tell us whether to change task framing, placement, capture behavior, retiming, or the target embodiment.
Learning comparisons become measurable
Once the motion is calibrated and physically validated, we can compare a robot-native baseline against a human-data-assisted condition on held-out task success, teleoperation minutes, interventions, cycle time, and recovery.
Use the audit to prepare task-specific capture.
A hardware-ready capture packet includes mapping, gripper calibration, and several demonstrations of one defined task without changing camera or mount geometry.
Those measurements make it possible to verify the physical task axis, reduce the motion envelope, retime the trajectory, and add collision and controller checks before any supervised hardware test.
The resulting evaluation can compare a robot-native baseline with a human-data-assisted condition on held-out task performance, robot-native demonstrations, and human correction time.