Field note 01 / Phase one
Can handheld video replace teleoperation?
Adding 16 handheld demos to 16 teleop demos performed about as well as 32 teleop demos — 67% against 72%.
the baseline
+51 points over the baseline
what doubling teleop buys
01 / What we asked
Teleoperation is the bottleneck, not the algorithm.
Every demonstration costs a human at a leader arm, driving the robot one repetition at a time. It does not parallelise, and you need the robot before you can collect anything.
The alternative: film a person doing the task with a GoPro on a 3D-printed gripper, then retarget that motion onto robot joints. Roughly 10× cheaper per demonstration, and no robot required to collect.
The question: does 16 teleop + 16 handheld behave like 32 teleop? We also trained 16 teleop alone as the baseline — without it, a mixed policy scoring well says nothing about the handheld half.
GoPro + UMI gripper → Pose track
ArUco + SLAM → Retarget to SO-101
IK + clearance guard + replay validation → Train + evaluate
ACT, three policies
In plain language
- Teleop demo
- A human drives the robot through the task with a leader arm. The robot records its own joints. This is the standard way to collect training data, and the expensive one.
- Handheld demo
- A human does the task holding a GoPro on a 3D-printed gripper — no robot involved. Software then solves that motion onto robot joints.
- Placement
- A marked cell on the mat where the box starts. Fixed in advance, so every policy is graded on the same board. The target square never moves.
- What 16% vs 67% means
- Out of 100 attempts, the 16-teleop policy finished the task 16 times. Keep those 16 demos, add 16 filmed by hand, and it finishes 67 times.
02 / What it looks like
From a hand to a joint trajectory.
Thirty clips at once, then one of them in detail at three stages: the handheld GoPro capture, the same motion solved onto SO-101 joints, and the robot's own wrist camera as it replays it. Every frame is recorded data — none of it is an animation.
The middle panel is drawn from forward kinematics on the retargeted joint angles, not a physics simulation — it is what the arm was commanded to do
03 / Results
Handheld demonstrations did the work of teleoperation.
Live success rate on the box pick-and-place task. The comparison that matters is the first two rows: the same 16 teleop demos, with and without handheld footage added.
Read the first two rows together. They share their robot data exactly. The only difference between 16% and 67% is 16 demos that nobody teleoperated.
The third row is the price of the alternative. 16 more teleop demos — about 10 minutes each with reset and review — buy 72%. The handheld 16 reach 67%, for footage that cost no robot time at all.
The intervals are wide. They are not wide enough to make 16% and 67% the same number.
The gains landed where the robot was failing.
The same policies, split by where the box started: the 4 central placements against the remaining 16.
The teleop-only policy was not uniformly weak — in the middle it was fine. It collapsed at the edges: 1 success in 80 attempts outside the centre. That is a policy that has seen the task but not the workspace.
The handheld demos did not make an already-solved placement more reliable. They extended where the policy worked at all — from 1 in 80 to about 6 in 10. Coverage is the expensive thing to teleoperate and the cheap thing to film.
In that outer region, 16 handheld demos (48/80) do not quite match 16 more teleop demos (52/80). They buy the same kind of improvement — coverage — not a strictly better trade.
04 / Whether the pipeline holds up
Retargeting was never the thing that failed.
The obvious worry: that converting human video to robot joints is lossy, and quietly poisons a fraction of the training data.
It did not. Across every retargeted execution recorded for this work, no episode was discarded because the retarget failed. Episodes were discarded, but for ordinary bench reasons — a servo fault mid-run, a placement miss, a recording not started.
That distinction decides whether this scales. A retargeting failure is a research problem. A power fault is a shopping list, and a lighting failure is a checklist.
05 / What this does not show
One task, one arm, one room.
-
01
One setup
One arm, one object, one room. Contact-rich and deformable tasks are the next test.
-
02
Adjacent bars overlap
16% against 67% is a wide separation. 67% against 72% is not — those two should be read as close, not ranked.
-
03
Policies ran on separate sittings
Not interleaved placement by placement, so day-to-day variation in lighting and servo state is folded into every comparison. The next phase runs them back to back.
-
04
Scoring is human
The operator scores each attempt. An automated jaw check runs alongside and occasionally disagrees, usually when the gripper closes on an edge. We keep both.
-
05
Handheld is cheaper, not free
It still needs a calibrated rig, a mapping pass, and lighting QA. The claim is a much lower cost per demo, not zero.
-
06
The infrastructure was the result
Small amounts of motor drift moved the success rate more than the choice of policy did — we would have written that up as a training failure. Start-pose verification now runs before every block, and a clip becomes training data only once the physical arm has replayed it. Read these numbers as a statement about the rig as much as about the method.
-
07
What's next
16 handheld demos took a 16% policy to 67%. The next blocks walk the handheld count up to find where the curve flattens, which turns this into a budget a team can plan against. The evaluation tightens alongside it: policies run back to back on the same placements in one session, and the task moves beyond a rigid box on a flat table.
06 / Reproduce it
Every number here is checkable.
The policies, the training data and the raw per-attempt records are public. Success rates are not transcribed by hand — a verifier script recomputes each one from the operator's records, and it is in the repository below.
Policies
- act_so101_t16b_u16 The result. 16 teleop + 16 handheld — 65%
- act_so101_t16b 16 teleop, the baseline — 17%
- act_so101_t16b_t16 32 teleop, the matched-count control — 70%
Datasets
- so101_t16b_u16 16 teleop + 16 handheld episodes
- so101_retargeted_umi The handheld GoPro demos, retargeted into joint space
- All five datasets → Teleop, handheld and the 50-episode set
Code & records
- robotfuel_research Method, raw per-attempt records and the charts on this page
- verify_results.py Recomputes every published rate from the raw records