Field note 01 / Phase one

Can handheld video replace teleoperation?

Adding 16 handheld demos to 16 teleop demos performed about as well as 32 teleop demos — 67% against 72%.

SO-101 · 5-DOF ACT policy Three policies Sep 2026
16% 16 teleop demos
the baseline
67% 16 teleop + 16 handheld
+51 points over the baseline
72% 32 teleop demos
what doubling teleop buys

Teleoperation is the bottleneck, not the algorithm.

Every demonstration costs a human at a leader arm, driving the robot one repetition at a time. It does not parallelise, and you need the robot before you can collect anything.

The alternative: film a person doing the task with a GoPro on a 3D-printed gripper, then retarget that motion onto robot joints. Roughly 10× cheaper per demonstration, and no robot required to collect.

The question: does 16 teleop + 16 handheld behave like 32 teleop? We also trained 16 teleop alone as the baseline — without it, a mixed policy scoring well says nothing about the handheld half.

Handheld capture
GoPro + UMI gripper
→ Pose track
ArUco + SLAM
→ Retarget to SO-101
IK + clearance guard + replay validation
→ Train + evaluate
ACT, three policies

In plain language

Teleop demo
A human drives the robot through the task with a leader arm. The robot records its own joints. This is the standard way to collect training data, and the expensive one.
Handheld demo
A human does the task holding a GoPro on a 3D-printed gripper — no robot involved. Software then solves that motion onto robot joints.
Placement
A marked cell on the mat where the box starts. Fixed in advance, so every policy is graded on the same board. The target square never moves.
What 16% vs 67% means
Out of 100 attempts, the 16-teleop policy finished the task 16 times. Keep those 16 demos, add 16 filmed by hand, and it finishes 67 times.

From a hand to a joint trajectory.

Thirty clips at once, then one of them in detail at three stages: the handheld GoPro capture, the same motion solved onto SO-101 joints, and the robot's own wrist camera as it replays it. Every frame is recorded data — none of it is an animation.

The middle panel is drawn from forward kinematics on the retargeted joint angles, not a physics simulation — it is what the arm was commanded to do

Fig. 1 — capture, retarget, replay Recorded data throughout — no animation

Handheld demonstrations did the work of teleoperation.

Live success rate on the box pick-and-place task. The comparison that matters is the first two rows: the same 16 teleop demos, with and without handheld footage added.

Live success rate for three policies: 16 teleop at 16%, 16 teleop plus 16 human at 67%, and 32 teleop at 72%.
Live success rate by training mix. Each bar is 100 attempts on the arm. Whiskers are Wilson 95% intervals — the range the true rate is likely to sit in.

Read the first two rows together. They share their robot data exactly. The only difference between 16% and 67% is 16 demos that nobody teleoperated.

The third row is the price of the alternative. 16 more teleop demos — about 10 minutes each with reset and review — buy 72%. The handheld 16 reach 67%, for footage that cost no robot time at all.

The intervals are wide. They are not wide enough to make 16% and 67% the same number.

The gains landed where the robot was failing.

The same policies, split by where the box started: the 4 central placements against the remaining 16.

Success rate split by workspace region: the 4 central placements versus the remaining 16.
Where the gains land. All 3 policies do well in the centre (15/20 for teleop-only, 19/20 and 20/20 for the rest). Almost the whole difference is the outer 16 placements: 1/80 for teleop-only, vs 48/80 with handheld added and 52/80 for doubled teleop.

The teleop-only policy was not uniformly weak — in the middle it was fine. It collapsed at the edges: 1 success in 80 attempts outside the centre. That is a policy that has seen the task but not the workspace.

The handheld demos did not make an already-solved placement more reliable. They extended where the policy worked at all — from 1 in 80 to about 6 in 10. Coverage is the expensive thing to teleoperate and the cheap thing to film.

In that outer region, 16 handheld demos (48/80) do not quite match 16 more teleop demos (52/80). They buy the same kind of improvement — coverage — not a strictly better trade.

Retargeting was never the thing that failed.

The obvious worry: that converting human video to robot joints is lossy, and quietly poisons a fraction of the training data.

It did not. Across every retargeted execution recorded for this work, no episode was discarded because the retarget failed. Episodes were discarded, but for ordinary bench reasons — a servo fault mid-run, a placement miss, a recording not started.

That distinction decides whether this scales. A retargeting failure is a research problem. A power fault is a shopping list, and a lighting failure is a checklist.

Share of recorded episodes kept versus discarded, per batch, with discard causes.
Retargeting is not the bottleneck. Episodes kept per batch, and why the rest were discarded. The figure that matters is the zero: no batch lost an episode to a failed retarget. Batch 1 kept 19, of which 16 were used and 3 held as spares.

One task, one arm, one room.

Every number here is checkable.

The policies, the training data and the raw per-attempt records are public. Success rates are not transcribed by hand — a verifier script recomputes each one from the operator's records, and it is in the repository below.

Policies

Datasets

Code & records

Find how much teleoperation you can skip.

Talk to us about your task →