Human data for physical intelligence
Human data that fits the way robots learn.
RobotFuel helps robot-learning teams collect, structure, and evaluate human experience
for the stage of training that needs it — from diverse pretraining data to
robot-aligned downstream supervision.
Early design partnerships with robotics teams and research groups.
01 / Point of view
Scale matters. So does what you collect.
Large-scale human video can provide environments, objects, and behaviors that are
expensive to reproduce with robots.
But usefulness depends on the training objective: broad pretraining can tolerate
weaker supervision, while downstream action learning may require precise motion,
interaction structure, and robot-specific alignment.
We work backwards from the learning objective to determine what experience,
representation, and fidelity are worth adding.
Human data shouldn't replace robot data. It should make expensive robot data go further.
Robot-native teleoperation and post-training are powerful because they match the
embodiment directly, but they're expensive to scale. Human data is one part of a
broader mix that can include robot logs, simulation, synthetic experience, and
existing video. We're testing where purpose-collected or selected human experience
adds coverage and physical structure before or alongside robot-native data.
02 / Integration surfaces
A qualification layer, not another data silo.
Use one piece or the full path. Each stage emits inspectable artifacts that can
sit around an existing SLAM, training, control, or evaluation stack.
01
Working
Capture preflight
Reject derivatives, missing telemetry, wrong capture profiles, and incomplete
session roles before expensive processing begins.
original media + manifest → input report
Fits before SLAM, labeling, or dataset generation.
02
Working
Embodiment audit
Match a trajectory against a robot model and task constraints, explore
placement and scale, and preserve every rejected frame.
trajectory + robot + task → QA + Rerun
Fits before playback, policy training, or hardware trials.
03
Next integration
Outcome scorecard
Compare the existing robot-data loop against a human-augmented condition on the
team's own held-out evaluator.
trial logs → success + effort + recovery delta
Fits around the evaluator a robotics team already trusts.
Keep your robot stack
→
Add evidence at the boundary where uncertainty enters
→
Return portable JSON, NPZ, logs, and visual review
03 / Pilot loop
From robot gap to evaluation.
01
Define the objective
Start with a robot task, training stage, or data bottleneck and agree on the
downstream metric.
02
Choose the source & spec
Determine whether to collect or select the needed experience, then define its
behaviors, representation, fidelity, and quality.
03
Curate, structure & align
Purpose-collect or identify usable episodes, then add only the physical
structure and robot alignment the experiment requires.
04
Train & evaluate
Put the data into the team's existing loop and use performance and failure cases
to determine what supervision is worth adding next.
The useful experience may already exist, or it may need to be collected for the task.
The downstream result determines what data comes next.
Learning objective
→
Data source / spec
→
Collect or select
→
Structure & align
→
Train / evaluate
→
Next data
↺
04 / Results
Handheld demonstrations did the work of teleoperation.
We trained three policies on the same box pick-and-place task and ran every one of
them live on the arm. Holding the robot data fixed and adding demonstrations filmed
by hand moved the policy from barely working to working — close to what
doubling teleoperation buys, for a fraction of the effort to collect.
The full write-up has the charts, the intervals, the pipeline yield, and what the
result does not show.
05 / Work with us
Have a robot task where data collection is the bottleneck?
Give us one task, one target robot, the current data-collection process, and the
evaluator you already use. We'll identify what can plug in without asking you to
replace the rest of your stack.
“Which of our human demonstrations actually transfer to this robot?”
“Can we reduce teleoperation and HITL while holding task success?”
“Can we reject unusable capture before processing the full dataset?”
06 / FAQ
Frequently asked questions.
Does human data replace robot data?
Usually not.
Robot-native data is already aligned to the robot's sensors, embodiment, and
action space. Human data is useful for a different reason: it can provide much
broader coverage of environments, objects, behaviors, and interaction strategies.
We're interested in where human data can complement robot-native demonstrations
and reduce how much expensive robot collection is required.
What kind of human data do you collect?
It depends on the learning objective.
Broad pretraining may only need scalable egocentric RGB and diverse behavior.
Manipulation experiments may need wrist or object motion. Humanoid learning may
require whole-body information. Robot-specific adaptation may require tighter
embodiment and environment alignment.
We start from the downstream task rather than a fixed sensor configuration.
Is RGB video enough?
Sometimes.
Recent robot-learning systems have shown useful results from monocular or
body-worn RGB video, particularly for large-scale pretraining.
Other applications require more precise physical information, such as
hand-object state, object motion, whole-body pose, or robot-specific alignment.
The useful question is not whether RGB is universally enough, but whether it
provides the information needed for a particular training objective.
Can you convert arbitrary human video directly into robot actions?
Not reliably in the general case.
Human and robot bodies have different kinematics, viewpoints, workspaces, and
grasp capabilities. Recent research addresses this through approaches such as
action retargeting, reconstruction, human-robot alignment, and robot-specific
adaptation.
We treat human-to-robot alignment as part of the problem rather than assuming
it is already solved.
Why not just use existing datasets or internet video?
We should, when they contain the right experience.
New collection is not automatically better. Existing human video, robot
datasets, simulation, or previously collected data may already contain what a
model needs.
Part of the problem is determining whether the relevant experience already
exists, whether it can be selected and structured effectively, or whether
targeted collection is actually necessary.
What is a targeted human-data pilot?
A small experiment built around a concrete robot-learning question.
We define the target task or capability, the behavior or distribution that may
be missing, the required signals and quality, the target robot or training stage
where relevant, and how the data will be evaluated downstream.
The goal is to learn whether the data changes something meaningful before
scaling collection.
What does success look like?
It depends on the team's existing evaluation loop.
Possible outcomes include improved task success or robustness, better
performance on unseen objects or environments, higher usable-data yield, reduced
robot teleoperation or correction time, or improved reconstruction and
retargeting quality.
A useful result can also be finding that the proposed human data does not help.
We prefer downstream evaluation over judging a dataset only by how polished the
labels look.