Human data for physical intelligence

Human data that fits the way robots learn.

RobotFuel helps robot-learning teams collect, structure, and evaluate human experience for the stage of training that needs it — from diverse pretraining data to robot-aligned downstream supervision.

Early design partnerships with robotics teams and research groups.

01 / Point of view

Scale matters. So does what you collect.

Large-scale human video can provide environments, objects, and behaviors that are expensive to reproduce with robots.

But usefulness depends on the training objective: broad pretraining can tolerate weaker supervision, while downstream action learning may require precise motion, interaction structure, and robot-specific alignment.

We work backwards from the learning objective to determine what experience, representation, and fidelity are worth adding.

Human data shouldn't replace robot data. It should make expensive robot data go further.

Robot-native teleoperation and post-training are powerful because they match the embodiment directly, but they're expensive to scale. Human data is one part of a broader mix that can include robot logs, simulation, synthetic experience, and existing video. We're testing where purpose-collected or selected human experience adds coverage and physical structure before or alongside robot-native data.

02 / Integration surfaces

A qualification layer, not another data silo.

Use one piece or the full path. Each stage emits inspectable artifacts that can sit around an existing SLAM, training, control, or evaluation stack.

01
Working

Capture preflight

Reject derivatives, missing telemetry, wrong capture profiles, and incomplete session roles before expensive processing begins.

original media + manifest → input report Fits before SLAM, labeling, or dataset generation.
02
Working

Embodiment audit

Match a trajectory against a robot model and task constraints, explore placement and scale, and preserve every rejected frame.

trajectory + robot + task → QA + Rerun Fits before playback, policy training, or hardware trials.
03
Next integration

Outcome scorecard

Compare the existing robot-data loop against a human-augmented condition on the team's own held-out evaluator.

trial logs → success + effort + recovery delta Fits around the evaluator a robotics team already trusts.
Keep your robot stack Add evidence at the boundary where uncertainty enters Return portable JSON, NPZ, logs, and visual review
03 / Pilot loop

From robot gap to evaluation.

01

Define the objective

Start with a robot task, training stage, or data bottleneck and agree on the downstream metric.

02

Choose the source & spec

Determine whether to collect or select the needed experience, then define its behaviors, representation, fidelity, and quality.

03

Curate, structure & align

Purpose-collect or identify usable episodes, then add only the physical structure and robot alignment the experiment requires.

04

Train & evaluate

Put the data into the team's existing loop and use performance and failure cases to determine what supervision is worth adding next.

The useful experience may already exist, or it may need to be collected for the task. The downstream result determines what data comes next.

Learning objective Data source / spec Collect or select Structure & align Train / evaluate Next data
04 / Results

Handheld demonstrations did the work of teleoperation.

We trained three policies on the same box pick-and-place task and ran every one of them live on the arm. Holding the robot data fixed and adding demonstrations filmed by hand moved the policy from barely working to working — close to what doubling teleoperation buys, for a fraction of the effort to collect.

The full write-up has the charts, the intervals, the pipeline yield, and what the result does not show.

05 / Work with us

Have a robot task where data collection is the bottleneck?

Give us one task, one target robot, the current data-collection process, and the evaluator you already use. We'll identify what can plug in without asking you to replace the rest of your stack.

“Which of our human demonstrations actually transfer to this robot?”

“Can we reduce teleoperation and HITL while holding task success?”

“Can we reject unusable capture before processing the full dataset?”

Submitting opens your email app with the request ready to send.

06 / FAQ

Frequently asked questions.

Does human data replace robot data?

Usually not.

Robot-native data is already aligned to the robot's sensors, embodiment, and action space. Human data is useful for a different reason: it can provide much broader coverage of environments, objects, behaviors, and interaction strategies.

We're interested in where human data can complement robot-native demonstrations and reduce how much expensive robot collection is required.

What kind of human data do you collect?

It depends on the learning objective.

Broad pretraining may only need scalable egocentric RGB and diverse behavior. Manipulation experiments may need wrist or object motion. Humanoid learning may require whole-body information. Robot-specific adaptation may require tighter embodiment and environment alignment.

We start from the downstream task rather than a fixed sensor configuration.

Is RGB video enough?

Sometimes.

Recent robot-learning systems have shown useful results from monocular or body-worn RGB video, particularly for large-scale pretraining.

Other applications require more precise physical information, such as hand-object state, object motion, whole-body pose, or robot-specific alignment.

The useful question is not whether RGB is universally enough, but whether it provides the information needed for a particular training objective.

Can you convert arbitrary human video directly into robot actions?

Not reliably in the general case.

Human and robot bodies have different kinematics, viewpoints, workspaces, and grasp capabilities. Recent research addresses this through approaches such as action retargeting, reconstruction, human-robot alignment, and robot-specific adaptation.

We treat human-to-robot alignment as part of the problem rather than assuming it is already solved.

Why not just use existing datasets or internet video?

We should, when they contain the right experience.

New collection is not automatically better. Existing human video, robot datasets, simulation, or previously collected data may already contain what a model needs.

Part of the problem is determining whether the relevant experience already exists, whether it can be selected and structured effectively, or whether targeted collection is actually necessary.

What is a targeted human-data pilot?

A small experiment built around a concrete robot-learning question.

We define the target task or capability, the behavior or distribution that may be missing, the required signals and quality, the target robot or training stage where relevant, and how the data will be evaluated downstream.

The goal is to learn whether the data changes something meaningful before scaling collection.

What does success look like?

It depends on the team's existing evaluation loop.

Possible outcomes include improved task success or robustness, better performance on unseen objects or environments, higher usable-data yield, reduced robot teleoperation or correction time, or improved reconstruction and retargeting quality.

A useful result can also be finding that the proposed human data does not help. We prefer downstream evaluation over judging a dataset only by how polished the labels look.