HuRo converts egocentric human video into robot-aligned observations and action trajectories. The pipeline retargets the human hand motion into robot joint trajectories, and it removes the human arms from the frames and renders the robot in their place. Intermediate signals a source does not provide are estimated, so videos at different annotation levels are converted into a common format. Using this pipeline we construct the HuRo dataset, comprising about 630K robotized episodes and 142M processed frames from five human-video sources. Across four real-world manipulation tasks, pretraining a vision-language-action (VLA) policy on increasing amounts of robotized human-video data improves overall completion after finetuning from 51.5% to 80.3% and out-of-distribution completion under spatial and visual shifts from 34.9% to 72.2%.
- [2026-09-11] π Code released.
- [2026-09-10] π Paper on arXiv.
- [2026-09-09] π Project page live.
- π’ News & Updates
- π Robotization Pipeline
- π§ Setup
- β‘ Run
- π¬ Input video
- π€ Target robot
- π Data format
- π₯ HuRo Dataset
- π License
- π Citation
Note
This repository releases HuRo's robotization pipeline, from raw video to a LeRobot V2.0 dataset.
HuRo's robotization pipeline converts raw, untrimmed egocentric video of everyday activity
into robotized episodes. It first estimates the camera and hand annotations,
identifies the manipulation segments within the untrimmed clip, and assigns a language instruction
to each. Then action conversion retargets the human hand motion into the target robot's joint
trajectory, and visual conversion removes the visible human embodiment and renders that robot
onto the cleaned scene. A segment becomes one episode of the target embodiment, comprising the
robotized video, the retargeted states and actions, and a language instruction, formatted as a
LeRobot V2.0 dataset. See pipeline/README.md for the pipeline in detail,
and the data format reference in examples/README.md for the outputs a
run writes and how to read them.
The pipeline requires Linux, an NVIDIA GPU with at least 24 GB of VRAM and a driver supporting
CUDA 12.8, and git. The robot overlay also requires a GPU with RT cores and a driver no newer
than R580. setup/README.md covers the installation and how to download
the off-the-shelf models the pipeline uses.
./run_pipeline.shThe script runs all ten stages in order over a directory of clips. Edit the settings block at
the top of it to configure a run. INPUT_DIR defaults to examples/clips/, and those two clips
already meet the requirements in Input video, so the script runs as it stands.
All stages are resumable.
The outputs are written beside the input directory, as <clips>_intr, _contact,
_contact_refined, _hand, _extr, _chunked and _lerobot.
examples/README.md documents what each of them holds.
Each stage can also be run on its own:
CUDA_VISIBLE_DEVICES=<gpu> python pipeline/stage<N>_<phase>_<name>.py \
--input_dir <clip-dir> --part <a>/<b> --no_tqdm
CUDA_VISIBLE_DEVICES=0 python pipeline/stage1_annot_intrinsics.py \
--input_dir examples/clips --part 2/4 --no_tqdmEvery stage accepts those three flags. The retargeting, overlay and LeRobot conversion stages
additionally accept --robot_name, which defaults to allex (configs/allex.yaml). See
Target robot for adding a custom robot.
Note
The released code does not read annotations that a source dataset provides. It estimates
the camera geometry, hand poses and language from the video itself. To start at stage N with
provided annotations, write what stage N-1 would write in the format that
examples/README.md documents, with the .done markers included.
Without the markers, a stage can treat a clip as dropped.
INPUT_DIR (in run_pipeline.sh) points at a directory of .mp4 files. Untrimmed video is
fine, because the pipeline finds its own segments. A file of roughly 30 seconds to 30 minutes
is the length to aim for. A very short clip can hold too little camera motion for a good
calibration. A longer clip takes more CPU memory.
Three properties of the video matter:
- Egocentric. The whole pipeline is designed around egocentric video.
- 30 fps. The pipeline's internal parameters assume that rate.
- A 256-pixel shorter side. A larger input also runs, but decoding the video and resizing its frames for the off-the-shelf models take longer, without a noticeable gain in quality.
Fisheye and rectilinear video are both accepted as they are. A clip whose camera calibration fails
is dropped without an error, and an empty <clips>_intr/<clip>.json is the record of that.
Convert each video to 30 fps and a 256-pixel shorter side with:
ffmpeg -i raw.mp4 -vf "fps=30,scale='if(gt(iw,ih),-2,256)':'if(gt(iw,ih),256,-2)'" \
-c:v libx264 -crf 18 -an clips/raw_30fps_256.mp4Allex is the only robot configured here. Adding another robot is mostly configuration. The
retargeting and the overlay read one YAML file and the URDF it names, and they need no code
change. The annotations before the retargeting do not depend on the robot either, so an existing
<clips>_chunked tree can be converted for another robot. See
configs/README.md for the full procedure.
The Parquet tables carry the annotations, and the LeRobot V2.0 dataset holds the
robotized episodes. See examples/README.md for both formats in full.
Two reader scripts come with them, one per format, and each is the reference implementation for its own format. Run them on the outputs of the example clips:
python examples/read_parquet.py examples/clips_chunked # the Parquet tables
python examples/load_lerobot.py examples/clips_lerobot/allex/192x342 # the LeRobot episodesComing soon.
HuRo's own code is Apache-2.0 (LICENSE).
THIRD_PARTY_NOTICES.md lists the licence of every
off-the-shelf component, what this repository redistributes of each, and the
attribution for the EPIC-KITCHENS material.
Warning
This pipeline cannot be run commercially. Its dependencies impose the restriction.
@misc{jeong2026huro,
title={HuRo: Robotizing Human Videos for Scalable VLA Pretraining},
author={Jinho Jeong and Se June Joo and Jaehyun Kang and Dongyun Kim and Yena Kim and Hanjung Kim and Seon Joo Kim},
year={2026},
eprint={2609.10706},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2609.10706},
}
