VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
- [2026-09-14] 🎉 We appreciate human five (AI椰青) covering our work—read the WeChat article!
- [2026-09-11] 🎉 We are thrilled that Heart Of Embodied AI (具身智能之心) and Lumina Embodied AI Community (Lumina 具身智能社区) featured our work—check out the WeChat article and Xiaohongshu post!
- [2026-09-07] 🎬 We thank Heart Of Embodied AI for recently featuring our work on WeChat Channels!
- [2026-09-06] 🚀 We are excited to release our paper together with the full VLA-Precision codebase!
The following projects can be used to collect demonstrations in LeRobot format:
| Robot | Teleoperation | Project | Branch |
|---|---|---|---|
| UR5e/UR7e | Keyboard | scy-v/lerobot_ur5e_keyteleop | main |
| UR5e/UR7e | Isomorphic master–slave | scy-v/lerobot_ur5e_isoteleop | main |
| Dual UR5e/UR7e | VR | scy-v/lerobot_ur_dual_vrteleop | main |
| Franka | 3D mouse/VR | Shenzhaolong1330/lerobot_franka_teleop | main |
| Franka | Keyboard | Shenzhaolong1330/lerobot_franka_teleop | vla-precision |
2.1 Install uv and Clone the Repository
curl -LsSf https://astral.sh/uv/install.sh | sh
git clone https://github.com/scy-v/vla-precision.git # Deploy on both server and local robot
cd vla-precisionln -s /path/to/storage/checkpoints ./checkpoints
ln -s /path/to/storage/train_data ./train_data- Policy training runs on the GPU server, with real-robot control provided by the local robot during Stage II.
| Location | Purpose | Command |
|---|---|---|
| GPU server | Stage I | uv sync --frozen --group stage1 |
| GPU server | Stage II | uv sync --frozen --group stage2 |
| Local real-robot | Stage II | uv sync --frozen --group real-robot |
Configuration layout:
configs/stage1/ # Stage I
configs/stage2/tasks/ # Stage II tasks and algorithm
configs/stage2/deployments/ # Network, GPUs, and hardware
You can find descriptions of all configuration parameters in CONFIGURATION.md.
Compute normalization statistics:
# GPU server
uv run --no-sync main.py \
--stage stage1 \
--mode norm-stats \
--config configs/stage1/insert_two_bottles_diagonal_rack.yamlStart OpenPI full-parameter fine-tuning:
# GPU server
uv run --no-sync main.py \
--stage stage1 \
--mode train \
--config configs/stage1/insert_two_bottles_diagonal_rack.yamlPreprocess part of the offline data to populate the Replay Buffer and Context Buffer:
# GPU server
uv run --no-sync main.py \
--stage stage2 \
--mode preprocess \
--config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
--deployment configs/stage2/deployments/single_ur.yamlAfter preprocessing, start the four processes below to begin online RL training:
Robot machine: robot control service
Robot machine: robot-agent communication bridge
GPU server: Learner
GPU server: Actor
Robot machine:
# Terminal 1: low-level robot control
uv run --no-sync main.py --stage stage2 --mode serve-robot \
--config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
--deployment configs/stage2/deployments/single_ur.yaml
# Terminal 2: connect the local robot system to the remote policy process
uv run --no-sync main.py --stage stage2 --mode robot-agent-bridge \
--config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
--deployment configs/stage2/deployments/single_ur.yamlGPU server:
# Terminal 1: Learner
uv run --no-sync main.py --stage stage2 --mode train --role learner \
--config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
--deployment configs/stage2/deployments/single_ur.yaml
# Terminal 2: Actor
uv run --no-sync main.py --stage stage2 --mode train --role actor \
--config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
--deployment configs/stage2/deployments/single_ur.yamlAll four processes use the same task and deployment YAMLs.
The evaluation model can run on either the GPU server or the robot machine. When it runs locally, install the corresponding model dependencies:
| Local evaluation model | Installation |
|---|---|
| Stage I full model | uv sync --frozen --group stage1 --group real-robot |
| Stage II ACoB model | uv sync --frozen --group stage2 --group real-robot |
# Robot machine · Terminal 1
uv run --no-sync main.py --stage stage2 --mode serve-robot \
--config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
--deployment configs/stage2/deployments/single_ur.yaml
# Robot machine · Terminal 2: connect the robot system to the evaluation policy
uv run --no-sync main.py --stage stage2 --mode robot-agent-bridge \
--config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
--deployment configs/stage2/deployments/single_ur.yamlRun evaluation on either the GPU server or robot machine.
# Stage I full model
uv run --no-sync main.py --stage stage1 --mode evaluate \
--config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
--deployment configs/stage2/deployments/single_ur.yaml
# Stage II ACoB model
uv run --no-sync main.py --stage stage2 --mode evaluate \
--config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
--deployment configs/stage2/deployments/single_ur.yamlThis path evaluates Stage I full models. Its configuration is:
src/vla_precision/integrations/openpi/inference/configs/ur.yaml
For a model on the robot machine, set policy.location: local:
# Robot machine
uv run --no-sync main.py --stage stage1 --mode openpi-inference \
--config src/vla_precision/integrations/openpi/inference/configs/ur.yamlFor a model on the GPU server, set policy.location: server and policy.host:
# GPU server
uv run --no-sync main.py --stage stage1 --mode serve-openpi-policy \
--config src/vla_precision/integrations/openpi/inference/configs/ur.yaml
# Robot machine
uv run --no-sync main.py --stage stage1 --mode openpi-inference \
--config src/vla_precision/integrations/openpi/inference/configs/ur.yamlevaluation.checkpoint_step: 0 selects the latest checkpoint. Results are updated after every episode:
results/<experiment>/vla/<time>.json
results/<experiment>/openpi-native/<time>.json
results/<experiment>/acob/<time>.json
VLA-Precision can be adapted to new robots, cameras, grippers, teleoperation devices, and tasks by following the extension guide.
@article{su2026vlaprecision,
title = {VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models},
author = {Su, Chenyu and Shen, Zhaolong and Qian, Yuan and Qian, Chen and Zhang, Rui and Yan, Feng and Chen, Weixing and Zhang, Fei and Wang, Jiamin and Cong, Shuang and Shang, Weiwei},
journal = {arXiv preprint arXiv:2609.04355},
year = {2026},
url = {https://arxiv.org/abs/2609.04355}
}