Skip to content

Repository files navigation

VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models

Project Website Paper Apache-2.0 License

VLA-Precision overview

Teleoperation to supervised fine-tuning to online RL post-training to real-robot deployment

🔥 News

  • [2026-09-14] 🎉 We appreciate human five (AI椰青) covering our work—read the WeChat article!
  • [2026-09-11] 🎉 We are thrilled that Heart Of Embodied AI (具身智能之心) and Lumina Embodied AI Community (Lumina 具身智能社区) featured our work—check out the WeChat article and Xiaohongshu post!
  • [2026-09-07] 🎬 We thank Heart Of Embodied AI for recently featuring our work on WeChat Channels!
  • [2026-09-06] 🚀 We are excited to release our paper together with the full VLA-Precision codebase!

🎮 1. Teleoperation and Data Collection

The following projects can be used to collect demonstrations in LeRobot format:

Robot Teleoperation Project Branch
UR5e/UR7e Keyboard scy-v/lerobot_ur5e_keyteleop main
UR5e/UR7e Isomorphic master–slave scy-v/lerobot_ur5e_isoteleop main
Dual UR5e/UR7e VR scy-v/lerobot_ur_dual_vrteleop main
Franka 3D mouse/VR Shenzhaolong1330/lerobot_franka_teleop main
Franka Keyboard Shenzhaolong1330/lerobot_franka_teleop vla-precision

📋 2. Environment Setup

2.1 Install uv and Clone the Repository

curl -LsSf https://astral.sh/uv/install.sh | sh
git clone https://github.com/scy-v/vla-precision.git  # Deploy on both server and local robot
cd vla-precision

2.2 Optional: Configure Storage Directories on the Server

ln -s /path/to/storage/checkpoints ./checkpoints
ln -s /path/to/storage/train_data ./train_data

2.3 Install Runtime Dependencies

  • Policy training runs on the GPU server, with real-robot control provided by the local robot during Stage II.
Location Purpose Command
GPU server Stage I uv sync --frozen --group stage1
GPU server Stage II uv sync --frozen --group stage2
Local real-robot Stage II uv sync --frozen --group real-robot

🚀 3. Training Pipeline

Configuration layout:

configs/stage1/               # Stage I
configs/stage2/tasks/         # Stage II tasks and algorithm
configs/stage2/deployments/   # Network, GPUs, and hardware

You can find descriptions of all configuration parameters in CONFIGURATION.md.

3.1 Stage I: OpenPI full-parameter fine-tuning

Compute normalization statistics:

# GPU server
uv run --no-sync main.py \
  --stage stage1 \
  --mode norm-stats \
  --config configs/stage1/insert_two_bottles_diagonal_rack.yaml

Start OpenPI full-parameter fine-tuning:

# GPU server
uv run --no-sync main.py \
  --stage stage1 \
  --mode train \
  --config configs/stage1/insert_two_bottles_diagonal_rack.yaml

3.2 Stage II: ACoB online post-training

Preprocess part of the offline data to populate the Replay Buffer and Context Buffer:

# GPU server
uv run --no-sync main.py \
  --stage stage2 \
  --mode preprocess \
  --config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
  --deployment configs/stage2/deployments/single_ur.yaml

After preprocessing, start the four processes below to begin online RL training:

Robot machine: robot control service
Robot machine: robot-agent communication bridge
GPU server: Learner
GPU server: Actor

Robot machine:

# Terminal 1: low-level robot control
uv run --no-sync main.py --stage stage2 --mode serve-robot \
  --config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
  --deployment configs/stage2/deployments/single_ur.yaml

# Terminal 2: connect the local robot system to the remote policy process
uv run --no-sync main.py --stage stage2 --mode robot-agent-bridge \
  --config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
  --deployment configs/stage2/deployments/single_ur.yaml

GPU server:

# Terminal 1: Learner
uv run --no-sync main.py --stage stage2 --mode train --role learner \
  --config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
  --deployment configs/stage2/deployments/single_ur.yaml

# Terminal 2: Actor
uv run --no-sync main.py --stage stage2 --mode train --role actor \
  --config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
  --deployment configs/stage2/deployments/single_ur.yaml

All four processes use the same task and deployment YAMLs.

📊 4. Evaluation

The evaluation model can run on either the GPU server or the robot machine. When it runs locally, install the corresponding model dependencies:

Local evaluation model Installation
Stage I full model uv sync --frozen --group stage1 --group real-robot
Stage II ACoB model uv sync --frozen --group stage2 --group real-robot

4.1 VLA-Precision evaluation

# Robot machine · Terminal 1
uv run --no-sync main.py --stage stage2 --mode serve-robot \
  --config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
  --deployment configs/stage2/deployments/single_ur.yaml

# Robot machine · Terminal 2: connect the robot system to the evaluation policy
uv run --no-sync main.py --stage stage2 --mode robot-agent-bridge \
  --config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
  --deployment configs/stage2/deployments/single_ur.yaml

Run evaluation on either the GPU server or robot machine.

# Stage I full model
uv run --no-sync main.py --stage stage1 --mode evaluate \
  --config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
  --deployment configs/stage2/deployments/single_ur.yaml

# Stage II ACoB model
uv run --no-sync main.py --stage stage2 --mode evaluate \
  --config configs/stage2/tasks/insert_two_bottles_diagonal_rack.yaml \
  --deployment configs/stage2/deployments/single_ur.yaml

4.2 Native OpenPI evaluation

This path evaluates Stage I full models. Its configuration is:

src/vla_precision/integrations/openpi/inference/configs/ur.yaml

For a model on the robot machine, set policy.location: local:

# Robot machine
uv run --no-sync main.py --stage stage1 --mode openpi-inference \
  --config src/vla_precision/integrations/openpi/inference/configs/ur.yaml

For a model on the GPU server, set policy.location: server and policy.host:

# GPU server
uv run --no-sync main.py --stage stage1 --mode serve-openpi-policy \
  --config src/vla_precision/integrations/openpi/inference/configs/ur.yaml

# Robot machine
uv run --no-sync main.py --stage stage1 --mode openpi-inference \
  --config src/vla_precision/integrations/openpi/inference/configs/ur.yaml

4.3 Results

evaluation.checkpoint_step: 0 selects the latest checkpoint. Results are updated after every episode:

results/<experiment>/vla/<time>.json
results/<experiment>/openpi-native/<time>.json
results/<experiment>/acob/<time>.json

🧩 5. Extensions and Customization

VLA-Precision can be adapted to new robots, cameras, grippers, teleoperation devices, and tasks by following the extension guide.

📝 6. Citation

@article{su2026vlaprecision,
  title   = {VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models},
  author  = {Su, Chenyu and Shen, Zhaolong and Qian, Yuan and Qian, Chen and Zhang, Rui and Yan, Feng and Chen, Weixing and Zhang, Fei and Wang, Jiamin and Cong, Shuang and Shang, Weiwei},
  journal = {arXiv preprint arXiv:2609.04355},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.04355}
}

About

VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models

Resources

Stars

81 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages