ACCEPTED AT CoRL 2026
BiCICLe
Bimanual Robot Manipulation via
Multi-Agent In-Context Learning
BiCICLe coordinates two robot arms through leader–follower in-context learning, using demonstrations instead of task-specific fine-tuning.
Paper
Abstract
Large Language Models (LLMs) have emerged as powerful reasoning engines for embodied control. In particular, In-Context Learning (ICL) enables off-the-shelf, text-only LLMs to predict robot actions without any task-specific training while preserving their generalization capabilities. Applying ICL to bimanual manipulation remains challenging as the high-dimensional joint action space and tight inter-arm coordination constraints rapidly overwhelm standard context windows. To address this, we introduce BiCICLe (Bimanual Coordinated In-Context Learning), the first framework that enables standard LLMs to perform few-shot bimanual manipulation without fine-tuning. BiCICLe frames bimanual control as a multi-agent leader-follower problem, decoupling the action space into sequential, conditioned single-arm predictions. Evaluated on 13 tasks from the TWIN benchmark, BiCICLe achieves 70.5% average success rate, outperforming the best training-free baseline by 6.1 percentage points and surpassing most supervised methods. We also demonstrate superior real-world performance on 3 tasks without hardware-specific retraining.
Video presentation
Watch the method in action.
A narrated overview of the framework, simulation tasks, and real robot experiments.
Method
A leader plans. A follower coordinates.
Predicting both arms jointly is complex. Predicting them independently loses coordination. BiCICLe keeps each prediction single-arm, while sharing the leader’s complete trajectory.
- 01
Read the examples
Serialize object positions and keypose demonstrations as text. The task is inferred from the examples, without a written task description.
- 02
Predict the leader
A frozen LLM predicts the leader arm’s full trajectory: position, rotation, and gripper state at each keypose.
- 03
Condition the follower
Add the leader’s plan to the follower’s observation. Predict the second trajectory, then compose both for motion-planned execution.
Bimanual demonstrations
Different tasks. Shared coordination.
Representative successful episodes from TWIN.
Handover
Transfer an object between arms.
Lift Tray
Lift together from opposite sides.
Pick Plate
Coordinate contact around a plate.
Sweep Dustpan
Assign complementary roles to two tools.
Push Box
Move a shared object with both arms.
Item Drawer
Open a drawer and place an item inside.
More TWIN demonstrations
Dual Buttons
Press with both arms.
Lift Ball
Maintain contact with a shared object.
Pick Laptop
Grasp a thin object with two arms.
Straighten Rope
Manipulate a deformable object.
Simulation results
70.5% average success on TWIN.
With ten demonstrations per prompt, BiCICLe improves on the strongest training-free baseline by 6.1 percentage points across 13 tasks.
Training-free ICL methods
Average success (%)New tasks, adapted through the prompt.
On two tasks outside TWIN, BiCICLe reaches 54.5% average success. For this experiment, 3DFA is fine-tuned on ten demonstrations per task; BiCICLe uses ten in-context examples.
Close Jar
Lid handover and jar stabilization.
Take Item Out of Box
Hold the lid and retrieve the object.
| Method | Jar | Box | Avg. |
|---|---|---|---|
| 3DFA full fine-tuning | 11.0 | 9.0 | 10.0 |
| 3DFA LoRA | 12.0 | 18.0 | 15.0 |
| BiCICLe | 61.0 | 48.0 | 54.5 |
Real-world experiments
From text predictions to physical manipulation.
Two Franka Panda Research 3 arms, a ZED X camera, FoundationPose object estimates, and MoveIt execution. The pose-based interface transfers without hardware-specific retraining.
Lift Box
Symmetric lifting from opposite sides.
Open Pot
Stabilize the pot, then lift the lid.
Cleanup
Place two objects into a shared basket.
53.3%
average success across three tasks
15 kinesthetic demonstrations collected per task; 10 sampled in each prompt. Evaluation: 10 trials per task with varied object locations.
| Method | Lift Box | Open Pot | Cleanup | Avg. |
|---|---|---|---|---|
| KAT-DA | 40.0 | 20.0 | 10.0 | 23.3 |
| RoboPrompt-DA | 30.0 | 40.0 | 30.0 | 33.3 |
| BiCICLe | 60.0 | 40.0 | 60.0 | 53.3 |
Citation
BibTeX
@misc{palma2026bimanual,
title = {Bimanual Robot Manipulation via Multi-Agent In-Context Learning},
author = {Alessio Palma and Indro Spinelli and Vignesh Prasad and Luca Scofano and Yufeng Jin and Georgia Chalvatzaki and Fabio Galasso},
year = {2026},
eprint = {2604.20348},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2604.20348}
}
Acknowledgments
This work was supported by the Sapienza grant RG123188B3EF6A80 (CENTS), the German Research Foundation Emmy Noether Programme (CH 2676/1-1), the EU’s Horizon Europe project “ARISE” (Grant 101135959), the German Federal Ministry of Research, Technology and Space of Germany project “RIG” (Grant 16ME1001) and the European Research Council project “SIREN” (Grant 101163933). We also thank CINECA for the allocation of computational resources.