ACCEPTED AT CoRL 2026

BiCICLe

Bimanual Robot Manipulation via
Multi-Agent In-Context Learning

Alessio Palma1,*, Indro Spinelli1, Vignesh Prasad2,3,4, Luca Scofano1, Yufeng Jin2, Georgia Chalvatzaki2,3,4,†, Fabio Galasso1,†
1 PINlab, Sapienza University of Rome, Italy 2 Interactive Robot Perception & Learning (PEARL) Lab, TU Darmstadt, Germany 3 Hessian.AI, Germany 4 Robotics Institute Germany

* Corresponding author · † Co-senior authors

Two arms. One coordinated plan. Successful executions on a physical Franka Panda system.

BiCICLe coordinates two robot arms through leader–follower in-context learning, using demonstrations instead of task-specific fine-tuning.

10demonstrations per prompt
No gradient updatesfrozen, text-only language models
Simulation + real robotthe same pose-based ICL interface

Paper

Abstract

Large Language Models (LLMs) have emerged as powerful reasoning engines for embodied control. In particular, In-Context Learning (ICL) enables off-the-shelf, text-only LLMs to predict robot actions without any task-specific training while preserving their generalization capabilities. Applying ICL to bimanual manipulation remains challenging as the high-dimensional joint action space and tight inter-arm coordination constraints rapidly overwhelm standard context windows. To address this, we introduce BiCICLe (Bimanual Coordinated In-Context Learning), the first framework that enables standard LLMs to perform few-shot bimanual manipulation without fine-tuning. BiCICLe frames bimanual control as a multi-agent leader-follower problem, decoupling the action space into sequential, conditioned single-arm predictions. Evaluated on 13 tasks from the TWIN benchmark, BiCICLe achieves 70.5% average success rate, outperforming the best training-free baseline by 6.1 percentage points and surpassing most supervised methods. We also demonstrate superior real-world performance on 3 tasks without hardware-specific retraining.

Video presentation

Watch the method in action.

A narrated overview of the framework, simulation tasks, and real robot experiments.

Method

A leader plans. A follower coordinates.

Predicting both arms jointly is complex. Predicting them independently loses coordination. BiCICLe keeps each prediction single-arm, while sharing the leader’s complete trajectory.

Ten demonstrations pair object positions with single-arm actions. The leader predicts its trajectory from the new scene. Its complete plan is added to the follower's observation. The follower predicts conditioned actions, and the two plans are composed into bimanual keyposes.
  1. 01

    Read the examples

    Serialize object positions and keypose demonstrations as text. The task is inferred from the examples, without a written task description.

  2. 02

    Predict the leader

    A frozen LLM predicts the leader arm’s full trajectory: position, rotation, and gripper state at each keypose.

  3. 03

    Condition the follower

    Add the leader’s plan to the follower’s observation. Predict the second trajectory, then compose both for motion-planned execution.

Bimanual demonstrations

Different tasks. Shared coordination.

Representative successful episodes from TWIN.

Handover

Transfer an object between arms.

Lift Tray

Lift together from opposite sides.

Pick Plate

Coordinate contact around a plate.

Sweep Dustpan

Assign complementary roles to two tools.

Push Box

Move a shared object with both arms.

Item Drawer

Open a drawer and place an item inside.

More TWIN demonstrations

Dual Buttons

Press with both arms.

Lift Ball

Maintain contact with a shared object.

Pick Laptop

Grasp a thin object with two arms.

Straighten Rope

Manipulate a deformable object.

Simulation results

70.5% average success on TWIN.

With ten demonstrations per prompt, BiCICLe improves on the strongest training-free baseline by 6.1 percentage points across 13 tasks.

Training-free ICL methods

Average success (%)
GPT-5-mini; 3 seeds × 100 episodes per task. SA predicts joint actions; DA predicts each arm independently. Simulation object centroids use RGB-D point clouds and ground-truth segmentation masks.

New tasks, adapted through the prompt.

On two tasks outside TWIN, BiCICLe reaches 54.5% average success. For this experiment, 3DFA is fine-tuned on ten demonstrations per task; BiCICLe uses ten in-context examples.

Close Jar

Lid handover and jar stabilization.

Take Item Out of Box

Hold the lid and retrieve the object.

Success (%) · 100 episodes per task
MethodJarBoxAvg.
3DFA full fine-tuning11.09.010.0
3DFA LoRA12.018.015.0
BiCICLe61.048.054.5

Real-world experiments

From text predictions to physical manipulation.

Two Franka Panda Research 3 arms, a ZED X camera, FoundationPose object estimates, and MoveIt execution. The pose-based interface transfers without hardware-specific retraining.

Lift Box

Symmetric lifting from opposite sides.

Open Pot

Stabilize the pot, then lift the lid.

Cleanup

Place two objects into a shared basket.

53.3%

average success across three tasks

15 kinesthetic demonstrations collected per task; 10 sampled in each prompt. Evaluation: 10 trials per task with varied object locations.

Real-world success (%)
MethodLift BoxOpen PotCleanupAvg.
KAT-DA40.020.010.023.3
RoboPrompt-DA30.040.030.033.3
BiCICLe60.040.060.053.3

Citation

BibTeX

@misc{palma2026bimanual,
  title         = {Bimanual Robot Manipulation via Multi-Agent In-Context Learning},
  author        = {Alessio Palma and Indro Spinelli and Vignesh Prasad and Luca Scofano and Yufeng Jin and Georgia Chalvatzaki and Fabio Galasso},
  year          = {2026},
  eprint        = {2604.20348},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2604.20348}
}

Acknowledgments

This work was supported by the Sapienza grant RG123188B3EF6A80 (CENTS), the German Research Foundation Emmy Noether Programme (CH 2676/1-1), the EU’s Horizon Europe project “ARISE” (Grant 101135959), the German Federal Ministry of Research, Technology and Space of Germany project “RIG” (Grant 16ME1001) and the European Research Council project “SIREN” (Grant 101163933). We also thank CINECA for the allocation of computational resources.