ParcelStow evaluates how the performance difference between an imitation learner and its expert varies with task conditions. Three simulated manipulation tasks instantiate this comparison through variation in execution timing.
3 tasks · 970,565 demonstration control steps · scripted experts + ACT, DP, and DAgger checkpoints · canonical evaluation records · one policy interface · CPU-only result reproduction from records
Project Page · Paper · Dataset · Runbook · Install · Reproduce Results on CPU · Evaluate a Policy · Submit Policy Results
Support ParcelStow: ★ Star on GitHub · ♥ Like on Hugging Face
Expert on the top row; ACT below. Each column shows one task at speedup
factor r=2, played at 2× speed. The ACT recordings use the checkpoints
reported below. Labels identify each recorded outcome.
Watch the full-resolution video.
A scalar speedup factor r divides the nominal durations of selected task
phases while acquisition and settling retain fixed durations. At r=1, the
schedule is nominal; at r=2, the scaled phases have half their nominal
duration. Expert and learner policies are evaluated over the same factor grid.
| Task | Expert at r=1 |
ACT at r=1 |
Higher r |
Expert | ACT |
|---|---|---|---|---|---|
| Parcel insertion | 100/100 | 100/100 | 2 | 84/100 | 53/100 |
| Upright placement | 185/200 | 194/200 | 2 | 105/200 | 24/200 |
| Keyed peg insertion | 182/200 | 191/200 | 1.5 | 150/200 | 3/200 |
Each ACT column reports one trained policy per task. Expert-minus-ACT
differences at the highlighted factors are 31, 40.5, and 73.5 percentage points,
with pointwise 95% paired bootstrap intervals [18,44], [32,49], and
[67.5,79.5]. Parcel r=2 is within its demonstrated range; upright r=2
and peg r=1.5 are outside theirs. Training replications are reported
individually in the evaluation records.
Diffusion Policy (DP) nominal success is 68/100, 181/200, and 95/200 on parcel, upright, and peg. Its parcel results support standalone success rates because episode-level pairing and distribution matching with the expert are not established. DAgger nominal success is 3/100, 1/200, and 0/200; these results support failure characterization but not execution-speed degradation claims. The record catalogs provide task-stage counts, terminal failures, hashes, pairing restrictions, and excluded series.
Clone main to use the current three-task benchmark:
git clone --branch main https://github.com/coenwerem/parcelstow.git
cd parcelstowSimulator execution requires Isaac Lab and a supported NVIDIA GPU. From the repository root, install the extension into the Isaac Lab Python environment:
uv pip install -p <isaaclab-venv>/bin/python -e source/parcelstowSimulator records use Python 3.11.14, Isaac Sim 5.1.0, Isaac Lab 0.54.2, PyTorch 2.7.0+cu128, SciPy 1.15.3, and NumPy 1.26.4 on an RTX 5070 Ti. Upright/peg ACT training used Python 3.12.3, PyTorch 2.8.0+cu128, and NumPy 2.3.5. The runbook distinguishes the environments.
Run direct module tests without Isaac Lab:
python -m pytest tests/ -qRun the simulator groups in separate processes:
python -m pytest tests/test_parcel_physics.py tests/test_relative_handoff.py --isaac -q
python -m pytest tests/test_upright_physics.py --isaac-upright -q
python -m pytest tests/test_peg_physics.py --isaac-peg -qChange only --task to run another scripted expert:
python scripts/run_task.py --task parcel
python scripts/run_task.py --task upright
python scripts/run_task.py --task pegEvaluate the released experts through the same public interface:
python scripts/evaluate.py --task parcel --actor expert
python scripts/evaluate.py --task upright --actor expert
python scripts/evaluate.py --task peg --actor expertThe --task value selects one task: parcel selects parcel insertion,
upright selects upright placement, and peg selects keyed peg insertion.
The canonical checkpoint bundle is explicit:
python3 scripts/download_artifacts.py --manuscript
python3 scripts/download_artifacts.py --manuscript --verifyFor an offline copy from a local Hugging Face checkout, add
--local-hf ../parcelstow-hf. Use the
runbook evaluation commands
with the listed checkpoint paths, episode counts, factor grids, and bank seeds.
The short commands above are for trying the interface; the runbook specifies
the complete evaluation configuration.
The canonical inventory and artifact index identify each condition, checkpoint, and dataset. The record guide explains how to resolve original source paths and interpret pairing status. Task specifications: parcel, upright, and peg.
Recompute the canonical counts, stages, failures and pairing checks without Isaac Lab or a GPU:
python3 scripts/reproduce_manuscript.py --output-dir outputs/reproduce/canonicalAdd --bootstrap with NumPy installed for 20,000-resample paired intervals.
The expected audit covers 136 conditions and 20 source series. The output
path must be new. See the detailed runbook for environment setup,
training, checkpoint selection, evaluation and all dataset/record locations.
The same Python class can be loaded for every task:
python scripts/evaluate.py --task parcel --actor examples.custom_policy:HoldPosturePolicy --rates 1 --episodes 5
python scripts/evaluate.py --task upright --actor examples.custom_policy:HoldPosturePolicy --rates 1 --episodes 5
python scripts/evaluate.py --task peg --actor examples.custom_policy:HoldPosturePolicy --rates 1 --episodes 5All tasks produce a 147-dimensional state observation and accept a 16-dimensional normalized joint-position action at 50 Hz. Task identity is selected by --task; it is not appended to the observation. Observation index 146 contains r. The pose slice at indices 118:125 represents parcel_pose for parcel insertion and object_pose for upright placement and keyed peg insertion. Phase values retain task-specific schedules. Policy Interface documents every slice and the adapter boundary.
HoldPosturePolicy commands the default posture and normally fails. It demonstrates loading and record generation, not task performance.
The Hugging Face repository
organizes demonstrations, training tensors, checkpoints, evaluation records,
and videos by artifact role. The manuscript_20260921 directories contain
the checkpoints, records, and illustrations used for the three-task results.
artifacts/manifest.json records download paths, sizes and SHA-256. Data and Checkpoints and the runbook identify the files used by each training and evaluation command.
Policy Results lists the included baselines and defines the evidence required to submit another policy. Contributing distinguishes bug reports, policy results, policy integrations, candidate tasks, and changes to fixed definitions. Candidate Task Authoring Protocol defines the scientific and software evidence required before a task can be listed as part of ParcelStow.
| Path | Content |
|---|---|
scripts/run_task.py, scripts/evaluate.py |
public simulator commands for all three tasks |
scripts/reproduce_manuscript.py |
canonical counts, intervals, and pairing from frozen records |
docs/RUNBOOK.md |
training, evaluation, environments, and artifact locations |
scripts/task_registry.py |
task aliases, gym IDs, defaults, stage keys, experts, monitors, and schedules |
source/parcelstow/ |
Isaac Lab extension and task definitions |
data/manuscript_20260921/ |
canonical records, pairing, exclusions, artifact index and provenance |
examples/custom_policy.py |
one policy class loadable on all tasks |
RESULTS.md |
included baselines and policy-result submission requirements |
docs/ |
current benchmark, policy, reproduction, contribution, and task specifications |
Cite the accompanying preprint:
@misc{enwerem2026parcelstow,
title = {Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds},
author = {Enwerem, Clinton and Baras, John S. and Belta, Calin},
year = {2026},
eprint = {2609.01453},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2609.01453}
}A software citation is available in CITATION.cff.
