IMPACT: Attention Is the Interaction Map for
Scalable Interaction-Aware World Model Training
Abstract
World models have made remarkable progress in action-conditioned future prediction for embodied agents, yet still struggle to model physically plausible interactions. Existing approaches address this limitation by constraining the generation process with external representations encoding motion, geometry, or semantics. Obtaining these spatiotemporally dense representations typically requires auxiliary estimators or manual annotations, limiting training scalability. We instead revisit the training objective and identify a supervision-allocation mismatch under the globally averaged mean squared error (MSE) denoising objective: prevalent static content dominates the optimization signal, leaving sparse dynamic-object regions critical to interaction generation disproportionately under-supervised. Motivated by this observation, we introduce IMPACT, a scalable Interaction-aware Model training framework with Prior-guided Attention Calibration and Targeting. IMPACT uses cross-attention associated with manipulated-object tokens as an internal spatiotemporal prior for action-conditioned changes. It samples candidate regions from this prior, calibrates them with detached local prediction errors to construct an interaction map, and uses the map to reweight denoising supervision, requiring neither external representations nor inference-time modifications. Extensive experiments on robot-arm and human-hand manipulation, spanning diverse control modalities and DiT backbones, show that IMPACT consistently outperforms the corresponding MSE-trained baselines, improving interaction fidelity, physical plausibility, and visual quality.
1University of Science and Technology of China 2Zhongguancun Academy
3Tsinghua University 4Manifold AI
chgao96@gmail.com, chenzhibo@ustc.edu.cnattr /Border [0 0 0] user /Subtype /Link /A << /S /URI /URI (https://embodiedcity.github.io/IMPACT/) >> Project Page attr /Border [0 0 0] user /Subtype /Link /A << /S /URI /URI (https://github.com/EmbodiedCity/IMPACT.code) >> Code
Introduction
name impact.sec:introduction xyz
World models have made remarkable progress in predicting future observations (Zhang et al. 2026a; Kim et al. 2026; Fang et al. 2026b) and supporting embodied simulation (Wang et al. 2026a; Gao et al. 2025). Building on large-scale pretrained video generation backbones(Yang et al. 2024; Kong et al. 2024; Wan Team et al. 2025; NVIDIA et al. 2025), they can simulate how an environment evolves under a commanded action, generating action responses in interactive environments (AlayaWorld Team et al. 2026; Hu et al. 2026; Fang et al. 2026a) and robotic manipulation tasks (Zhang et al. 2026a; Bi et al. 2026; Zhang et al. 2026b). Their imagined futures provide scalable experience for long-horizon planning (Wang et al. 2026a; Gao et al. 2025), policy learning (Hu et al. 2025; Zhen et al. 2025; Su et al. 2026), and policy evaluation (Shang et al. 2026; Zhu et al. 2025). In this role, world models should go beyond visual coherence to simulate physically plausible interactions, yet still suffer from object deformation, discontinuous motion, weak action coupling and inconsistent contact (Zhang et al. 2026b).
As shown in attr /Border [0 0 0] goto name impact.fig:overviewFigure 1a, existing approaches typically address these failures by constraining generation with external representations of interaction dynamics. Motion-based methods use optical flow or point trajectories to describe object dynamics (Gao et al. 2025; Zhang et al. 2026b) and geometry-based methods introduce depth, surface normals, reconstructed scenes, or articulated hand meshes (Zhen et al. 2025; Kim et al. 2026). However, obtaining such spatiotemporally dense representations through auxiliary estimators or manual annotations is costly, and the resulting supervision is bounded by the accuracy of these external signals, limiting both training scalability and the achievable interaction quality.
We instead revisit the standard training objective of video world models and identify a supervision-allocation mismatch, as illustrated in attr /Border [0 0 0] goto name impact.fig:overviewFigure 1b. Inherited from general video generation, this globally averaged mean squared error (MSE) denoising objective uniformly weights all spatiotemporal positions (Po et al. 2025; Zhu et al. 2025; Kim et al. 2026). Under uniform weighting, each region contributes to optimization according to its spatial extent rather than its functional importance. Prevalent static content thus dominates the training signal, leaving sparse dynamic-object regions that carry action-conditioned changes disproportionately under-supervised. As a result, models may reduce the global denoising loss and generate visually coherent videos while leaving interaction regions under-optimized.
Our key insight is that world models built on large-scale pretrained video generation backbones already contain a spatiotemporal prior for interaction regions. For manipulation instructions, cross-attention aligns language tokens with spatiotemporal video representations, enabling the attention maps associated with manipulated-object tokens to serve as a spatiotemporal prior for regions likely to undergo action-conditioned changes. The refined prior provides a natural basis for reweighting denoising supervision, allowing sparse interaction regions to contribute to gradient optimization according to their functional importance rather than their spatial extent, thereby improving interaction generation. Because this prior is obtained from the model’s standard forward pass, it requires no external dense representations, scales readily with training data, and leaves inference unchanged.
Based on this insight, we introduce IMPACT, a scalable Interaction-aware Model training framework with Prior-guided Attention Calibration and Targeting. As shown in attr /Border [0 0 0] goto name impact.fig:overviewFigure 1c, the core idea is to treat object-conditioned cross-attention as an interaction prior, turn it into an interaction map, and use this map to reweight denoising supervision toward interaction regions, thereby mitigating the supervision-allocation mismatch and improving interaction generation. IMPACT realizes this within a single training step through two complementary components. In the forward pass, Attention Distribution Sampling (ADS) aggregates the cross-attention of the manipulated-object tokens into a proposal distribution, samples multiple candidate regions, and weights each by its detached local prediction error, calibrating the attention prior with the model’s current prediction difficulty to form a precise interaction map. In the backward pass, Interaction-Weighted Supervision (IWS) uses this map to strengthen denoising supervision on interaction regions while preserving the global objective, and routes gradients so that the interaction-weighted objective updates the non-cross-attention DiT parameters while the attention-producing parameters follow the global objective, preventing the prior from collapsing onto its own signal.
We evaluate IMPACT in two settings: robotic-arm manipulation on WorldArena (Shang et al. 2026) and human-hand manipulation on EgoDex (Hoque et al. 2025). Across both settings, IMPACT surpasses standard uniformly-supervised training (MSE) under the same backbone as well as various baselines, delivering higher interaction fidelity and physical plausibility across diverse embodied scenarios and action conditions.
In summary:
- •
We identify a supervision-allocation mismatch in world model training and introduce IMPACT, a scalable training framework that leverages the object-conditioned cross-attention prior to reweight denoising supervision toward interaction regions, without requiring any external dense representations.
- •
We realize IMPACT through two complementary designs: ADS evaluates object-conditioned attention proposals using detached local prediction errors and converts these errors into weights to form an interaction map, while IWS uses this map to target denoising supervision, preserve the global objective, and decouple region estimation from regional optimization.
- •
We demonstrate across robotic-arm and human-hand manipulation with different control signals and DiT backbones that IMPACT consistently outperforms uniform-MSE training, improving both the physical consistency and visual quality of generated videos.
Related Work
name impact.sec:related xyz
Interactive video generation and world models
Building on large-scale pretrained video generation backbones, including CogVideoX, HunyuanVideo, Wan, and Cosmos (Yang et al. 2024; Kong et al. 2024; Wan Team et al. 2025; NVIDIA et al. 2025), recent world models inherit rich representations of visual appearance and temporal dynamics, enabling high-fidelity prediction and coherent motion modeling, but these capabilities alone are insufficient to make them useful embodied simulators. To bridge this gap, recent work introduces diverse control conditions that make world models interactive, including natural-language instructions (Xiang et al. 2024; Zhang et al. 2026a; Zhao et al. 2025; Zhao et al. 2026), game and camera controls (Bruce et al. 2024; He et al. 2025), robot action trajectories (NVIDIA et al. 2025; Zhu et al. 2025), and articulated hand or body controls (Wang et al. 2026b; Gao et al. 2026). However, these methods primarily target overall visual quality and controllable generation, while providing limited constraints on interaction dynamics, resulting in physically implausible behavior within interaction regions.
External representations priors
To address this limitation, recent methods introduce spatiotemporally dense representations such as optical flow, depth maps, or reconstructed 3D structure to explicitly constrain the generation process. Motion-based approaches exploit optical flow, point trajectories, or latent temporal discrepancies to emphasize dynamic regions (Fang et al. 2025; Zhang et al. 2026b; Wu et al. 2026). Geometry-based methods introduce depth, cross-view 3D structure, or projected robot kinematics as auxiliary targets or structured conditions (Tian et al. 2026; Yang et al. 2026; Liu et al. 2025). However, constructing these representations often requires external motion, depth, segmentation, and video-understanding models, sometimes together with manually verified annotations (Fang et al. 2025; Zhang et al. 2026b; Yan et al. 2025; Luo et al. 2026). These preprocessing costs grow with the size and duration of the training corpus, thereby limiting training scalability.
Method
name impact.sec:method xyz
attr /Border [0 0 0] goto name impact.fig:frameworkFigure 2 presents an overview of IMPACT. Given a training sample, IMPACT identifies the manipulated-object tokens and uses their cross-attention as a spatial prior. Attention Distribution Sampling(ADS) samples candidate regions from this prior and weights them by detached local prediction errors to construct an interaction map, which Interaction-Weighted Supervision(IWS) uses to target denoising supervision toward interaction-relevant tokens. Gradient routing optimizes the cross-attention parameters with the original global objective and the remaining DiT parameters with the interaction-weighted objective, while all additional operations are training-only and incur no inference-time overhead.
Preliminaries
name impact.sec:preliminaries xyz
We build on a latent video diffusion transformer trained with flow matching (Lipman et al. 2023). Given a video , a frozen variational autoencoder maps it into a latent representation , where , , and denote the temporal and spatial dimensions. The conditioning information is denoted by , where is the language instruction, is the reference observation, and represents the corresponding control signal, such as hand poses or robot trajectories.
For a sampled noise level and Gaussian noise , the clean latent is interpolated with noise as
| (1) |
The diffusion transformer takes as input and predicts the velocity target . Let
| (2) |
denote the set of spatiotemporal latent positions. The prediction error at position is
| (3) |
Standard flow-matching training uniformly averages this error over all spatiotemporal positions:
| (4) |
Consequently, the contribution of a region to optimization is largely determined by its spatial extent. Small interaction regions therefore receive no additional emphasis despite containing rapid motion, contact transitions, and action-conditioned state changes. IMPACT addresses this supervision-allocation mismatch without changing the underlying flow-matching formulation.
Object-Token Grounding
name impact.sec:object_grounding xyz
The interaction region depends on the object being manipulated. We therefore begin by grounding the manipulated object in the language instruction. For an instruction , a frozen Qwen2.5-0.5B model (Qwen et al. 2025) extracts the noun phrase that denotes the manipulated object. For example, given the instruction “Right gripper stacks blue bowl on white tabletop,” the extracted phrase is “blue bowl.”
Let
| (5) |
denote the sequence of text embeddings produced by the tokenizer and text encoder. We align the extracted object phrase with this token sequence and denote its token positions by
| (6) |
When an object phrase is divided into multiple subword tokens, all matched positions are included in . This object-token set provides the semantic anchor used by ADS to extract an object-conditioned spatial distribution. The grounding stage operates only on the instruction and therefore requires neither visual annotations nor external spatial estimators.
Forward: Attention Distribution Sampling
name impact.sec:ads xyz
ADS converts the semantic object grounding into an interaction map through three operations: object-conditioned attention aggregation, candidate-region sampling, and prediction-error-based weighting.
Object-conditioned attention distribution.
Within each DiT cross-attention layer, visual queries attend to the text-token sequence. For block , attention head , visual position , and text-token position , the cross-attention probability is
| (7) |
where is the key dimension. We sum the probability mass assigned to the object-token set and average it over the selected blocks and attention heads :
| (8) |
The resulting map forms an object-conditioned proposal distribution over the video latent. Because it is obtained from the standard conditional forward pass, constructing requires no additional visual model or spatial annotation.
Candidate-region sampling.
Instead of using a single deterministic region, ADS samples spatially coherent candidates around the attention distribution. We first detach and temper the attention map:
| (9) |
where denotes stop-gradient and controls the concentration of the proposal distribution. We then transform it into logits:
| (10) |
For each candidate , we sample Gaussian noise on a coarse spatiotemporal grid and trilinearly interpolate it to the resolution of , producing a smooth perturbation field . A soft candidate is generated as
| (11) |
where controls the perturbation magnitude and is the sampling temperature. Sampling noise on a coarse grid produces coherent spatial variations rather than independent token-wise perturbations.
To prevent candidate scores from being dominated by differences in region size, every candidate is converted into a fixed-area binary region. Specifically, we retain the largest values of :
| (12) |
yielding . All candidates consequently cover the same number of latent positions and remain directly comparable.
Prediction-error-based weighting.
ADS evaluates every candidate using the mean local prediction error within that region:
| (13) |
A higher value of indicates that the candidate covers content that is currently more difficult for the model to predict. The candidate errors are converted into normalized weights:
| (14) |
where controls the concentration of the weighting distribution. The candidate regions are then weighted to form the interaction map
| (15) |
Through this process, the object-conditioned attention determines where candidate regions are sampled, while the current local prediction error determines their relative contribution to . The entire ADS pipeline is computed without gradient tracking, preventing the model from directly modifying the proposal distribution or candidate weights to reduce the regional objective.
In our implementation, we use , an attention tempering exponent , perturbation scale , sampling temperature , and a coarse noise grid of size . Each candidate retains of the spatiotemporal latent positions, and candidate errors are converted into weights using .
Backward: Interaction-Weighted Supervision
name impact.sec:iws xyz
IWS uses the interaction map to increase the contribution of interaction-relevant positions during denoising optimization. For each position , we define
| (16) |
where controls the maximum regional emphasis. Because , the resulting weight satisfies . Positions outside the estimated interaction region retain their original unit weight, while positions with high interaction-map values receive stronger supervision.
The interaction-weighted denoising objective is
| (17) |
The normalization by the total weight prevents the loss scale from growing with the size or magnitude of the emphasized region. Unlike a masked regional loss, attr /Border [0 0 0] goto name impact.eq:iws_objectiveEquation 17 retains supervision over the complete spatiotemporal field and only changes its spatial allocation. We use in the main experiments.
Gradient-decoupled optimization.
The interaction map is derived from cross-attention, creating a dependency between region estimation and the objective guided by that region. Although is detached, allowing to update the attention-producing parameters could still cause the cross-attention representation itself to adapt to the regional objective in subsequent training steps. We therefore separate the DiT parameters into the cross-attention parameters and all remaining parameters . Their gradients are routed according to
| (18) |
and
| (19) |
In practice, this routing is implemented through two backward passes with parameter-group gradient hooks. The global backward pass retains gradients only for , while the IWS backward pass retains gradients only for . Consequently, the cross-attention used to estimate interaction regions remains governed by the original uniformly supervised objective, whereas the remaining DiT parameters learn from the spatially reallocated supervision. This separation prevents the regional objective from directly optimizing its own localization signal and decouples interaction-region estimation from interaction-focused optimization.
Training and inference cost.
IMPACT reuses the cross-attention probabilities and token-wise prediction errors already produced during standard world-model training. Its additional computation consists primarily of sampling candidate masks, evaluating their masked mean errors, and constructing the interaction-weighted objective. No component of ADS or IWS is used at inference time, so the trained world model preserves the original architecture and inference procedure.
Experiments
name impact.sec:experiments xyz
| Model | EWMScore | Visual Quality | Motion Quality | Content Consist. | Physics Adher. | 3D Accuracy | Control- lability |
|---|---|---|---|---|---|---|---|
| General world models | |||||||
| CogVideoX | 57.90 | 55.81 | 42.49 | 68.12 | 47.33 | 84.62 | 54.40 |
| Veo 3.1 | 58.87 | 56.44 | 46.12 | 66.12 | 45.52 | 78.48 | 62.62 |
| Wan 2.6 | 61.86 | 61.62 | 68.31 | 60.36 | 42.31 | 75.88 | 60.85 |
| Embodied world models | |||||||
| GigaWorld-0 | 53.39 | 44.82 | 58.79 | 58.74 | 34.60 | 69.56 | 52.94 |
| Genie Envisioner | 43.65 | 29.78 | 49.17 | 62.63 | 13.66 | 69.73 | 35.60 |
| Vidar | 51.60 | 46.07 | 40.55 | 60.93 | 36.38 | 77.32 | 51.86 |
| IRASim | 58.12 | 54.81 | 44.25 | 69.67 | 46.48 | 85.50 | 53.26 |
| CtrlWorld | 59.70 | 55.33 | 50.28 | 63.99 | 54.89 | 86.30 | 54.65 |
| Representation-guided models | |||||||
| TesserAct | 53.23 | 41.64 | 50.59 | 66.60 | 35.98 | 75.39 | 50.82 |
| RoboMaster | 51.84 | 34.32 | 48.49 | 69.25 | 32.61 | 79.62 | 49.62 |
| WoW | 54.88 | 52.98 | 50.02 | 64.52 | 38.11 | 74.77 | 49.89 |
| Wan 2.2-based models | |||||||
| Wan 2.2 | 50.79 | 51.41 | 42.12 | 54.02 | 34.05 | 77.14 | 49.22 |
| Wan 2.2-AC | 58.65 | 56.57 | 48.23 | 60.55 | 53.41 | 86.16 | 54.40 |
| Wan 2.2-AC + IMPACT | 62.46 | 60.60 | 53.28 | 60.29 | 55.87 | 92.56 | 60.00 |
| Cosmos-based models | |||||||
| Cosmos-Predict 2.5 (text) | 50.81 | 47.65 | 60.32 | 57.94 | 23.44 | 75.08 | 39.38 |
| Cosmos-Predict 2.5 (action) | 55.91 | 57.87 | 38.89 | 68.73 | 42.23 | 82.53 | 49.51 |
| Cosmos-Predict 2.5 (action) + IMPACT | 62.53 | 56.32 | 68.22 | 57.90 | 41.82 | 77.99 | 71.19 |
| Visual | Hand interaction | |||
| Model | FVD | FID | CLIP- Hand | Hand IoU |
| General world models | ||||
| HunyuanVideo-1.5 | 541.83 | 56.73 | 0.902 | 0.328 |
| Cosmos-Predict 2.5 | 615.42 | 50.12 | 0.914 | 0.386 |
| Pose control | ||||
| MimicMotion | 612.75 | 48.55 | 0.882 | 0.492 |
| MagicPose | 1456.20 | 212.94 | 0.864 | 0.298 |
| VACE | 358.42 | 50.65 | 0.895 | 0.493 |
| LOME | 1748.29 | 66.03 | 0.745 | 0.087 |
| Wan 2.2-based models | ||||
| Wan 2.2 | 1463.05 | 199.35 | 0.876 | 0.557 |
| Wan 2.2-AC | 366.12 | 44.71 | 0.921 | 0.693 |
| Wan 2.2-AC + IMPACT | 110.94 | 5.79 | 0.952 | 0.772 |
Setup
name impact.sec:setup xyz
Implementation details.
We evaluate IMPACT in two manipulation settings to demonstrate the generality of our method across interaction types, applying it to the Wan2.2 and Cosmos-Predict 2.5 backbone in both. For robot-arm manipulation, the model is trained on 350K 17-frame videos collected from RoboTwin (Chen et al. 2025), conditioned on 14-DoF dual-arm action trajectories, injected through an additional action encoder. For human-hand manipulation, we train on 256K clips drawn from EgoDex (Hoque et al. 2025) across 118 tasks, using 81-frame videos conditioned on hand-pose videos temporally aligned with the RGB frames, injected through the shared VAE encoder without any additional encoder. Both settings are trained at 720p and optimization is identical across the two settings: we use bf16 mixed precision with FSDP over 8 GPUs, a per-device batch size of 1 with 4-step gradient accumulation (effective global batch size ), a constant learning rate of after 100 warm-up steps, and we train for one epoch. IMPACT hyperparameters are fixed across both settings: , , coarse grid , , , , , .
Benchmarks and metrics.
For robot-arm manipulation, we evaluate on WorldArena (Shang et al. 2026), which scores dual-arm manipulation along six dimensions: Visual Quality, Motion Quality, Content Consistency, Physics Adherence, 3D Accuracy, and Controllability, spanning 16 normalized metrics, and condenses overall generation quality into a single EWMScore (the mean of the 16 metrics). We report EWMScore together with the six aggregate dimensions in attr /Border [0 0 0] goto name impact.tab:mainTable 1, and provide all 16 metrics in the technical appendix. For human-hand manipulation, we evaluate on the EgoDex (Hoque et al. 2025) test set along two axes (attr /Border [0 0 0] goto name impact.tab:egodexTable 2): visual metrics: FVD (Unterthiner et al. 2018) and FID (Heusel et al. 2017), covering temporal coherence and per-frame appearance quality; and hand-interaction metrics: CLIP-Hand for the local appearance and semantics of the hand and nearby manipulated object, and Hand IoU for the coarse 2D position and scale of the generated hand (Sun et al. 2026).
Baselines.
For robot-arm manipulation, all models follow the WorldArena evaluation settings. Existing baselines comprise general world models: CogVideoX (Yang et al. 2024), Wan 2.6 (Wan Team et al. 2025), and Veo 3.1; embodied world models: GigaWorld-0, Genie Envisioner, Vidar, IRASim (Zhu et al. 2025), and CtrlWorld; and representation-guided models: TesserAct (Zhen et al. 2025), RoboMaster, and WoW. We additionally report two backbone-specific groups to evaluate IMPACT: Wan 2.2-based models include Wan 2.2 (Wan Team et al. 2025), MSE-trained Wan 2.2-AC, and Wan 2.2-AC with IMPACT; Cosmos-based models include Cosmos-Predict 2.5 (text) (NVIDIA et al. 2025), Cosmos-Predict 2.5 (action), and its IMPACT variant. For human-hand manipulation, baselines comprise general video world models: HunyuanVideo-1.5 (Wu et al. 2025) and Cosmos-Predict 2.5 (NVIDIA et al. 2025) and pose-controlled models: MimicMotion (Zhang et al. 2025), MagicPose (Chang et al. 2024), VACE (Jiang et al. 2025), and LOME (Gao et al. 2026).


Quantitative Analysis
name impact.sec:quantitative_analysis xyz
Robot-arm manipulation.
attr /Border [0 0 0] goto name impact.tab:mainTable 1 shows that IMPACT improves both backbone families. On Cosmos-Predict 2.5 (action), it raises EWMScore from 55.91 to 62.53 (+6.62 points, 11.8%), achieving the best overall result, the best Controllability (71.19), and the second-best Motion Quality (68.22). Interaction Quality increases from 0.5500 to 0.6360 and Action Following from 0.0133 to 0.6260, although the remaining aggregate dimensions decline. On Wan 2.2-AC, IMPACT raises EWMScore from 58.65 to 62.46 (+3.81 points, 6.5%) and improves Visual Quality, Motion Quality, Physics Adherence, 3D Accuracy, and Controllability, attaining the best Physics Adherence (55.87) and 3D Accuracy (92.56) and the second-best Visual Quality (60.60). The best IMPACT result exceeds Wan 2.6, CtrlWorld, and WoW by 0.67, 2.83, and 7.65 points, respectively.
Human-hand manipulation.
For human-hand manipulation, attr /Border [0 0 0] goto name impact.tab:egodexTable 2 reports that IMPACT attains the best results on visual metrics across both general video world models and pose-controlled models, cutting FVD from 366.12 to 110.94 and FID from 44.71 to 5.79 over the action-conditioned Wan 2.2-AC on the same backbone. It also leads on both hand-interaction metrics, improving CLIP-Hand from 0.921 to 0.952 and Hand IoU from 0.693 to 0.772 over Wan 2.2-AC on the same backbone. These gains indicate stronger local interaction fidelity and hand localization than both the MSE-trained counterpart and the pose-controlled baselines. Together with the robot-arm results, these improvements demonstrate that IMPACT delivers consistent gains across different DiT backbones and control signals.
Qualitative analysis
name impact.sec:Qualitative analysis xyz
Generation comparison.
attr /Border [0 0 0] goto name impact.fig:qualitativeFigure 4 compares generations in robot-arm and human-hand manipulation. In both cases, the baselines follow the instruction loosely and tend to blur or distort the contact region, whether the gripper–object contact for the robot arm or the hand–object contact for the human hand, and some drift away from the target. In contrast, IMPACT produces the specified interaction with a sharper contact region and more coherent object dynamics, while keeping the surrounding scene stable. This clearly demonstrates the effectiveness of IMPACT for interaction-region generation.
ADS calibration.
attr /Border [0 0 0] goto name impact.fig:ads_calibrationFigure 4 illustrates how ADS calibrates the attention prior in a representative training example. The raw object-conditioned cross-attention map is diffuse, spreading across the robot arm, workspace, and background. ADS samples candidate regions from this prior and evaluates them using detached local prediction errors (attr /Border [0 0 0] goto name impact.eq:candidate_scoreEq. 13). Candidates covering the contact region receive higher weights than those dominated by the static background. Their weighted aggregation (attr /Border [0 0 0] goto name impact.eq:interaction_mapEq. 15) produces an interaction map that is more concentrated around the interacting arm, gripper, and manipulated object. This example illustrates how ADS refines a coarse attention prior into a more targeted supervision map for IWS.
Ablation Studies
name impact.sec:ablation xyz
Component ablation: IWS and ADS.
attr /Border [0 0 0] goto name impact.tab:ads_samplingTable 3 shows the effect of separating the components on WorldArena and reveals their complementarity. Starting from the AC backbone (58.65 EWMScore), IWS provides the more direct gain (+2.89 to 61.54), since it primarily addresses the supervision-allocation mismatch. On top of this, ADS further calibrates the cross-attention prior and delivers an additional improvement (+0.92 to 62.46) over weighting the raw, coarse prior directly. Integrated within a single training step, the two components jointly strengthen interaction-region generation, improving EWMScore by 3.81 points overall.
| Method | EWMScore | Visual Quality | Motion Quality | Physics Adher. |
|---|---|---|---|---|
| Wan 2.2-AC | 58.65 | 56.57 | 48.23 | 53.41 |
| + IWS | 61.54 | 60.44 | 49.16 | 55.16 |
| + IWS + ADS (IMPACT) | 62.46 | 60.60 | 53.28 | 55.87 |
Conclusion
name impact.sec:conclusion xyz
In this work, we introduced IMPACT, a scalable framework that addresses the supervision-allocation mismatch by converting object-conditioned cross-attention into targeted denoising supervision. ADS calibrates attention proposals with detached local prediction errors, while IWS reweights training with the resulting interaction map, requiring neither external spatial signals nor inference-time changes. Experiments on robot-arm and human-hand manipulation show consistent gains over uniform MSE training and strong baselines, demonstrating the effectiveness and scalability of IMPACT for interaction-aware world model training.
References
- AlayaWorld: long-horizon and playable video world generation. External Links: 2607.06291, Link Cited by: Introduction.
- Motus: a unified latent action world model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 35101–35113. External Links: Link Cited by: Introduction.
- Genie: generative interactive environments. arXiv preprint arXiv:2402.15391. Cited by: Interactive video generation and world models.
- MagicPose: realistic human poses and facial expressions retargeting with identity-aware diffusion. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, pp. 6263–6285. External Links: Link Cited by: Baselines..
- RoboTwin 2.0: a scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation. arXiv preprint arXiv:2506.18088. Cited by: Implementation details..
- iWorld-Bench: a benchmark for interactive world models with a unified action generation framework. External Links: 2605.03941, Link Cited by: Introduction.
- Worldscape-MoE: a unified mixture-of-experts world model for scalable heterogeneous action control. External Links: 2607.03964, Link Cited by: Introduction.
- Robotic VLA benefits from joint learning with motion image diffusion. arXiv preprint arXiv:2512.18007. Cited by: External representations priors.
- FLIP: flow-centric generative planning as general-purpose manipulation world model. In International Conference on Learning Representations, External Links: Link Cited by: Introduction, Introduction.
- LOME: learning human-object manipulation with action-conditioned egocentric world model. External Links: 2603.27449, Link Cited by: Interactive video generation and world models, Baselines..
- Matrix-Game 2.0: an open-source, real-time, and streaming interactive world model. External Links: 2508.13009, Link Cited by: Interactive video generation and world models.
- GANs trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems, Cited by: Benchmarks and metrics..
- EgoDex: learning dexterous manipulation from large-scale egocentric video. External Links: 2505.11709, Link Cited by: Introduction, Implementation details., Benchmarks and metrics..
- Multiplayer interactive world models with representation autoencoders. External Links: 2607.05352, Link Cited by: Introduction.
- Video prediction policy: a generalist robot policy with predictive visual representations. In Proceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 267, pp. 24328–24346. External Links: Link Cited by: Introduction.
- VACE: all-in-one video creation and editing. arXiv preprint arXiv:2503.07598. Cited by: Baselines..
- Dexterous world models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 29663–29673. External Links: Link Cited by: Introduction, Introduction, Introduction.
- HunyuanVideo: a systematic framework for large video generative models. arXiv preprint arXiv:2412.03603. Cited by: Introduction, Interactive video generation and world models.
- Flow matching for generative modeling. In International Conference on Learning Representations, Cited by: Preliminaries.
- Geometry-aware 4d video generation for robot manipulation. arXiv preprint arXiv:2507.01099. Cited by: External representations priors.
- CoInteract: physically-consistent human-object interaction video synthesis via spatially-structured co-generation. arXiv preprint arXiv:2604.19636. Cited by: External representations priors.
- Cosmos world foundation model platform for physical ai. arXiv preprint arXiv:2501.03575. Cited by: Introduction, Interactive video generation and world models, Baselines..
- Long-context state-space video world models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 8733–8744. External Links: Link Cited by: Introduction.
- Qwen2.5 technical report. arXiv preprint arXiv:2412.15115. External Links: Link Cited by: Object-Token Grounding.
- WorldArena: a unified benchmark for evaluating perception and functional utility of embodied world models. External Links: 2602.08971, Link Cited by: Introduction, Introduction, Benchmarks and metrics..
- WorldScape Policy 2.0: empowering steerable world action modeling with reasoning-augmented memory. External Links: 2607.18840, Link Cited by: Introduction.
- HandWorld: hand-centric unified video action generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15976–15985. Cited by: Benchmarks and metrics..
- STARRY: spatial-temporal action-centric world modeling for robotic manipulation. arXiv preprint arXiv:2604.26848. Cited by: External representations priors.
- Towards accurate generative models of video: a new metric and challenges. arXiv preprint arXiv:1812.01717. Cited by: Benchmarks and metrics..
- Wan: open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314. Cited by: Introduction, Interactive video generation and world models, Baselines..
- Lifting embodied world models for planning and control. External Links: 2604.26182, Link Cited by: Introduction.
- Hand2World: autoregressive egocentric interaction generation via free-space hand gestures. External Links: 2602.09600, Link Cited by: Interactive video generation and world models.
- HunyuanVideo 1.5 technical report. External Links: 2511.18870, Link Cited by: Baselines..
- Latent temporal discrepancy as motion prior: a loss-weighting strategy for dynamic fidelity in t2v. arXiv preprint arXiv:2601.20504. Cited by: External representations priors.
- Pandora: towards general world model with natural language actions and video states. External Links: 2406.09455, Link Cited by: Interactive video generation and world models.
- Open-world hand-object interaction video generation based on structure and contact-aware representation. arXiv preprint arXiv:2512.01677. Cited by: External representations priors.
- EA-WM: event-aware generative world model with structured kinematic-to-visual action fields. arXiv preprint arXiv:2605.06192. Cited by: External representations priors.
- CogVideoX: text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072. Cited by: Introduction, Interactive video generation and world models, Baselines..
- Qwen-RobotWorld technical report: unifying embodied world modeling through language-conditioned video generation. External Links: 2606.17030, Link Cited by: Introduction, Interactive video generation and world models.
- PhysisForcing: physics reinforced world simulator for robotic manipulation. External Links: 2606.28128, Link Cited by: Introduction, Introduction, External representations priors.
- MimicMotion: high-quality human motion video generation with confidence-aware pose guidance. In Proceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 267, pp. 74896–74910. External Links: Link Cited by: Baselines..
- AirScape: an aerial generative world model with motion controllability. External Links: 2507.08885, Link Cited by: Interactive video generation and world models.
- WorldVLN: autoregressive world action model for aerial vision-language navigation. External Links: 2605.15964, Link Cited by: Interactive video generation and world models.
- Learning 4d embodied world models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 5337–5347. External Links: Link Cited by: Introduction, Introduction, Baselines..
- IRASim: a fine-grained world model for robot manipulation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 9834–9844. External Links: Link Cited by: Introduction, Introduction, Interactive video generation and world models, Baselines..
A. Data Construction and Process Details
name impact.sec:data_details xyz
A.1 RoboTwin Data
name impact.sec:robotwin_data xyz
Data source and modalities.
The robot-arm training corpus comprises simulated bimanual manipulation rollouts from RoboTwin 2.0, generated with the SAPIEN physics engine and the dual-arm Aloha-AgileX embodiment. Each episode provides synchronized RGB observations, camera calibration, robot states, end-effector poses, gripper states, language instructions, and planned action trajectories. We use the head-camera stream at resolution and 10 fps, obtained by sampling every third frame from the 30 fps source video. The synchronized action at each frame is
| (20) |
where and denote the six joint positions of the left and right arms, and and denote their gripper coordinates.
Task and visual diversity.
The training set contains both clean and domain-randomized rollouts. The randomized regime introduces collision-aware distractors, samples table and background appearances from more than 12,000 textures, varies illumination, perturbs the tabletop height by up to 0.03 m, and displaces the head camera. Clean backgrounds and extreme illumination each occur in 2% of randomized rollouts. The language instructions also vary across equivalent task executions.
The corpus covers 50 tasks spanning pick-and-place, container placement, stacking, object ranking, tool use, switch operation, articulated-object manipulation, bimanual handover, and coordinated dual-arm manipulation. The task distribution is non-uniform and ranges from 99 to 28,117 training windows per task. The four largest tasks are color-based block ranking (28,117 windows), size-based block ranking (27,170), bottle placement into a container (16,881), and block handover (16,723).
Window construction and action normalization.
Each training example contains 17 consecutive frames with stride one and the corresponding 17 action vectors. Boundary examples are temporally resampled or padded to preserve this fixed length. For action dimension , we normalize using the first and ninety-ninth dataset percentiles, and :
| (21) |
RGB frames and actions always share the same temporal indices. Spatial bucket sampling and matched crop-and-resize augmentation produce the 720p training inputs.
A.2 EgoDex Data
name impact.sec:egodex_data xyz
Data source and modalities.
The human-hand training corpus contains approximately 256K egocentric clips from EgoDex spanning 118 tasks. Each clip is paired with camera intrinsics, per-frame camera poses, articulated hand transforms and a language instruction. Each RGB clip is paired with a rendered hand-pose video constructed from the capture-time 3D hand tracks.
The tasks cover tabletop setup and cleanup, pick-and-place, stacking, assembly and disassembly, folding and wrapping, cooking, washing, device insertion and removal, writing, drawing, and tool use.
Temporal clip construction.
Each source clip is converted into an 81-frame training clip. Clips longer than 81 frames are represented by 81 uniformly spaced indices including both endpoints; clips shorter than 81 frames repeat the final index. The same index sequence is applied to RGB and hand-pose videos, preserving frame-level correspondence between the target and condition. Source videos are at 30 fps, and contiguous 81-frame clips therefore span 2.7 seconds. Training uses matched spatial transformations at 720p.
A.3 Object-Token Grounding
name impact.sec:grounding_details xyz
Manipulated-object extraction.
We annotate each training example once with Qwen2.5-0.5B-Instruct. For both RoboTwin and EgoDex, the task instruction is the sole annotation input; no video frame, pose signal, temporal metadata, or external visual annotation is used. The model extracts compact noun phrases for the manipulated objects while preserving discriminative attributes such as color, material, shape, label, or container type. Agents, body parts, cameras, supporting surfaces, backgrounds, and non-interacting scene elements are excluded.
Complete annotation prompt.
RoboTwin and EgoDex use the same prompt template below, with {instruction} replaced by the task instruction for the current training clip.
System message
You are a precise language annotation assistant for manipulation
training. Given only a task instruction, identify the concrete physical
objects directly manipulated by the agent. Return only valid JSON, without
Markdown, explanations, or code fences.
Use compact English noun phrases and preserve discriminative attributes
stated in the instruction, including color, material, shape, printed label,
object part, and container type. Exclude the robot, hands, fingers, arms,
people, cameras, frames, tables, workspaces, backgrounds, and generic scene
regions unless the instruction explicitly identifies one of them as the
manipulated object. Do not infer objects that are not stated in the
instruction.
User message
Task instruction:
{instruction}
Return the following JSON object:
{
“interacting_objects”: [
{“object”: “compact manipulated-object phrase”}
]
}
Include every manipulated object explicitly stated in the instruction.
Preserve the order in which the objects appear. If the instruction refers to
the same object more than once, return it once. Return JSON only.
Mapping object phrases to text tokens.
The extracted object phrase is aligned with its occurrence in the original task instruction after tokenization. All subword tokens associated with the object phrase are included in the object-token set . We then aggregate the cross-attention associated with across attention heads and transformer layers to obtain the object-conditioned attention map used by ADS.
| Setting | Task instruction | Extracted object phrase |
|---|---|---|
| Robot arm | “Take the bottle with white printed label from the table and keep upright.” | white printed-label bottle. |
| Human hand | “Pick up the orange plate from the table.” | orange plate. |
A.4 Hand-Pose Video Construction
name impact.sec:hand_pose_construction xyz
Pose representation.
EgoDex provides capture-time 3D hand and finger tracks together with per-frame camera calibration. We use these annotations directly and do not estimate hand pose from the compressed RGB video. Each hand is represented by 21 three-dimensional joints, accompanied by a presence mask. The camera metadata provides a world-to-camera transformation and four intrinsic parameters for every frame.
Projection and rendering.
For frame , a homogeneous world point is transformed by the world-to-camera matrix and projected using the intrinsics :
| (22) |
Points with depth are excluded. The 21 joints follow the wrist, thumb, index, middle, ring, and little-finger ordering. Each finger is connected to the wrist and rendered as an anti-aliased skeleton on a black background. We draw green edges with width 2, blue joints with radius 3, and fingertips with radius 5. Pose videos preserve the RGB resolution and are encoded at 30 fps using H.264 with the YUV420p pixel format and a constant-rate factor of 23.
Temporal and spatial alignment.
RGB and pose videos contain the same number of decoded frames and use the same 81-frame index sequence for every training segment. The two modalities therefore remain aligned by frame index throughout temporal sampling. Identical crop and resize parameters are subsequently applied to both modalities.
B. Implementation Details
name impact.sec:implementation_details xyz
B.1 Robot-Arm Models
name impact.sec:robot_architecture xyz
Backbone.
We instantiate the robot-arm setting with two action-conditioned diffusion transformer backbones: Wan 2.2 TI2V 5B and Cosmos-Predict 2.5 (action). The Wan 2.2 backbone contains 30 transformer blocks with hidden width , 24 attention heads, and an FFN width of 14,336. It receives a 17-frame RGB sequence together with the first-frame visual condition and the synchronized 14-DoF action trajectory. The second model is initialized from Cosmos-Predict2.5-2B and retains its action-conditioning interface for the same robot-arm setting.
Action encoder.
Let denote the normalized action trajectory. We flatten the complete trajectory into 238 scalars and process it with two independent multilayer perceptrons of identical topology:
| (23) |
with hidden width . The first encoder produces a -dimensional vector that is added to the timestep embedding. The second produces values, reshaped to , which modulate the six adaptive normalization components in every transformer block. Learned binary condition embeddings distinguish action-conditioned and action-dropped examples. Action information is therefore injected globally through the timestep and adaptive normalization pathways.
IMPACT configuration.
For both backbones, IMPACT constructs its attention anchor from the object-token cross-attention produced by the native transformer. ADS samples candidate masks by adding Gaussian perturbations with to the logit of the detached attention anchor on a coarse grid. The perturbed maps are trilinearly upsampled to the latent resolution, transformed with sampling temperature and anchor power , and thresholded to retain the top of positions. Detached regional MSEs score the candidate masks, and centered scores are converted into aggregation weights with coefficient . IWS uses the resulting soft interaction map to assign per-position weights in , with .
B.2 Human-Hand Model
name impact.sec:human_architecture xyz
Pose-conditioned input.
The human-hand model uses the same 5B transformer and conditions on an 81-frame hand-pose video. RGB targets, the first-frame reference, and pose frames are encoded by the shared Wan VAE without an additional pose encoder. The VAE produces 48-channel latents with temporal and spatial compression factors of 4, 16, and 16, yielding 21 latent timesteps for an 81-frame sequence.
The first latent timestep retains the RGB reference condition. At subsequent timesteps, the 48-channel reference latent is replaced by the temporally aligned 48-channel pose latent. Four binary mask channels are concatenated to form a 52-channel condition . Before 3D patch embedding, the model concatenates the 48-channel noisy video latent with , giving 100 input channels:
| (24) |
The action mask applies pose replacement only to conditioned samples. RGB and pose inputs share the same temporal indices and spatial transformation.
C. Evaluation Details
name impact.sec:evaluation_details xyz
C.1 Evaluation Protocols and Baseline Versions
name impact.sec:baseline_protocol xyz
WorldArena evaluates 500 held-out episodes from 50 RoboTwin 2.0 tasks. All decoded submissions have a minimum resolution of and a frame rate of 24 fps. Text-conditioned submissions contain 121 frames. Action-conditioned submissions follow the benchmark action sequence and match the corresponding reference trajectory length. EgoDex models are evaluated at their release-specific inference settings, as detailed in attr /Border [0 0 0] goto name impact.tab:baseline_specsTable 5.
| Method | Version | Resolution | Sec. / FPS |
| General video generation | |||
| HunyuanVideo-1.5 | 25.11.20 | 5 / 24 | |
| Cosmos-Predict 2.5 | 25.10.06 | 5 / 16 | |
| Pose-controlled video generation | |||
| MimicMotion | 24.07.08 | 4.8 / 15 | |
| MagicPose | 24.04.03 | 4 / 15 | |
| VACE | 25.03.11 | 5 / 16 | |
| LOME | 26.04.05 | 5 / 15 | |
| Wan 2.2-based models | |||
| Wan 2.2 | 25.07.28 | 5 / 24 | |
| Wan 2.2-AC / +IMPACT | – | 3.4 / 24 | |
Metric computation uses the ordered decoded frames from each output. For Wan 2.2-AC and IMPACT, the 81 output frames and the pose condition use identical temporal indices. Each model uses its official sampling schedule and guidance configuration.
C.2 WorldArena Metric Definitions
name impact.sec:worldarena_metrics xyz
WorldArena normalizes each raw metric to using empirically selected boundaries and reports EWMScore as times the arithmetic mean of the 16 normalized metrics. Thus, all displayed entries are higher-is-better, including metrics whose underlying raw quantity is an error.
| Dimension | Metric | implementation |
|---|---|---|
| Visual | IQ (Image Quality) | Frame clarity and distortion quality from the no-reference MUSIQ image-quality model; averaged over frames. |
| AQ (Aesthetic Quality) | Per-frame visual appeal (lighting, color, composition) from the LAION aesthetic predictor; averaged over frames. | |
| JS (JEPA Similarity) | Distributional similarity between generated and ground-truth V-JEPA video features, computed using MMD with a second-order polynomial kernel. | |
| Motion | DD (Dynamic Degree) | Salient motion intensity: RAFT flow between adjacent frames, averaging the top 5% motion magnitudes and applying a resolution-adaptive sigmoid. |
| FS (Flow Score) | Overall motion intensity: mean RAFT optical-flow magnitude over all pixels and adjacent-frame pairs. | |
| MS (Motion Smoothness) | Temporal smoothness: SSIM between each actual middle frame and a VFI-Mamba interpolation from its neighbors, weighted by log motion magnitude to avoid rewarding static video. | |
| Content | SC (Subject Consistency) | Subject identity/structure stability from DINO cosine similarity of each frame to both the first and previous frame, penalized for near-static video. |
| BC (Background Consistency) | Global background/scene stability from CLIP image-feature cosine similarity to the first and previous frames, with the same low-motion penalty. | |
| PC (Photometric Consistency) | Pixel-level texture stability from forward–backward SEA-RAFT flow cycle error (raw AEPE is lower-better; the reported consistency score is inverted/normalized). | |
| Physics | Inter. (Interaction Quality) | Qwen3-VL 1–5 judgment of physically plausible robot–object contact, force transmission, and interaction, divided by five. |
| Traj. (Trajectory Accuracy) | Robot-arm trajectory agreement with ground truth: SAM 3 arm boxes/trajectories compared through normalized dynamic time warping. | |
| 3D | Depth (Depth Accuracy) | Monocular depth agreement to ground truth after per-video median-scale alignment; based on depth error and converted to a higher-is-better accuracy. Up to 40 frames are sampled uniformly. |
| Persp. (Perspectivity) | Qwen3-VL judgment of 3D plausibility, including scale-versus-depth, lighting, and occlusion relations. | |
| Control | Instr. (Instruction Following) | Qwen3-VL judgment of whether action type, target object, and resulting task state follow the instruction. |
| Sem. (Semantic Alignment) | Cosine similarity between Qwen2.5-VL descriptions of generated and reference videos. | |
| Act. (Action Following) | Response diversity under three distinct instructions sharing one initial frame; average pairwise feature dissimilarity between the three generated videos. |
Full 16-metric results.
name impact.sec:full_worldarena_results xyz
Tables attr /Border [0 0 0] goto name impact.tab:worldarena_full_quality8 and attr /Border [0 0 0] goto name impact.tab:worldarena_full_task8 report all 16 normalized WorldArena metrics for the same models and ordering. The first covers the generation-quality dimensions (visual quality, motion quality, content consistency) and the overall EWMScore; the second covers the task-oriented dimensions (physics adherence, 3D accuracy, controllability).
| Model | EWMScore | Visual Quality | Motion Quality | Content Consistency | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| IQ | AQ | JS | DD | FS | MS | SC | BC | PC | ||
| General world models | ||||||||||
| CogVideoX | 57.90 | 0.3582 | 0.3777 | 0.9384 | 0.3166 | 0.2189 | 0.7391 | 0.8083 | 0.8773 | 0.3580 |
| Veo 3.1 | 58.87 | 0.6605 | 0.4632 | 0.5694 | 0.5450 | 0.1396 | 0.6989 | 0.7878 | 0.8710 | 0.3247 |
| Wan 2.6 | 61.86 | 0.6824 | 0.4433 | 0.7229 | 0.7421 | 0.4532 | 0.8539 | 0.7517 | 0.8687 | 0.1904 |
| Embodied world models | ||||||||||
| GigaWorld-0 | 53.39 | 0.5041 | 0.3991 | 0.4413 | 0.6709 | 0.3118 | 0.7811 | 0.7303 | 0.8563 | 0.1756 |
| Genie Envisioner | 43.65 | 0.2305 | 0.3289 | 0.3340 | 0.6930 | 0.0855 | 0.6966 | 0.7760 | 0.9024 | 0.2006 |
| Vidar | 51.60 | 0.4145 | 0.4068 | 0.5608 | 0.2767 | 0.1426 | 0.7973 | 0.7629 | 0.8300 | 0.2350 |
| IRASim | 58.12 | 0.3489 | 0.3623 | 0.9330 | 0.4139 | 0.2083 | 0.7052 | 0.8312 | 0.9068 | 0.3522 |
| CtrlWorld | 59.70 | 0.3522 | 0.3893 | 0.9185 | 0.4257 | 0.3449 | 0.7377 | 0.8411 | 0.9057 | 0.1729 |
| Representation-guided models | ||||||||||
| TesserAct | 53.23 | 0.3322 | 0.4590 | 0.4579 | 0.5150 | 0.2447 | 0.7579 | 0.8250 | 0.9238 | 0.2491 |
| RoboMaster | 51.84 | 0.3487 | 0.3842 | 0.2966 | 0.6124 | 0.1484 | 0.6940 | 0.8295 | 0.9123 | 0.3356 |
| WoW | 54.88 | 0.4587 | 0.3868 | 0.7440 | 0.4608 | 0.2706 | 0.7692 | 0.8161 | 0.9025 | 0.2170 |
| Wan 2.2-based models | ||||||||||
| Wan 2.2 | 50.79 | 0.3884 | 0.3963 | 0.7575 | 0.4349 | 0.1269 | 0.7019 | 0.7400 | 0.8000 | 0.0806 |
| Wan 2.2-AC | 58.65 | 0.4053 | 0.3575 | 0.9343 | 0.4262 | 0.2488 | 0.7719 | 0.8250 | 0.9095 | 0.0821 |
| Wan 2.2-AC + IMPACT (w/o ADS) | 61.54 | 0.5236 | 0.4078 | 0.8819 | 0.4381 | 0.2611 | 0.7756 | 0.8233 | 0.8970 | 0.1090 |
| Wan 2.2-AC + IMPACT | 62.46 | 0.4906 | 0.4223 | 0.9050 | 0.4630 | 0.3328 | 0.8026 | 0.8191 | 0.8953 | 0.0944 |
| Cosmos-based models | ||||||||||
| Cosmos-Predict 2.5 (text) | 50.81 | 0.6668 | 0.4501 | 0.3126 | 0.5911 | 0.4302 | 0.7882 | 0.7488 | 0.8511 | 0.1383 |
| Cosmos-Predict 2.5 (action) | 55.91 | 0.4489 | 0.3576 | 0.9296 | 0.3994 | 0.0573 | 0.7100 | 0.8197 | 0.8894 | 0.3528 |
| + IMPACT | 62.53 | 0.5588 | 0.3941 | 0.7366 | 0.5810 | 0.5816 | 0.8839 | 0.8026 | 0.8862 | 0.0482 |
| Model | Physics Adherence | 3D Accuracy | Controllability | ||||
|---|---|---|---|---|---|---|---|
| Inter. | Traj. | Depth | Persp. | Instr. | Sem. | Act. | |
| General world models | |||||||
| CogVideoX | 0.5940 | 0.3526 | 0.9097 | 0.7828 | 0.7268 | 0.8977 | 0.0076 |
| Veo 3.1 | 0.7872 | 0.1231 | 0.7421 | 0.8276 | 0.9328 | 0.8607 | 0.0852 |
| Wan 2.6 | 0.7280 | 0.1182 | 0.7144 | 0.8032 | 0.8536 | 0.8728 | 0.0992 |
| Embodied world models | |||||||
| GigaWorld-0 | 0.5368 | 0.1552 | 0.6316 | 0.7596 | 0.6156 | 0.8591 | 0.1134 |
| Genie Envisioner | 0.2052 | 0.0679 | 0.8663 | 0.5284 | 0.2028 | 0.8544 | 0.0109 |
| Vidar | 0.5348 | 0.1928 | 0.7872 | 0.7592 | 0.5912 | 0.8826 | 0.0819 |
| IRASim | 0.5656 | 0.3639 | 0.9312 | 0.7788 | 0.6604 | 0.8849 | 0.0526 |
| CtrlWorld | 0.6212 | 0.4766 | 0.9300 | 0.7960 | 0.7272 | 0.8912 | 0.0210 |
| Representation-guided models | |||||||
| TesserAct | 0.5800 | 0.1396 | 0.7159 | 0.7920 | 0.6152 | 0.8783 | 0.0311 |
| RoboMaster | 0.5364 | 0.1158 | 0.8335 | 0.7588 | 0.5772 | 0.8761 | 0.0352 |
| WoW | 0.5564 | 0.2058 | 0.7283 | 0.7672 | 0.5692 | 0.8842 | 0.0434 |
| Wan 2.2-based models | |||||||
| Wan 2.2 | 0.5184 | 0.1627 | 0.7768 | 0.7660 | 0.5376 | 0.8877 | 0.0512 |
| Wan 2.2-AC | 0.6672 | 0.4009 | 0.8657 | 0.8574 | 0.7394 | 0.8837 | 0.0089 |
| Wan 2.2-AC + IMPACT (w/o ADS) | 0.7546 | 0.3485 | 0.8974 | 0.9780 | 0.8472 | 0.8886 | 0.0143 |
| Wan 2.2-AC + IMPACT | 0.7516 | 0.3657 | 0.9026 | 0.9486 | 0.8560 | 0.8881 | 0.0559 |
| Cosmos-based models | |||||||
| Cosmos-Predict 2.5 (text) | 0.3872 | 0.0816 | 0.7051 | 0.7964 | 0.2664 | 0.7733 | 0.1418 |
| Cosmos-Predict 2.5 (action) | 0.5500 | 0.2945 | 0.8862 | 0.7644 | 0.5840 | 0.8879 | 0.0133 |
| + IMPACT | 0.6360 | 0.2003 | 0.6557 | 0.9040 | 0.6360 | 0.8738 | 0.6260 |
C.3 Additional Robot-Arm Cases
name impact.sec:additional_robot_cases xyz In this section, we present more qualitative results of robot-arm manipulation on WorldArena.
Prompt: Pick up the printed sneaker and place it on the blue mat.
Prompt: Pick up the brown-and-white bottle and place it on the blue mat.

Prompt: Pick up the red block and place it on the blue target.

Prompt: Grasp the lidded pot with both grippers and lift it.

Prompt: Pick up the green block and stack it on the red block.

Prompt: Pick up the brown shoe and place it on the blue mat.

Prompt: Pick up one blue bowl and stack it inside the other blue bowl.

Prompt: Press the blue service bell.

Prompt: Place the toy hamburger and French fries in the tray.

Prompt: Pick up the blue elephant toy from beside the black case.
C.4 Additional Human-Hand Cases
name impact.sec:additional_hand_cases xyz In this section, we present more qualitative results of human-hand manipulation on EgoDex.
Prompt: Grasp the transparent container and lift it from the shelf.
Prompt: Grasp the white cylindrical object and lift it from the base.

Prompt: Pick up the two white drawers and stack them together.

Prompt: Grasp and adjust the colorful block structure.

Prompt: Pick up the red book from the table.

Prompt: Stack the orange bowl on the white bowl.