arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.00161v1 [cs.AI] 31 Aug 2026

IMPACT: Attention Is the Interaction Map for
Scalable Interaction-Aware World Model Training

\aaai@corrmultitrueRongze Tang1,2, Jianjie Fang3, Zhaolu Wang3, Ziyou Wang3, Xvyuan Liu3,
Haisheng Su4, Xin Zhang4, Wei Wu4, Chen Gao2,3\corresponding, Yong Li3, Zhibo Chen1,2\corresponding
Abstract

World models have made remarkable progress in action-conditioned future prediction for embodied agents, yet still struggle to model physically plausible interactions. Existing approaches address this limitation by constraining the generation process with external representations encoding motion, geometry, or semantics. Obtaining these spatiotemporally dense representations typically requires auxiliary estimators or manual annotations, limiting training scalability. We instead revisit the training objective and identify a supervision-allocation mismatch under the globally averaged mean squared error (MSE) denoising objective: prevalent static content dominates the optimization signal, leaving sparse dynamic-object regions critical to interaction generation disproportionately under-supervised. Motivated by this observation, we introduce IMPACT, a scalable Interaction-aware Model training framework with Prior-guided Attention Calibration and Targeting. IMPACT uses cross-attention associated with manipulated-object tokens as an internal spatiotemporal prior for action-conditioned changes. It samples candidate regions from this prior, calibrates them with detached local prediction errors to construct an interaction map, and uses the map to reweight denoising supervision, requiring neither external representations nor inference-time modifications. Extensive experiments on robot-arm and human-hand manipulation, spanning diverse control modalities and DiT backbones, show that IMPACT consistently outperforms the corresponding MSE-trained baselines, improving interaction fidelity, physical plausibility, and visual quality.

1University of Science and Technology of China  2Zhongguancun Academy

3Tsinghua University  4Manifold AI

chgao96@gmail.com, chenzhibo@ustc.edu.cnattr /Border [0 0 0] user /Subtype /Link /A << /S /URI /URI (https://embodiedcity.github.io/IMPACT/) >> Project Page   attr /Border [0 0 0] user /Subtype /Link /A << /S /URI /URI (https://github.com/EmbodiedCity/IMPACT.code) >> Code

Refer to caption
Figure 1: Overview of IMPACT. (a) Existing approaches rely on costly and unscalable external dense representations. (b) Standard world-model training with uniformly averaged MSE suffers from a supervision-allocation mismatch, under-supervising sparse interaction regions. (c) IMPACT calibrates object-conditioned cross-attention with prediction errors into an interaction map for efficient and scalable interaction-targeted training.\pdfdestname impact.fig:overview xyz

Introduction

\pdfdest

name impact.sec:introduction xyz

World models have made remarkable progress in predicting future observations (Zhang et al. 2026a; Kim et al. 2026; Fang et al. 2026b) and supporting embodied simulation (Wang et al. 2026a; Gao et al. 2025). Building on large-scale pretrained video generation backbones(Yang et al. 2024; Kong et al. 2024; Wan Team et al. 2025; NVIDIA et al. 2025), they can simulate how an environment evolves under a commanded action, generating action responses in interactive environments (AlayaWorld Team et al. 2026; Hu et al. 2026; Fang et al. 2026a) and robotic manipulation tasks (Zhang et al. 2026a; Bi et al. 2026; Zhang et al. 2026b). Their imagined futures provide scalable experience for long-horizon planning (Wang et al. 2026a; Gao et al. 2025), policy learning (Hu et al. 2025; Zhen et al. 2025; Su et al. 2026), and policy evaluation (Shang et al. 2026; Zhu et al. 2025). In this role, world models should go beyond visual coherence to simulate physically plausible interactions, yet still suffer from object deformation, discontinuous motion, weak action coupling and inconsistent contact (Zhang et al. 2026b).

As shown in attr /Border [0 0 0] goto name impact.fig:overviewFigure 1a, existing approaches typically address these failures by constraining generation with external representations of interaction dynamics. Motion-based methods use optical flow or point trajectories to describe object dynamics (Gao et al. 2025; Zhang et al. 2026b) and geometry-based methods introduce depth, surface normals, reconstructed scenes, or articulated hand meshes (Zhen et al. 2025; Kim et al. 2026). However, obtaining such spatiotemporally dense representations through auxiliary estimators or manual annotations is costly, and the resulting supervision is bounded by the accuracy of these external signals, limiting both training scalability and the achievable interaction quality.

We instead revisit the standard training objective of video world models and identify a supervision-allocation mismatch, as illustrated in attr /Border [0 0 0] goto name impact.fig:overviewFigure 1b. Inherited from general video generation, this globally averaged mean squared error (MSE) denoising objective uniformly weights all spatiotemporal positions (Po et al. 2025; Zhu et al. 2025; Kim et al. 2026). Under uniform weighting, each region contributes to optimization according to its spatial extent rather than its functional importance. Prevalent static content thus dominates the training signal, leaving sparse dynamic-object regions that carry action-conditioned changes disproportionately under-supervised. As a result, models may reduce the global denoising loss and generate visually coherent videos while leaving interaction regions under-optimized.

Our key insight is that world models built on large-scale pretrained video generation backbones already contain a spatiotemporal prior for interaction regions. For manipulation instructions, cross-attention aligns language tokens with spatiotemporal video representations, enabling the attention maps associated with manipulated-object tokens to serve as a spatiotemporal prior for regions likely to undergo action-conditioned changes. The refined prior provides a natural basis for reweighting denoising supervision, allowing sparse interaction regions to contribute to gradient optimization according to their functional importance rather than their spatial extent, thereby improving interaction generation. Because this prior is obtained from the model’s standard forward pass, it requires no external dense representations, scales readily with training data, and leaves inference unchanged.

Based on this insight, we introduce IMPACT, a scalable Interaction-aware Model training framework with Prior-guided Attention Calibration and Targeting. As shown in attr /Border [0 0 0] goto name impact.fig:overviewFigure 1c, the core idea is to treat object-conditioned cross-attention as an interaction prior, turn it into an interaction map, and use this map to reweight denoising supervision toward interaction regions, thereby mitigating the supervision-allocation mismatch and improving interaction generation. IMPACT realizes this within a single training step through two complementary components. In the forward pass, Attention Distribution Sampling (ADS) aggregates the cross-attention of the manipulated-object tokens into a proposal distribution, samples multiple candidate regions, and weights each by its detached local prediction error, calibrating the attention prior with the model’s current prediction difficulty to form a precise interaction map. In the backward pass, Interaction-Weighted Supervision (IWS) uses this map to strengthen denoising supervision on interaction regions while preserving the global objective, and routes gradients so that the interaction-weighted objective updates the non-cross-attention DiT parameters while the attention-producing parameters follow the global objective, preventing the prior from collapsing onto its own signal.

We evaluate IMPACT in two settings: robotic-arm manipulation on WorldArena (Shang et al. 2026) and human-hand manipulation on EgoDex (Hoque et al. 2025). Across both settings, IMPACT surpasses standard uniformly-supervised training (MSE) under the same backbone as well as various baselines, delivering higher interaction fidelity and physical plausibility across diverse embodied scenarios and action conditions.

In summary:

  • •

    We identify a supervision-allocation mismatch in world model training and introduce IMPACT, a scalable training framework that leverages the object-conditioned cross-attention prior to reweight denoising supervision toward interaction regions, without requiring any external dense representations.

  • •

    We realize IMPACT through two complementary designs: ADS evaluates object-conditioned attention proposals using detached local prediction errors and converts these errors into weights to form an interaction map, while IWS uses this map to target denoising supervision, preserve the global objective, and decouple region estimation from regional optimization.

  • •

    We demonstrate across robotic-arm and human-hand manipulation with different control signals and DiT backbones that IMPACT consistently outperforms uniform-MSE training, improving both the physical consistency and visual quality of generated videos.

Related Work

\pdfdest

name impact.sec:related xyz

Interactive video generation and world models

Building on large-scale pretrained video generation backbones, including CogVideoX, HunyuanVideo, Wan, and Cosmos (Yang et al. 2024; Kong et al. 2024; Wan Team et al. 2025; NVIDIA et al. 2025), recent world models inherit rich representations of visual appearance and temporal dynamics, enabling high-fidelity prediction and coherent motion modeling, but these capabilities alone are insufficient to make them useful embodied simulators. To bridge this gap, recent work introduces diverse control conditions that make world models interactive, including natural-language instructions (Xiang et al. 2024; Zhang et al. 2026a; Zhao et al. 2025; Zhao et al. 2026), game and camera controls (Bruce et al. 2024; He et al. 2025), robot action trajectories (NVIDIA et al. 2025; Zhu et al. 2025), and articulated hand or body controls (Wang et al. 2026b; Gao et al. 2026). However, these methods primarily target overall visual quality and controllable generation, while providing limited constraints on interaction dynamics, resulting in physically implausible behavior within interaction regions.

External representations priors

To address this limitation, recent methods introduce spatiotemporally dense representations such as optical flow, depth maps, or reconstructed 3D structure to explicitly constrain the generation process. Motion-based approaches exploit optical flow, point trajectories, or latent temporal discrepancies to emphasize dynamic regions (Fang et al. 2025; Zhang et al. 2026b; Wu et al. 2026). Geometry-based methods introduce depth, cross-view 3D structure, or projected robot kinematics as auxiliary targets or structured conditions (Tian et al. 2026; Yang et al. 2026; Liu et al. 2025). However, constructing these representations often requires external motion, depth, segmentation, and video-understanding models, sometimes together with manually verified annotations (Fang et al. 2025; Zhang et al. 2026b; Yan et al. 2025; Luo et al. 2026). These preprocessing costs grow with the size and duration of the training corpus, thereby limiting training scalability.

Method

\pdfdest

name impact.sec:method xyz

attr /Border [0 0 0] goto name impact.fig:frameworkFigure 2 presents an overview of IMPACT. Given a training sample, IMPACT identifies the manipulated-object tokens and uses their cross-attention as a spatial prior. Attention Distribution Sampling(ADS) samples candidate regions from this prior and weights them by detached local prediction errors to construct an interaction map, which Interaction-Weighted Supervision(IWS) uses to target denoising supervision toward interaction-relevant tokens. Gradient routing optimizes the cross-attention parameters with the original global objective and the remaining DiT parameters with the interaction-weighted objective, while all additional operations are training-only and incur no inference-time overhead.

Refer to caption
Figure 2: Framework of IMPACT. Object-token grounding (left): a frozen Qwen2.5-0.5B extracts the manipulated-object phrase from the instruction and aligns it to the corresponding object tokens. Training pipeline (right): within the DiT blocks, the cross-attention of these object tokens forms a proposal distribution (cross-attention map), from which 𝐌1,…,𝐌8\mathbf{M}_{1},\dots,\mathbf{M}_{8} candidate regions are sampled and calibrated by their detached local prediction errors into an interaction map. The map targets denoising supervision toward interaction regions, and gradient routing optimizes the cross-attention parameters with the global objective ℒMSE\mathcal{L}_{\mathrm{MSE}} and the remaining DiT parameters with the interaction-weighted objective ℒIWS\mathcal{L}_{\mathrm{IWS}}. \pdfdestname impact.fig:framework xyz

Preliminaries

\pdfdest

name impact.sec:preliminaries xyz

We build on a latent video diffusion transformer trained with flow matching (Lipman et al. 2023). Given a video 𝐱0\mathbf{x}_{0}, a frozen variational autoencoder maps it into a latent representation 𝐳0∈ℝC×T×H×W\mathbf{z}_{0}\in\mathbb{R}^{C\times T\times H\times W}, where TT, HH, and WW denote the temporal and spatial dimensions. The conditioning information is denoted by 𝐜=(𝐲,𝐱ref,𝐚)\mathbf{c}=(\mathbf{y},\mathbf{x}^{\mathrm{ref}},\mathbf{a}), where 𝐲\mathbf{y} is the language instruction, 𝐱ref\mathbf{x}^{\mathrm{ref}} is the reference observation, and 𝐚\mathbf{a} represents the corresponding control signal, such as hand poses or robot trajectories.

For a sampled noise level σt∈[0,1]\sigma_{t}\in[0,1] and Gaussian noise ϵ∼𝒩⁡(𝟎,𝐈)\boldsymbol{\epsilon}\sim\mathcal{N}(\mathbf{0},\mathbf{I}), the clean latent is interpolated with noise as

𝐳t=(1−σt)​𝐳0+σt​ϵ.\pdfdest​n​a​m​e​i​m​p​a​c​t.e​q:f​l​o​wi​n​t​e​r​p​o​l​a​t​i​o​n​x​y​z\mathbf{z}_{t}=(1-\sigma_{t})\mathbf{z}_{0}+\sigma_{t}\boldsymbol{\epsilon}.\pdfdest name{impact.eq:flow_{i}nterpolation}xyz (1)

The diffusion transformer vθv_{\theta} takes (𝐳t,t,𝐜)(\mathbf{z}_{t},t,\mathbf{c}) as input and predicts the velocity target ϵ−𝐳0\boldsymbol{\epsilon}-\mathbf{z}_{0}. Let

Ω={1,…,T}×{1,…,H}×{1,…,W}\Omega=\{1,\ldots,T\}\times\{1,\ldots,H\}\times\{1,\ldots,W\} (2)

denote the set of spatiotemporal latent positions. The prediction error at position p∈Ωp\in\Omega is

ℓp=1C​‖vθ​(𝐳t,t,𝐜)p−(ϵ−𝐳0)p‖22.\pdfdest​n​a​m​e​i​m​p​a​c​t.e​q:l​o​c​a​le​r​r​o​r​x​y​z\ell_{p}=\frac{1}{C}\left\|v_{\theta}(\mathbf{z}_{t},t,\mathbf{c})_{p}-(\boldsymbol{\epsilon}-\mathbf{z}_{0})_{p}\right\|_{2}^{2}.\pdfdest name{impact.eq:local_{e}rror}xyz (3)

Standard flow-matching training uniformly averages this error over all spatiotemporal positions:

ℒglobal​(θ)=𝔼𝐳0,ϵ,t​[1|Ω|​∑p∈Ωℓp].\pdfdest​n​a​m​e​i​m​p​a​c​t.e​q:g​l​o​b​a​lo​b​j​e​c​t​i​v​e​x​y​z\mathcal{L}_{\mathrm{global}}(\theta)=\mathbb{E}_{\mathbf{z}_{0},\boldsymbol{\epsilon},t}\left[\frac{1}{|\Omega|}\sum_{p\in\Omega}\ell_{p}\right].\pdfdest name{impact.eq:global_{o}bjective}xyz (4)

Consequently, the contribution of a region to optimization is largely determined by its spatial extent. Small interaction regions therefore receive no additional emphasis despite containing rapid motion, contact transitions, and action-conditioned state changes. IMPACT addresses this supervision-allocation mismatch without changing the underlying flow-matching formulation.

Object-Token Grounding

\pdfdest

name impact.sec:object_grounding xyz

The interaction region depends on the object being manipulated. We therefore begin by grounding the manipulated object in the language instruction. For an instruction 𝐲\mathbf{y}, a frozen Qwen2.5-0.5B model (Qwen et al. 2025) extracts the noun phrase that denotes the manipulated object. For example, given the instruction “Right gripper stacks blue bowl on white tabletop,” the extracted phrase is “blue bowl.”

Let

𝐞=(𝐞1,…,𝐞N)\mathbf{e}=(\mathbf{e}_{1},\ldots,\mathbf{e}_{N}) (5)

denote the sequence of text embeddings produced by the tokenizer and text encoder. We align the extracted object phrase with this token sequence and denote its token positions by

𝒪⊆{1,…,N}.\pdfdest​n​a​m​e​i​m​p​a​c​t.e​q:o​b​j​e​c​tt​o​k​e​ns​e​t​x​y​z\mathcal{O}\subseteq\{1,\ldots,N\}.\pdfdest name{impact.eq:object_{t}oken_{s}et}xyz (6)

When an object phrase is divided into multiple subword tokens, all matched positions are included in 𝒪\mathcal{O}. This object-token set provides the semantic anchor used by ADS to extract an object-conditioned spatial distribution. The grounding stage operates only on the instruction and therefore requires neither visual annotations nor external spatial estimators.

Forward: Attention Distribution Sampling

\pdfdest

name impact.sec:ads xyz

ADS converts the semantic object grounding into an interaction map through three operations: object-conditioned attention aggregation, candidate-region sampling, and prediction-error-based weighting.

Object-conditioned attention distribution.

Within each DiT cross-attention layer, visual queries attend to the text-token sequence. For block ll, attention head hh, visual position pp, and text-token position jj, the cross-attention probability is

Ap,j(l,h)=softmaxj⁡(𝐪p(l,h)​𝐤j(l,h)⊤d),\pdfdest​n​a​m​e​i​m​p​a​c​t.e​q:c​r​o​s​sa​t​t​e​n​t​i​o​n​x​y​zA_{p,j}^{(l,h)}=\operatorname{softmax}_{j}\left(\frac{\mathbf{q}_{p}^{(l,h)}\mathbf{k}_{j}^{(l,h)\top}}{\sqrt{d}}\right),\pdfdest name{impact.eq:cross_{a}ttention}xyz (7)

where dd is the key dimension. We sum the probability mass assigned to the object-token set 𝒪\mathcal{O} and average it over the selected blocks ℬ\mathcal{B} and attention heads ℋ\mathcal{H}:

Ap=1|ℬ|​|ℋ|​∑l∈ℬ∑h∈ℋ∑j∈𝒪Ap,j(l,h).\pdfdest​n​a​m​e​i​m​p​a​c​t.e​q:o​b​j​e​c​ta​t​t​e​n​t​i​o​n​x​y​zA_{p}=\frac{1}{|\mathcal{B}|\,|\mathcal{H}|}\sum_{l\in\mathcal{B}}\sum_{h\in\mathcal{H}}\sum_{j\in\mathcal{O}}A_{p,j}^{(l,h)}.\pdfdest name{impact.eq:object_{a}ttention}xyz (8)

The resulting map 𝐀∈[0,1]T×H×W\mathbf{A}\in[0,1]^{T\times H\times W} forms an object-conditioned proposal distribution over the video latent. Because it is obtained from the standard conditional forward pass, constructing 𝐀\mathbf{A} requires no additional visual model or spatial annotation.

Candidate-region sampling.

Instead of using a single deterministic region, ADS samples KK spatially coherent candidates around the attention distribution. We first detach and temper the attention map:

𝐀~=clip⁡(sg⁡[𝐀]κ,ε,1−ε),\pdfdest​n​a​m​e​i​m​p​a​c​t.e​q:a​t​t​e​n​t​i​o​nt​e​m​p​e​r​i​n​g​x​y​z\widetilde{\mathbf{A}}=\operatorname{clip}\left(\operatorname{sg}[\mathbf{A}]^{\,\kappa},\varepsilon,1-\varepsilon\right),\pdfdest name{impact.eq:attention_{t}empering}xyz (9)

where sg⁡[⋅]\operatorname{sg}[\cdot] denotes stop-gradient and κ\kappa controls the concentration of the proposal distribution. We then transform it into logits:

𝐙=log⁡𝐀~1−𝐀~.\pdfdest​n​a​m​e​i​m​p​a​c​t.e​q:a​t​t​e​n​t​i​o​nl​o​g​i​t​s​x​y​z\mathbf{Z}=\log\frac{\widetilde{\mathbf{A}}}{1-\widetilde{\mathbf{A}}}.\pdfdest name{impact.eq:attention_{l}ogits}xyz (10)

For each candidate k∈{1,…,K}k\in\{1,\ldots,K\}, we sample Gaussian noise on a coarse spatiotemporal grid and trilinearly interpolate it to the resolution of 𝐀\mathbf{A}, producing a smooth perturbation field 𝜼k\boldsymbol{\eta}_{k}. A soft candidate is generated as

𝐒k=sigmoid⁡(𝐙+σ​𝜼kτ),\pdfdest​n​a​m​e​i​m​p​a​c​t.e​q:s​o​f​tc​a​n​d​i​d​a​t​e​x​y​z\mathbf{S}_{k}=\operatorname{sigmoid}\left(\frac{\mathbf{Z}+\sigma\boldsymbol{\eta}_{k}}{\tau}\right),\pdfdest name{impact.eq:soft_{c}andidate}xyz (11)

where σ\sigma controls the perturbation magnitude and τ\tau is the sampling temperature. Sampling noise on a coarse grid produces coherent spatial variations rather than independent token-wise perturbations.

To prevent candidate scores from being dominated by differences in region size, every candidate is converted into a fixed-area binary region. Specifically, we retain the largest ρ​|Ω|\rho|\Omega| values of 𝐒k\mathbf{S}_{k}:

Mk,p=𝕀[Sk,p≥TopKThreshold(𝐒k,ρ|Ω|)],\pdfdestnameimpact.eq:fixedareacandidatexyzM_{k,p}=\mathbb{I}\left[S_{k,p}\geq\operatorname{TopKThreshold}(\mathbf{S}_{k},\rho|\Omega|)\right],\pdfdest name{impact.eq:fixed_{a}rea_{c}andidate}xyz (12)

yielding 𝐌k∈{0,1}T×H×W\mathbf{M}_{k}\in\{0,1\}^{T\times H\times W}. All candidates consequently cover the same number of latent positions and remain directly comparable.

Prediction-error-based weighting.

ADS evaluates every candidate using the mean local prediction error within that region:

dk=∑p∈ΩMk,p​sg⁡[ℓp]∑p∈ΩMk,p+ε.\pdfdest​n​a​m​e​i​m​p​a​c​t.e​q:c​a​n​d​i​d​a​t​es​c​o​r​e​x​y​zd_{k}=\frac{\sum_{p\in\Omega}M_{k,p}\operatorname{sg}[\ell_{p}]}{\sum_{p\in\Omega}M_{k,p}+\varepsilon}.\pdfdest name{impact.eq:candidate_{s}core}xyz (13)

A higher value of dkd_{k} indicates that the candidate covers content that is currently more difficult for the model to predict. The candidate errors are converted into normalized weights:

αk=exp⁡(β​dk)∑j=1Kexp⁡(β​dj),\pdfdest​n​a​m​e​i​m​p​a​c​t.e​q:c​a​n​d​i​d​a​t​ew​e​i​g​h​t​x​y​z\alpha_{k}=\frac{\exp(\beta d_{k})}{\sum_{j=1}^{K}\exp(\beta d_{j})},\pdfdest name{impact.eq:candidate_{w}eight}xyz (14)

where β\beta controls the concentration of the weighting distribution. The candidate regions are then weighted to form the interaction map

𝐆=clip⁡(∑k=1Kαk​𝐌k,0,1).\pdfdest​n​a​m​e​i​m​p​a​c​t.e​q:i​n​t​e​r​a​c​t​i​o​nm​a​p​x​y​z\mathbf{G}=\operatorname{clip}\left(\sum_{k=1}^{K}\alpha_{k}\mathbf{M}_{k},0,1\right).\pdfdest name{impact.eq:interaction_{m}ap}xyz (15)

Through this process, the object-conditioned attention determines where candidate regions are sampled, while the current local prediction error determines their relative contribution to 𝐆\mathbf{G}. The entire ADS pipeline is computed without gradient tracking, preventing the model from directly modifying the proposal distribution or candidate weights to reduce the regional objective.

In our implementation, we use K=8K=8, an attention tempering exponent κ=0.65\kappa=0.65, perturbation scale σ=0.75\sigma=0.75, sampling temperature τ=1\tau=1, and a coarse noise grid of size 5×8×125\times 8\times 12. Each candidate retains ρ=0.10\rho=0.10 of the spatiotemporal latent positions, and candidate errors are converted into weights using β=0.7\beta=0.7.

Backward: Interaction-Weighted Supervision

\pdfdest

name impact.sec:iws xyz

IWS uses the interaction map 𝐆\mathbf{G} to increase the contribution of interaction-relevant positions during denoising optimization. For each position p∈Ωp\in\Omega, we define

wp=1+(γ−1)​sg⁡[Gp],\pdfdest​n​a​m​e​i​m​p​a​c​t.e​q:i​n​t​e​r​a​c​t​i​o​nw​e​i​g​h​t​x​y​zw_{p}=1+(\gamma-1)\operatorname{sg}[G_{p}],\pdfdest name{impact.eq:interaction_{w}eight}xyz (16)

where γ≥1\gamma\geq 1 controls the maximum regional emphasis. Because Gp∈[0,1]G_{p}\in[0,1], the resulting weight satisfies wp∈[1,γ]w_{p}\in[1,\gamma]. Positions outside the estimated interaction region retain their original unit weight, while positions with high interaction-map values receive stronger supervision.

The interaction-weighted denoising objective is

ℒIWS​(θ)=𝔼𝐳0,ϵ,t​[∑p∈Ωwp​ℓp∑p∈Ωwp].\pdfdest​n​a​m​e​i​m​p​a​c​t.e​q:i​w​so​b​j​e​c​t​i​v​e​x​y​z\mathcal{L}_{\mathrm{IWS}}(\theta)=\mathbb{E}_{\mathbf{z}_{0},\boldsymbol{\epsilon},t}\left[\frac{\sum_{p\in\Omega}w_{p}\ell_{p}}{\sum_{p\in\Omega}w_{p}}\right].\pdfdest name{impact.eq:iws_{o}bjective}xyz (17)

The normalization by the total weight prevents the loss scale from growing with the size or magnitude of the emphasized region. Unlike a masked regional loss, attr /Border [0 0 0] goto name impact.eq:iws_objectiveEquation 17 retains supervision over the complete spatiotemporal field and only changes its spatial allocation. We use γ=5\gamma=5 in the main experiments.

Gradient-decoupled optimization.

The interaction map is derived from cross-attention, creating a dependency between region estimation and the objective guided by that region. Although 𝐆\mathbf{G} is detached, allowing ℒIWS\mathcal{L}_{\mathrm{IWS}} to update the attention-producing parameters could still cause the cross-attention representation itself to adapt to the regional objective in subsequent training steps. We therefore separate the DiT parameters into the cross-attention parameters θca\theta_{\mathrm{ca}} and all remaining parameters θrest\theta_{\mathrm{rest}}. Their gradients are routed according to

∇θcaℒI​M​P​A​C​T=∇θcaℒglobal,\pdfdest​n​a​m​e​i​m​p​a​c​t.e​q:c​r​o​s​sa​t​t​e​n​t​i​o​nr​o​u​t​i​n​g​x​y​z\nabla_{\theta_{\mathrm{ca}}}\mathcal{L}_{IMPACT}=\nabla_{\theta_{\mathrm{ca}}}\mathcal{L}_{\mathrm{global}},\pdfdest name{impact.eq:cross_{a}ttention_{r}outing}xyz (18)

and

∇θrestℒI​M​P​A​C​T=∇θrestℒIWS.\pdfdest​n​a​m​e​i​m​p​a​c​t.e​q:r​e​m​a​i​n​i​n​gp​a​r​a​m​e​t​e​rr​o​u​t​i​n​g​x​y​z\nabla_{\theta_{\mathrm{rest}}}\mathcal{L}_{IMPACT}=\nabla_{\theta_{\mathrm{rest}}}\mathcal{L}_{\mathrm{IWS}}.\pdfdest name{impact.eq:remaining_{p}arameter_{r}outing}xyz (19)

In practice, this routing is implemented through two backward passes with parameter-group gradient hooks. The global backward pass retains gradients only for θca\theta_{\mathrm{ca}}, while the IWS backward pass retains gradients only for θrest\theta_{\mathrm{rest}}. Consequently, the cross-attention used to estimate interaction regions remains governed by the original uniformly supervised objective, whereas the remaining DiT parameters learn from the spatially reallocated supervision. This separation prevents the regional objective from directly optimizing its own localization signal and decouples interaction-region estimation from interaction-focused optimization.

Training and inference cost.

IMPACT reuses the cross-attention probabilities and token-wise prediction errors already produced during standard world-model training. Its additional computation consists primarily of sampling KK candidate masks, evaluating their masked mean errors, and constructing the interaction-weighted objective. No component of ADS or IWS is used at inference time, so the trained world model preserves the original architecture and inference procedure.

Experiments

\pdfdest

name impact.sec:experiments xyz

Model EWMScore Visual Quality Motion Quality Content Consist. Physics Adher. 3D Accuracy Control- lability
General world models
CogVideoX 57.90 55.81 42.49 68.12 47.33 84.62 54.40
Veo 3.1 58.87 56.44 46.12 66.12 45.52 78.48 62.62
Wan 2.6 61.86 61.62 68.31 60.36 42.31 75.88 60.85
Embodied world models
GigaWorld-0 53.39 44.82 58.79 58.74 34.60 69.56 52.94
Genie Envisioner 43.65 29.78 49.17 62.63 13.66 69.73 35.60
Vidar 51.60 46.07 40.55 60.93 36.38 77.32 51.86
IRASim 58.12 54.81 44.25 69.67 46.48 85.50 53.26
CtrlWorld 59.70 55.33 50.28 63.99 54.89 86.30 54.65
Representation-guided models
TesserAct 53.23 41.64 50.59 66.60 35.98 75.39 50.82
RoboMaster 51.84 34.32 48.49 69.25 32.61 79.62 49.62
WoW 54.88 52.98 50.02 64.52 38.11 74.77 49.89
Wan 2.2-based models
Wan 2.2 50.79 51.41 42.12 54.02 34.05 77.14 49.22
Wan 2.2-AC 58.65 56.57 48.23 60.55 53.41 86.16 54.40
Wan 2.2-AC + IMPACT 62.46 60.60 53.28 60.29 55.87 92.56 60.00
Cosmos-based models
Cosmos-Predict 2.5 (text) 50.81 47.65 60.32 57.94 23.44 75.08 39.38
Cosmos-Predict 2.5 (action) 55.91 57.87 38.89 68.73 42.23 82.53 49.51
Cosmos-Predict 2.5 (action) + IMPACT 62.53 56.32 68.22 57.90 41.82 77.99 71.19
Table 1: Generation quality on WorldArena. Models are grouped into general, embodied, and representation-guided models, with Wan 2.2-based and Cosmos-based variants listed separately. We report the overall EWMScore and six aggregate dimensions, each on a [0,100][0,100] scale where higher is better (↑\uparrow). Boldface and underlining denote the best and second-best results in each column.\pdfdestname impact.tab:main xyz
Visual Hand interaction
Model FVD ↓\downarrow FID ↓\downarrow CLIP- Hand ↑\uparrow Hand IoU ↑\uparrow
General world models
HunyuanVideo-1.5 541.83 56.73 0.902 0.328
Cosmos-Predict 2.5 615.42 50.12 0.914 0.386
Pose control
MimicMotion 612.75 48.55 0.882 0.492
MagicPose 1456.20 212.94 0.864 0.298
VACE 358.42 50.65 0.895 0.493
LOME 1748.29 66.03 0.745 0.087
Wan 2.2-based models
Wan 2.2 1463.05 199.35 0.876 0.557
Wan 2.2-AC 366.12 44.71 0.921 0.693
Wan 2.2-AC + IMPACT 110.94 5.79 0.952 0.772
Table 2: Generation quality on EgoDex. Boldface and underlining denote the best and second-best results in each column.\pdfdestname impact.tab:egodex xyz

Setup

\pdfdest

name impact.sec:setup xyz

Implementation details.

We evaluate IMPACT in two manipulation settings to demonstrate the generality of our method across interaction types, applying it to the Wan2.2 and Cosmos-Predict 2.5 backbone in both. For robot-arm manipulation, the model is trained on ∼\sim350K 17-frame videos collected from RoboTwin (Chen et al. 2025), conditioned on 14-DoF dual-arm action trajectories, injected through an additional action encoder. For human-hand manipulation, we train on ∼\sim256K clips drawn from EgoDex (Hoque et al. 2025) across 118 tasks, using 81-frame videos conditioned on hand-pose videos temporally aligned with the RGB frames, injected through the shared VAE encoder without any additional encoder. Both settings are trained at 720p and optimization is identical across the two settings: we use bf16 mixed precision with FSDP over 8 GPUs, a per-device batch size of 1 with 4-step gradient accumulation (effective global batch size 1×8×4=321\times 8\times 4=32), a constant learning rate of 2×10−52\times 10^{-5} after 100 warm-up steps, and we train for one epoch. IMPACT hyperparameters are fixed across both settings: K=8K=8, σ=0.75\sigma=0.75, coarse grid 5×8×125\times 8\times 12, ρ=0.10\rho=0.10, κ=0.65\kappa=0.65, β=0.7\beta=0.7, τ=1\tau=1, γ=5\gamma=5.

Benchmarks and metrics.

For robot-arm manipulation, we evaluate on WorldArena (Shang et al. 2026), which scores dual-arm manipulation along six dimensions: Visual Quality, Motion Quality, Content Consistency, Physics Adherence, 3D Accuracy, and Controllability, spanning 16 normalized metrics, and condenses overall generation quality into a single EWMScore (the mean of the 16 metrics). We report EWMScore together with the six aggregate dimensions in attr /Border [0 0 0] goto name impact.tab:mainTable 1, and provide all 16 metrics in the technical appendix. For human-hand manipulation, we evaluate on the EgoDex (Hoque et al. 2025) test set along two axes (attr /Border [0 0 0] goto name impact.tab:egodexTable 2): visual metrics: FVD (Unterthiner et al. 2018) and FID (Heusel et al. 2017), covering temporal coherence and per-frame appearance quality; and hand-interaction metrics: CLIP-Hand for the local appearance and semantics of the hand and nearby manipulated object, and Hand IoU for the coarse 2D position and scale of the generated hand (Sun et al. 2026).

Baselines.

For robot-arm manipulation, all models follow the WorldArena evaluation settings. Existing baselines comprise general world models: CogVideoX (Yang et al. 2024), Wan 2.6 (Wan Team et al. 2025), and Veo 3.1; embodied world models: GigaWorld-0, Genie Envisioner, Vidar, IRASim (Zhu et al. 2025), and CtrlWorld; and representation-guided models: TesserAct (Zhen et al. 2025), RoboMaster, and WoW. We additionally report two backbone-specific groups to evaluate IMPACT: Wan 2.2-based models include Wan 2.2 (Wan Team et al. 2025), MSE-trained Wan 2.2-AC, and Wan 2.2-AC with IMPACT; Cosmos-based models include Cosmos-Predict 2.5 (text) (NVIDIA et al. 2025), Cosmos-Predict 2.5 (action), and its IMPACT variant. For human-hand manipulation, baselines comprise general video world models: HunyuanVideo-1.5 (Wu et al. 2025) and Cosmos-Predict 2.5 (NVIDIA et al. 2025) and pose-controlled models: MimicMotion (Zhang et al. 2025), MagicPose (Chang et al. 2024), VACE (Jiang et al. 2025), and LOME (Gao et al. 2026).

Refer to caption
Figure 3: Qualitative comparison on both interaction settings. Left: robot-arm manipulation on WorldArena; right: first-person human-hand manipulation on EgoDex. For each setting we show generated frames for a representative instruction against various baselines. In both settings, IMPACT follows the instruction more faithfully and renders the interaction. Red boxes highlight artifacts and inconsistencies in the interaction regions of the baseline generations.\pdfdestname impact.fig:qualitative xyz
Refer to caption
Figure 4: Visualization of ADS. On one robot-arm training forward pass, the light-to-dark color bar encodes increasing candidate errors and thus the relative contribution of each candidate to the aggregate.\pdfdestname impact.fig:ads_calibration xyz

Quantitative Analysis

\pdfdest

name impact.sec:quantitative_analysis xyz

Robot-arm manipulation.

attr /Border [0 0 0] goto name impact.tab:mainTable 1 shows that IMPACT improves both backbone families. On Cosmos-Predict 2.5 (action), it raises EWMScore from 55.91 to 62.53 (+6.62 points, 11.8%), achieving the best overall result, the best Controllability (71.19), and the second-best Motion Quality (68.22). Interaction Quality increases from 0.5500 to 0.6360 and Action Following from 0.0133 to 0.6260, although the remaining aggregate dimensions decline. On Wan 2.2-AC, IMPACT raises EWMScore from 58.65 to 62.46 (+3.81 points, 6.5%) and improves Visual Quality, Motion Quality, Physics Adherence, 3D Accuracy, and Controllability, attaining the best Physics Adherence (55.87) and 3D Accuracy (92.56) and the second-best Visual Quality (60.60). The best IMPACT result exceeds Wan 2.6, CtrlWorld, and WoW by 0.67, 2.83, and 7.65 points, respectively.

Human-hand manipulation.

For human-hand manipulation, attr /Border [0 0 0] goto name impact.tab:egodexTable 2 reports that IMPACT attains the best results on visual metrics across both general video world models and pose-controlled models, cutting FVD from 366.12 to 110.94 and FID from 44.71 to 5.79 over the action-conditioned Wan 2.2-AC on the same backbone. It also leads on both hand-interaction metrics, improving CLIP-Hand from 0.921 to 0.952 and Hand IoU from 0.693 to 0.772 over Wan 2.2-AC on the same backbone. These gains indicate stronger local interaction fidelity and hand localization than both the MSE-trained counterpart and the pose-controlled baselines. Together with the robot-arm results, these improvements demonstrate that IMPACT delivers consistent gains across different DiT backbones and control signals.

Qualitative analysis

\pdfdest

name impact.sec:Qualitative analysis xyz

Generation comparison.

attr /Border [0 0 0] goto name impact.fig:qualitativeFigure 4 compares generations in robot-arm and human-hand manipulation. In both cases, the baselines follow the instruction loosely and tend to blur or distort the contact region, whether the gripper–object contact for the robot arm or the hand–object contact for the human hand, and some drift away from the target. In contrast, IMPACT produces the specified interaction with a sharper contact region and more coherent object dynamics, while keeping the surrounding scene stable. This clearly demonstrates the effectiveness of IMPACT for interaction-region generation.

ADS calibration.

attr /Border [0 0 0] goto name impact.fig:ads_calibrationFigure 4 illustrates how ADS calibrates the attention prior in a representative training example. The raw object-conditioned cross-attention map 𝐀\mathbf{A} is diffuse, spreading across the robot arm, workspace, and background. ADS samples K=8K{=}8 candidate regions 𝐌1,…,𝐌8\mathbf{M}_{1},\dots,\mathbf{M}_{8} from this prior and evaluates them using detached local prediction errors (attr /Border [0 0 0] goto name impact.eq:candidate_scoreEq. 13). Candidates covering the contact region receive higher weights than those dominated by the static background. Their weighted aggregation (attr /Border [0 0 0] goto name impact.eq:interaction_mapEq. 15) produces an interaction map 𝐆\mathbf{G} that is more concentrated around the interacting arm, gripper, and manipulated object. This example illustrates how ADS refines a coarse attention prior into a more targeted supervision map for IWS.

Ablation Studies

\pdfdest

name impact.sec:ablation xyz

Component ablation: IWS and ADS.

attr /Border [0 0 0] goto name impact.tab:ads_samplingTable 3 shows the effect of separating the components on WorldArena and reveals their complementarity. Starting from the AC backbone (58.65 EWMScore), IWS provides the more direct gain (+2.89 to 61.54), since it primarily addresses the supervision-allocation mismatch. On top of this, ADS further calibrates the cross-attention prior and delivers an additional improvement (+0.92 to 62.46) over weighting the raw, coarse prior directly. Integrated within a single training step, the two components jointly strengthen interaction-region generation, improving EWMScore by 3.81 points overall.

Method EWMScore Visual Quality Motion Quality Physics Adher.
Wan 2.2-AC 58.65 56.57 48.23 53.41
+ IWS 61.54 60.44 49.16 55.16
+ IWS + ADS (IMPACT) 62.46 60.60 53.28 55.87
Table 3: Per-component ablation on WorldArena.\pdfdestname impact.tab:ads_sampling xyz

Conclusion

\pdfdest

name impact.sec:conclusion xyz

In this work, we introduced IMPACT, a scalable framework that addresses the supervision-allocation mismatch by converting object-conditioned cross-attention into targeted denoising supervision. ADS calibrates attention proposals with detached local prediction errors, while IWS reweights training with the resulting interaction map, requiring neither external spatial signals nor inference-time changes. Experiments on robot-arm and human-hand manipulation show consistent gains over uniform MSE training and strong baselines, demonstrating the effectiveness and scalability of IMPACT for interaction-aware world model training.

References

  • AlayaWorld Team et al. (2026) AlayaWorld Team, K. Zhang, C. Li, Y. Zhan, Y. Ge, Y. Yin, J. Tan, K. He, L. Fan, R. Liu, X. Xu, X. Chu, Z. Li, Z. Lin, Z. Wang, Z. Meng, and Z. Gao AlayaWorld: long-horizon and playable video world generation. External Links: 2607.06291, Link Cited by: Introduction.
  • Bi et al. (2026) H. Bi, H. Tan, S. Xie, Z. Wang, S. Huang, H. Liu, R. Zhao, Y. Feng, C. Xiang, Y. Rong, H. Zhao, H. Liu, Z. Su, L. Ma, H. Su, and J. Zhu Motus: a unified latent action world model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 35101–35113. External Links: Link Cited by: Introduction.
  • Bruce et al. (2024) J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder, Y. Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, Y. Aytar, S. Bechtle, F. Behbahani, S. C. Y. Chan, N. Heess, L. Gonzalez, S. Osindero, S. Ozair, S. Reed, J. Zhang, K. Zolna, J. Clune, N. d. Freitas, T. Rocktaschel, and D. Hafner Genie: generative interactive environments. arXiv preprint arXiv:2402.15391. Cited by: Interactive video generation and world models.
  • Chang et al. (2024) D. Chang, Y. Shi, Q. Gao, H. Xu, J. Fu, G. Song, Q. Yan, Y. Zhu, X. Yang, and M. Soleymani MagicPose: realistic human poses and facial expressions retargeting with identity-aware diffusion. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, pp. 6263–6285. External Links: Link Cited by: Baselines..
  • Chen et al. (2025) T. Chen, Z. Chen, B. Chen, Z. Cai, Y. Liu, Z. Li, Q. Liang, X. Lin, Y. Ge, Z. Gu, W. Deng, Y. Guo, T. Nian, X. Xie, Q. Chen, K. Su, T. Xu, G. Liu, M. Hu, H. Gao, K. Wang, Z. Liang, Y. Qin, X. Yang, P. Luo, and Y. Mu RoboTwin 2.0: a scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation. arXiv preprint arXiv:2506.18088. Cited by: Implementation details..
  • Fang et al. (2026a) J. Fang, Y. Lei, Q. Wan, Z. Wang, Y. Huang, Y. Xu, B. Zhao, W. Zhang, C. Gao, X. Chen, and Y. Li iWorld-Bench: a benchmark for interactive world models with a unified action generation framework. External Links: 2605.03941, Link Cited by: Introduction.
  • Fang et al. (2026b) J. Fang, Y. Xu, Z. Wang, C. Gao, Y. Huang, Z. Wang, R. Tang, M. Jia, B. Zhao, W. Zhang, X. Zhang, H. Su, Y. Shang, W. Wu, X. Chen, and Y. Li Worldscape-MoE: a unified mixture-of-experts world model for scalable heterogeneous action control. External Links: 2607.03964, Link Cited by: Introduction.
  • Fang et al. (2025) Y. Fang, K. Ranasinghe, L. Xue, H. Zhou, J. Tan, R. Xu, S. Heinecke, C. Xiong, S. Savarese, D. Szafir, M. Ding, M. S. Ryoo, and J. C. Niebles Robotic VLA benefits from joint learning with motion image diffusion. arXiv preprint arXiv:2512.18007. Cited by: External representations priors.
  • Gao et al. (2025) C. Gao, H. Zhang, Z. Xu, Z. Cai, and L. Shao FLIP: flow-centric generative planning as general-purpose manipulation world model. In International Conference on Learning Representations, External Links: Link Cited by: Introduction, Introduction.
  • Gao et al. (2026) Q. Gao, J. Yang, Q. Xu, L. Chen, and Y. Wang LOME: learning human-object manipulation with action-conditioned egocentric world model. External Links: 2603.27449, Link Cited by: Interactive video generation and world models, Baselines..
  • He et al. (2025) X. He, C. Peng, Z. Liu, B. Wang, Y. Zhang, Q. Cui, F. Kang, B. Jiang, M. An, Y. Ren, B. Xu, H. Guo, K. Gong, C. Wu, W. Li, X. Song, Y. Liu, E. Li, and Y. Zhou Matrix-Game 2.0: an open-source, real-time, and streaming interactive world model. External Links: 2508.13009, Link Cited by: Interactive video generation and world models.
  • Heusel et al. (2017) M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter GANs trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems, Cited by: Benchmarks and metrics..
  • Hoque et al. (2025) R. Hoque, P. Huang, D. J. Yoon, M. Sivapurapu, and J. Zhang EgoDex: learning dexterous manipulation from large-scale egocentric video. External Links: 2505.11709, Link Cited by: Introduction, Implementation details., Benchmarks and metrics..
  • Hu et al. (2026) A. Hu, V. Volhejn, A. R. Rahary, C. Mulder, A. Makkar, A. Royer, M. Orsini, A. Liao, A. Jelley, E. Alonso, F. Laurent, F. Norén, J. Swingos, J. Hünermann, K. Rollins, L. Hosseini, M. Le Cauchois, M. Peter, P. de Witte, T. Brown, V. Micheli, M. Böhle, G. de Marmiesse, V. Sharmanska, L. Specia, M. Black, and P. Pérez Multiplayer interactive world models with representation autoencoders. External Links: 2607.05352, Link Cited by: Introduction.
  • Hu et al. (2025) Y. Hu, Y. Guo, P. Wang, X. Chen, Y. Wang, J. Zhang, K. Sreenath, C. Lu, and J. Chen Video prediction policy: a generalist robot policy with predictive visual representations. In Proceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 267, pp. 24328–24346. External Links: Link Cited by: Introduction.
  • Jiang et al. (2025) Z. Jiang, Z. Han, C. Mao, J. Zhang, Y. Pan, and Y. Liu VACE: all-in-one video creation and editing. arXiv preprint arXiv:2503.07598. Cited by: Baselines..
  • Kim et al. (2026) B. Kim, T. Kim, J. Lee, and H. Joo Dexterous world models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 29663–29673. External Links: Link Cited by: Introduction, Introduction, Introduction.
  • Kong et al. (2024) W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhang, et al. HunyuanVideo: a systematic framework for large video generative models. arXiv preprint arXiv:2412.03603. Cited by: Introduction, Interactive video generation and world models.
  • Lipman et al. (2023) Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le Flow matching for generative modeling. In International Conference on Learning Representations, Cited by: Preliminaries.
  • Liu et al. (2025) Z. Liu, S. Li, E. Cousineau, S. Feng, B. Burchfiel, and S. Song Geometry-aware 4d video generation for robot manipulation. arXiv preprint arXiv:2507.01099. Cited by: External representations priors.
  • Luo et al. (2026) X. Luo, X. Xin, T. Feng, X. Guo, M. Jin, and J. Ma CoInteract: physically-consistent human-object interaction video synthesis via spatially-structured co-generation. arXiv preprint arXiv:2604.19636. Cited by: External representations priors.
  • NVIDIA et al. (2025) NVIDIA, N. Agarwal, A. Ali, M. Bala, Y. Balaji, E. Barker, T. Cai, P. Chattopadhyay, Y. Chen, Y. Cui, et al. Cosmos world foundation model platform for physical ai. arXiv preprint arXiv:2501.03575. Cited by: Introduction, Interactive video generation and world models, Baselines..
  • Po et al. (2025) R. Po, Y. Nitzan, R. Zhang, B. Chen, T. Dao, E. Shechtman, G. Wetzstein, and X. Huang Long-context state-space video world models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 8733–8744. External Links: Link Cited by: Introduction.
  • Qwen et al. (2025) Qwen, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu Qwen2.5 technical report. arXiv preprint arXiv:2412.15115. External Links: Link Cited by: Object-Token Grounding.
  • Shang et al. (2026) Y. Shang, Z. Li, Y. Ma, W. Su, X. Jin, Z. Wang, L. Jin, X. Zhang, Y. Tang, H. Su, C. Gao, W. Wu, X. Liu, D. Shah, Z. Zhang, Z. Chen, J. Zhu, Y. Tian, T. Chua, W. Zhu, and Y. Li WorldArena: a unified benchmark for evaluating perception and functional utility of embodied world models. External Links: 2602.08971, Link Cited by: Introduction, Introduction, Benchmarks and metrics..
  • Su et al. (2026) H. Su, Z. Liu, X. Jin, H. Dou, C. Hu, B. Li, Z. Liu, R. Xu, J. Fang, X. Zhang, Z. Yang, X. Yang, C. Gao, J. Yan, Y. Li, and W. Wu WorldScape Policy 2.0: empowering steerable world action modeling with reasoning-augmented memory. External Links: 2607.18840, Link Cited by: Introduction.
  • Sun et al. (2026) Z. Sun, Z. Du, X. Yang, and Z. Wu HandWorld: hand-centric unified video action generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15976–15985. Cited by: Benchmarks and metrics..
  • Tian et al. (2026) Y. Tian, Y. Jin, B. Yu, Y. Shi, H. Wu, C. H. Liu, K. Chen, and C. Huang STARRY: spatial-temporal action-centric world modeling for robotic manipulation. arXiv preprint arXiv:2604.26848. Cited by: External representations priors.
  • Unterthiner et al. (2018) T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly Towards accurate generative models of video: a new metric and challenges. arXiv preprint arXiv:1812.01717. Cited by: Benchmarks and metrics..
  • Wan Team et al. (2025) Wan Team, A. Wang, B. Ai, B. Wen, C. Mao, C. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, et al. Wan: open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314. Cited by: Introduction, Interactive video generation and world models, Baselines..
  • Wang et al. (2026a) A. N. Wang, T. Darrell, P. Izmailov, Y. Bai, and A. Bar Lifting embodied world models for planning and control. External Links: 2604.26182, Link Cited by: Introduction.
  • Wang et al. (2026b) Y. Wang, W. Ouyang, T. Wei, Y. Dong, Z. Shen, and X. Pan Hand2World: autoregressive egocentric interaction generation via free-space hand gestures. External Links: 2602.09600, Link Cited by: Interactive video generation and world models.
  • Wu et al. (2025) B. Wu, C. Zou, C. Li, D. Huang, F. Yang, H. Tan, J. Peng, J. Wu, J. Xiong, J. Jiang, Linus, Patrol, P. Zhang, P. Chen, P. Zhao, Q. Tian, S. Liu, W. Kong, W. Wang, X. He, X. Li, X. Deng, X. Zhe, Y. Li, Y. Long, Y. Peng, Y. Wu, Y. Liu, Z. Wang, Z. Dai, B. Peng, C. Li, G. Gong, G. Xiao, J. Tian, J. Lin, J. Liu, J. Zhang, J. Lian, K. Pan, L. Wang, L. Niu, M. Chen, M. Chen, M. Zheng, M. Yang, Q. Hu, Q. Yang, Q. Xiao, R. Wu, R. Xu, R. Yuan, S. Sang, S. Huang, S. Gong, S. Huang, W. Guo, X. Yuan, X. Chen, X. Hu, W. Sun, X. Wu, X. Ren, X. Yuan, X. Mi, Y. Zhang, Y. Sun, Y. Lu, Y. Li, Y. Huang, Y. Tang, Y. Li, Y. Deng, Y. Zhou, Z. Hu, Z. Liu, Z. Yang, Z. Yang, Z. Lu, Z. Zhou, and Z. Zhong HunyuanVideo 1.5 technical report. External Links: 2511.18870, Link Cited by: Baselines..
  • Wu et al. (2026) M. Wu, B. Song, R. Lin, C. Zhu, X. Feng, J. Wu, X. Chu, and K. Huang Latent temporal discrepancy as motion prior: a loss-weighting strategy for dynamic fidelity in t2v. arXiv preprint arXiv:2601.20504. Cited by: External representations priors.
  • Xiang et al. (2024) J. Xiang, G. Liu, Y. Gu, Q. Gao, Y. Ning, Y. Zha, Z. Feng, T. Tao, S. Hao, Y. Shi, Z. Liu, E. P. Xing, and Z. Hu Pandora: towards general world model with natural language actions and video states. External Links: 2406.09455, Link Cited by: Interactive video generation and world models.
  • Yan et al. (2025) H. Yan, H. Yu, Z. Zhong, W. Yuan, X. Gong, Z. Luo, C. Heyu, J. Li, W. Song, S. Zhou, and H. Li Open-world hand-object interaction video generation based on structure and contact-aware representation. arXiv preprint arXiv:2512.01677. Cited by: External representations priors.
  • Yang et al. (2026) Z. Yang, Y. Jin, L. Qi, C. Huang, and K. Chen EA-WM: event-aware generative world model with structured kinematic-to-visual action fields. arXiv preprint arXiv:2605.06192. Cited by: External representations priors.
  • Yang et al. (2024) Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng, D. Yin, X. Gu, Y. Zhang, W. Wang, Y. Cheng, T. Liu, B. Xu, Y. Dong, and J. Tang CogVideoX: text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072. Cited by: Introduction, Interactive video generation and world models, Baselines..
  • Zhang et al. (2026a) J. Zhang, X. Chen, A. Chen, C. Lv, D. Li, G. Zhou, H. Yin, H. Yuan, H. Li, J. Li, J. Zhang, J. Zhou, K. Gao, K. Yan, L. Jiang, N. Tang, P. Lin, Q. Peng, S. Yin, T. Wu, T. Yan, X. Xu, Y. Shu, Y. Zhang, Y. Wang, Y. Wang, Y. Chen, Y. Xu, Y. Huang, Y. Chen, Z. Zhang, Z. Wang, Z. Lei, Z. Liang, Z. Liu, Z. Zhou, X. Chen, and C. Wu Qwen-RobotWorld technical report: unifying embodied world modeling through language-conditioned video generation. External Links: 2606.17030, Link Cited by: Introduction, Interactive video generation and world models.
  • Zhang et al. (2026b) P. Zhang, Y. Deng, S. Sun, J. Ma, D. Wang, J. Du, Z. Pan, Y. Huang, H. Liang, S. Huang, R. Zhang, E. Xie, M. Liu, and D. Zhou PhysisForcing: physics reinforced world simulator for robotic manipulation. External Links: 2606.28128, Link Cited by: Introduction, Introduction, External representations priors.
  • Zhang et al. (2025) Y. Zhang, J. Gu, L. Wang, H. Wang, J. Cheng, Y. Zhu, and F. Zou MimicMotion: high-quality human motion video generation with confidence-aware pose guidance. In Proceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 267, pp. 74896–74910. External Links: Link Cited by: Baselines..
  • Zhao et al. (2025) B. Zhao, R. Tang, M. Jia, Z. Wang, F. Man, X. Zhang, Y. Shang, W. Zhang, W. Wu, C. Gao, X. Chen, and Y. Li AirScape: an aerial generative world model with motion controllability. External Links: 2507.08885, Link Cited by: Interactive video generation and world models.
  • Zhao et al. (2026) B. Zhao, J. Xu, W. Feng, X. Zhang, Z. Wang, H. Wang, S. Ji, Z. Wang, J. Fang, Z. Zheng, W. Zhang, Y. Shang, W. Wu, C. Gao, X. Chen, and Y. Li WorldVLN: autoregressive world action model for aerial vision-language navigation. External Links: 2605.15964, Link Cited by: Interactive video generation and world models.
  • Zhen et al. (2025) H. Zhen, Q. Sun, H. Zhang, J. Li, S. Zhou, Y. Du, and C. Gan Learning 4d embodied world models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 5337–5347. External Links: Link Cited by: Introduction, Introduction, Baselines..
  • Zhu et al. (2025) F. Zhu, H. Wu, S. Guo, Y. Liu, C. Cheang, and T. Kong IRASim: a fine-grained world model for robot manipulation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 9834–9844. External Links: Link Cited by: Introduction, Introduction, Interactive video generation and world models, Baselines..

A. Data Construction and Process Details

\pdfdest

name impact.sec:data_details xyz

A.1 RoboTwin Data

\pdfdest

name impact.sec:robotwin_data xyz

Data source and modalities.

The robot-arm training corpus comprises simulated bimanual manipulation rollouts from RoboTwin 2.0, generated with the SAPIEN physics engine and the dual-arm Aloha-AgileX embodiment. Each episode provides synchronized RGB observations, camera calibration, robot states, end-effector poses, gripper states, language instructions, and planned action trajectories. We use the head-camera stream at 1280×7041280\times 704 resolution and 10 fps, obtained by sampling every third frame from the 30 fps source video. The synchronized action at each frame is

[q1:6L,gL,q1:6R,gR]∈ℝ14,\pdfdestnameimpact.eq:robotwinactionorderxyz[\,q^{L}_{1:6},\ g^{L},\ q^{R}_{1:6},\ g^{R}\,]\in\mathbb{R}^{14},\pdfdest name{impact.eq:robotwin_{a}ction_{o}rder}xyz (20)

where qLq^{L} and qRq^{R} denote the six joint positions of the left and right arms, and gLg^{L} and gRg^{R} denote their gripper coordinates.

Task and visual diversity.

The training set contains both clean and domain-randomized rollouts. The randomized regime introduces collision-aware distractors, samples table and background appearances from more than 12,000 textures, varies illumination, perturbs the tabletop height by up to 0.03 m, and displaces the head camera. Clean backgrounds and extreme illumination each occur in 2% of randomized rollouts. The language instructions also vary across equivalent task executions.

The corpus covers 50 tasks spanning pick-and-place, container placement, stacking, object ranking, tool use, switch operation, articulated-object manipulation, bimanual handover, and coordinated dual-arm manipulation. The task distribution is non-uniform and ranges from 99 to 28,117 training windows per task. The four largest tasks are color-based block ranking (28,117 windows), size-based block ranking (27,170), bottle placement into a container (16,881), and block handover (16,723).

Window construction and action normalization.

Each training example contains 17 consecutive frames with stride one and the corresponding 17 action vectors. Boundary examples are temporally resampled or padded to preserve this fixed length. For action dimension jj, we normalize using the first and ninety-ninth dataset percentiles, p01,jp_{01,j} and p99,jp_{99,j}:

a~j=clip⁡(2​aj−p01,jp99,j−p01,j−1,−1, 1).\pdfdest​n​a​m​e​i​m​p​a​c​t.e​q:a​c​t​i​o​nn​o​r​m​a​l​i​z​a​t​i​o​n​x​y​z\widetilde{a}_{j}=\operatorname{clip}\!\left(2\frac{a_{j}-p_{01,j}}{p_{99,j}-p_{01,j}}-1,\,-1,\,1\right).\pdfdest name{impact.eq:action_{n}ormalization}xyz (21)

RGB frames and actions always share the same temporal indices. Spatial bucket sampling and matched crop-and-resize augmentation produce the 720p training inputs.

A.2 EgoDex Data

\pdfdest

name impact.sec:egodex_data xyz

Data source and modalities.

The human-hand training corpus contains approximately 256K egocentric clips from EgoDex spanning 118 tasks. Each clip is paired with camera intrinsics, per-frame camera poses, articulated hand transforms and a language instruction. Each RGB clip is paired with a rendered hand-pose video constructed from the capture-time 3D hand tracks.

The tasks cover tabletop setup and cleanup, pick-and-place, stacking, assembly and disassembly, folding and wrapping, cooking, washing, device insertion and removal, writing, drawing, and tool use.

Temporal clip construction.

Each source clip is converted into an 81-frame training clip. Clips longer than 81 frames are represented by 81 uniformly spaced indices including both endpoints; clips shorter than 81 frames repeat the final index. The same index sequence is applied to RGB and hand-pose videos, preserving frame-level correspondence between the target and condition. Source videos are 1920×10801920\times 1080 at 30 fps, and contiguous 81-frame clips therefore span 2.7 seconds. Training uses matched spatial transformations at 720p.

A.3 Object-Token Grounding

\pdfdest

name impact.sec:grounding_details xyz

Manipulated-object extraction.

We annotate each training example once with Qwen2.5-0.5B-Instruct. For both RoboTwin and EgoDex, the task instruction is the sole annotation input; no video frame, pose signal, temporal metadata, or external visual annotation is used. The model extracts compact noun phrases for the manipulated objects while preserving discriminative attributes such as color, material, shape, label, or container type. Agents, body parts, cameras, supporting surfaces, backgrounds, and non-interacting scene elements are excluded.

Complete annotation prompt.

RoboTwin and EgoDex use the same prompt template below, with {instruction} replaced by the task instruction for the current training clip.

System message
You are a precise language annotation assistant for manipulation training. Given only a task instruction, identify the concrete physical objects directly manipulated by the agent. Return only valid JSON, without Markdown, explanations, or code fences.
Use compact English noun phrases and preserve discriminative attributes stated in the instruction, including color, material, shape, printed label, object part, and container type. Exclude the robot, hands, fingers, arms, people, cameras, frames, tables, workspaces, backgrounds, and generic scene regions unless the instruction explicitly identifies one of them as the manipulated object. Do not infer objects that are not stated in the instruction. User message
Task instruction:
{instruction}
Return the following JSON object: {  “interacting_objects”: [ {“object”: “compact manipulated-object phrase”} ] } Include every manipulated object explicitly stated in the instruction. Preserve the order in which the objects appear. If the instruction refers to the same object more than once, return it once. Return JSON only.

Mapping object phrases to text tokens.

The extracted object phrase is aligned with its occurrence in the original task instruction after tokenization. All subword tokens associated with the object phrase are included in the object-token set 𝒪\mathcal{O}. We then aggregate the cross-attention associated with 𝒪\mathcal{O} across attention heads and transformer layers to obtain the object-conditioned attention map 𝐀\mathbf{A} used by ADS.

Setting Task instruction Extracted object phrase
Robot arm “Take the bottle with white printed label from the table and keep upright.” white printed-label bottle.
Human hand “Pick up the orange plate from the table.” orange plate.
Table 4: Representative instruction-only object-grounding examples.\pdfdestname impact.tab:grounding_examples xyz

A.4 Hand-Pose Video Construction

\pdfdest

name impact.sec:hand_pose_construction xyz

Pose representation.

EgoDex provides capture-time 3D hand and finger tracks together with per-frame camera calibration. We use these annotations directly and do not estimate hand pose from the compressed RGB video. Each hand is represented by 21 three-dimensional joints, accompanied by a presence mask. The camera metadata provides a 4×44\times 4 world-to-camera transformation and four intrinsic parameters for every frame.

Projection and rendering.

For frame tt, a homogeneous world point 𝐩¯\bar{\mathbf{p}} is transformed by the world-to-camera matrix 𝐄t\mathbf{E}_{t} and projected using the intrinsics (fx,fy,cx,cy)t(f_{x},f_{y},c_{x},c_{y})_{t}:

𝐩c=𝐄t​𝐩¯,u=fx​pcx/pcz+cx,v=fy​pcy/pcz+cy.\mathbf{p}_{c}=\mathbf{E}_{t}\bar{\mathbf{p}},\qquad u=f_{x}p_{c}^{x}/p_{c}^{z}+c_{x},\quad v=f_{y}p_{c}^{y}/p_{c}^{z}+c_{y}. (22)

Points with depth pcz≤0.01p_{c}^{z}\leq 0.01 are excluded. The 21 joints follow the wrist, thumb, index, middle, ring, and little-finger ordering. Each finger is connected to the wrist and rendered as an anti-aliased skeleton on a black background. We draw green edges with width 2, blue joints with radius 3, and fingertips with radius 5. Pose videos preserve the RGB resolution and are encoded at 30 fps using H.264 with the YUV420p pixel format and a constant-rate factor of 23.

Temporal and spatial alignment.

RGB and pose videos contain the same number of decoded frames and use the same 81-frame index sequence for every training segment. The two modalities therefore remain aligned by frame index throughout temporal sampling. Identical crop and resize parameters are subsequently applied to both modalities.

B. Implementation Details

\pdfdest

name impact.sec:implementation_details xyz

B.1 Robot-Arm Models

\pdfdest

name impact.sec:robot_architecture xyz

Backbone.

We instantiate the robot-arm setting with two action-conditioned diffusion transformer backbones: Wan 2.2 TI2V 5B and Cosmos-Predict 2.5 (action). The Wan 2.2 backbone contains 30 transformer blocks with hidden width d=3072d=3072, 24 attention heads, and an FFN width of 14,336. It receives a 17-frame RGB sequence together with the first-frame visual condition and the synchronized 14-DoF action trajectory. The second model is initialized from Cosmos-Predict2.5-2B and retains its action-conditioning interface for the same robot-arm setting.

Action encoder.

Let 𝐚∈ℝ17×14\mathbf{a}\in\mathbb{R}^{17\times 14} denote the normalized action trajectory. We flatten the complete trajectory into 238 scalars and process it with two independent multilayer perceptrons of identical topology:

MLPq⁡(𝐚)=Wq,2​GELU⁡(Wq,1​vec⁡(𝐚)+bq,1)+bq,2\operatorname{MLP}_{q}(\mathbf{a})=W_{q,2}\,\operatorname{GELU}(W_{q,1}\operatorname{vec}(\mathbf{a})+b_{q,1})+b_{q,2} (23)

with hidden width 4​d=12,2884d=12{,}288. The first encoder produces a dd-dimensional vector that is added to the timestep embedding. The second produces 6​d=18,4326d=18{,}432 values, reshaped to 6×d6\times d, which modulate the six adaptive normalization components in every transformer block. Learned binary condition embeddings distinguish action-conditioned and action-dropped examples. Action information is therefore injected globally through the timestep and adaptive normalization pathways.

IMPACT configuration.

For both backbones, IMPACT constructs its attention anchor from the object-token cross-attention produced by the native transformer. ADS samples K=8K=8 candidate masks by adding Gaussian perturbations with σ=0.75\sigma=0.75 to the logit of the detached attention anchor on a 5×8×125\times 8\times 12 coarse grid. The perturbed maps are trilinearly upsampled to the latent resolution, transformed with sampling temperature τ=1\tau=1 and anchor power κ=0.65\kappa=0.65, and thresholded to retain the top ρ=10%\rho=10\% of positions. Detached regional MSEs score the candidate masks, and centered scores are converted into aggregation weights with coefficient β=0.7\beta=0.7. IWS uses the resulting soft interaction map to assign per-position weights in [1,γ][1,\gamma], with γ=5\gamma=5.

B.2 Human-Hand Model

\pdfdest

name impact.sec:human_architecture xyz

Pose-conditioned input.

The human-hand model uses the same 5B transformer and conditions on an 81-frame hand-pose video. RGB targets, the first-frame reference, and pose frames are encoded by the shared Wan VAE without an additional pose encoder. The VAE produces 48-channel latents with temporal and spatial compression factors of 4, 16, and 16, yielding 21 latent timesteps for an 81-frame sequence.

The first latent timestep retains the RGB reference condition. At subsequent timesteps, the 48-channel reference latent is replaced by the temporally aligned 48-channel pose latent. Four binary mask channels are concatenated to form a 52-channel condition 𝐲\mathbf{y}. Before 3D patch embedding, the model concatenates the 48-channel noisy video latent 𝐱\mathbf{x} with 𝐲\mathbf{y}, giving 100 input channels:

𝐡0=PatchEmbed⁡([𝐱;𝐲mask+pose/ref]).\mathbf{h}_{0}=\operatorname{PatchEmbed}\bigl([\mathbf{x};\mathbf{y}_{\rm mask+pose/ref}]\bigr). (24)

The action mask applies pose replacement only to conditioned samples. RGB and pose inputs share the same temporal indices and spatial transformation.

C. Evaluation Details

\pdfdest

name impact.sec:evaluation_details xyz

C.1 Evaluation Protocols and Baseline Versions

\pdfdest

name impact.sec:baseline_protocol xyz

WorldArena evaluates 500 held-out episodes from 50 RoboTwin 2.0 tasks. All decoded submissions have a minimum resolution of 640×480640\times 480 and a frame rate of 24 fps. Text-conditioned submissions contain 121 frames. Action-conditioned submissions follow the benchmark action sequence and match the corresponding reference trajectory length. EgoDex models are evaluated at their release-specific inference settings, as detailed in attr /Border [0 0 0] goto name impact.tab:baseline_specsTable 5.

Method Version Resolution Sec. / FPS
General video generation
HunyuanVideo-1.5 25.11.20 848×480848\times 480 5 / 24
Cosmos-Predict 2.5 25.10.06 1280×7041280\times 704 5 / 16
Pose-controlled video generation
MimicMotion 24.07.08 1024×5761024\times 576 4.8 / 15
MagicPose 24.04.03 512×512512\times 512 4 / 15
VACE 25.03.11 720×1080720\times 1080 5 / 16
LOME 26.04.05 832×480832\times 480 5 / 15
Wan 2.2-based models
Wan 2.2 25.07.28 1280×7041280\times 704 5 / 24
Wan 2.2-AC / +IMPACT – 1280×7041280\times 704 3.4 / 24
Table 5: Detailed inference configurations of video generation models evaluated on EgoDex. All models use image-to-video generation; “–” denotes our adapted variants.\pdfdestname impact.tab:baseline_specs xyz

Metric computation uses the ordered decoded frames from each output. For Wan 2.2-AC and IMPACT, the 81 output frames and the pose condition use identical temporal indices. Each model uses its official sampling schedule and guidance configuration.

C.2 WorldArena Metric Definitions

\pdfdest

name impact.sec:worldarena_metrics xyz

WorldArena normalizes each raw metric to [0,1][0,1] using empirically selected boundaries and reports EWMScore as 100100 times the arithmetic mean of the 16 normalized metrics. Thus, all displayed entries are higher-is-better, including metrics whose underlying raw quantity is an error.

Dimension Metric implementation
Visual IQ (Image Quality) Frame clarity and distortion quality from the no-reference MUSIQ image-quality model; averaged over frames.
AQ (Aesthetic Quality) Per-frame visual appeal (lighting, color, composition) from the LAION aesthetic predictor; averaged over frames.
JS (JEPA Similarity) Distributional similarity between generated and ground-truth V-JEPA video features, computed using MMD with a second-order polynomial kernel.
Motion DD (Dynamic Degree) Salient motion intensity: RAFT flow between adjacent frames, averaging the top 5% motion magnitudes and applying a resolution-adaptive sigmoid.
FS (Flow Score) Overall motion intensity: mean RAFT optical-flow magnitude over all pixels and adjacent-frame pairs.
MS (Motion Smoothness) Temporal smoothness: SSIM between each actual middle frame and a VFI-Mamba interpolation from its neighbors, weighted by log motion magnitude to avoid rewarding static video.
Content SC (Subject Consistency) Subject identity/structure stability from DINO cosine similarity of each frame to both the first and previous frame, penalized for near-static video.
BC (Background Consistency) Global background/scene stability from CLIP image-feature cosine similarity to the first and previous frames, with the same low-motion penalty.
PC (Photometric Consistency) Pixel-level texture stability from forward–backward SEA-RAFT flow cycle error (raw AEPE is lower-better; the reported consistency score is inverted/normalized).
Physics Inter. (Interaction Quality) Qwen3-VL 1–5 judgment of physically plausible robot–object contact, force transmission, and interaction, divided by five.
Traj. (Trajectory Accuracy) Robot-arm trajectory agreement with ground truth: SAM 3 arm boxes/trajectories compared through normalized dynamic time warping.
3D Depth (Depth Accuracy) Monocular depth agreement to ground truth after per-video median-scale alignment; based on depth error and converted to a higher-is-better accuracy. Up to 40 frames are sampled uniformly.
Persp. (Perspectivity) Qwen3-VL judgment of 3D plausibility, including scale-versus-depth, lighting, and occlusion relations.
Control Instr. (Instruction Following) Qwen3-VL judgment of whether action type, target object, and resulting task state follow the instruction.
Sem. (Semantic Alignment) Cosine similarity between Qwen2.5-VL descriptions of generated and reference videos.
Act. (Action Following) Response diversity under three distinct instructions sharing one initial frame; average pairwise feature dissimilarity between the three generated videos.
Table 6: Meaning and implementation of all 16 WorldArena metrics.\pdfdestname impact.tab:worldarena_metric_definitions xyz

Full 16-metric results.

\pdfdest

name impact.sec:full_worldarena_results xyz

Tables attr /Border [0 0 0] goto name impact.tab:worldarena_full_quality8 and attr /Border [0 0 0] goto name impact.tab:worldarena_full_task8 report all 16 normalized WorldArena metrics for the same models and ordering. The first covers the generation-quality dimensions (visual quality, motion quality, content consistency) and the overall EWMScore; the second covers the task-oriented dimensions (physics adherence, 3D accuracy, controllability).

Model EWMScore Visual Quality Motion Quality Content Consistency
IQ AQ JS DD FS MS SC BC PC
General world models
CogVideoX 57.90 0.3582 0.3777 0.9384 0.3166 0.2189 0.7391 0.8083 0.8773 0.3580
Veo 3.1 58.87 0.6605 0.4632 0.5694 0.5450 0.1396 0.6989 0.7878 0.8710 0.3247
Wan 2.6 61.86 0.6824 0.4433 0.7229 0.7421 0.4532 0.8539 0.7517 0.8687 0.1904
Embodied world models
GigaWorld-0 53.39 0.5041 0.3991 0.4413 0.6709 0.3118 0.7811 0.7303 0.8563 0.1756
Genie Envisioner 43.65 0.2305 0.3289 0.3340 0.6930 0.0855 0.6966 0.7760 0.9024 0.2006
Vidar 51.60 0.4145 0.4068 0.5608 0.2767 0.1426 0.7973 0.7629 0.8300 0.2350
IRASim 58.12 0.3489 0.3623 0.9330 0.4139 0.2083 0.7052 0.8312 0.9068 0.3522
CtrlWorld 59.70 0.3522 0.3893 0.9185 0.4257 0.3449 0.7377 0.8411 0.9057 0.1729
Representation-guided models
TesserAct 53.23 0.3322 0.4590 0.4579 0.5150 0.2447 0.7579 0.8250 0.9238 0.2491
RoboMaster 51.84 0.3487 0.3842 0.2966 0.6124 0.1484 0.6940 0.8295 0.9123 0.3356
WoW 54.88 0.4587 0.3868 0.7440 0.4608 0.2706 0.7692 0.8161 0.9025 0.2170
Wan 2.2-based models
Wan 2.2 50.79 0.3884 0.3963 0.7575 0.4349 0.1269 0.7019 0.7400 0.8000 0.0806
Wan 2.2-AC 58.65 0.4053 0.3575 0.9343 0.4262 0.2488 0.7719 0.8250 0.9095 0.0821
Wan 2.2-AC + IMPACT (w/o ADS) 61.54 0.5236 0.4078 0.8819 0.4381 0.2611 0.7756 0.8233 0.8970 0.1090
Wan 2.2-AC + IMPACT 62.46 0.4906 0.4223 0.9050 0.4630 0.3328 0.8026 0.8191 0.8953 0.0944
Cosmos-based models
Cosmos-Predict 2.5 (text) 50.81 0.6668 0.4501 0.3126 0.5911 0.4302 0.7882 0.7488 0.8511 0.1383
Cosmos-Predict 2.5 (action) 55.91 0.4489 0.3576 0.9296 0.3994 0.0573 0.7100 0.8197 0.8894 0.3528
+ IMPACT 62.53 0.5588 0.3941 0.7366 0.5810 0.5816 0.8839 0.8026 0.8862 0.0482
Table 7: Full WorldArena results (1/2): EWMScore and the nine generation-quality metrics. EWMScore is the mean of all 16 metrics on a [0,100][0,100] scale; components are normalized to [0,1][0,1]. Higher is better (↑\uparrow); boldface and underlining mark the best and second-best per column.\pdfdestname impact.tab:worldarena_full_quality xyz
Model Physics Adherence 3D Accuracy Controllability
Inter. Traj. Depth Persp. Instr. Sem. Act.
General world models
CogVideoX 0.5940 0.3526 0.9097 0.7828 0.7268 0.8977 0.0076
Veo 3.1 0.7872 0.1231 0.7421 0.8276 0.9328 0.8607 0.0852
Wan 2.6 0.7280 0.1182 0.7144 0.8032 0.8536 0.8728 0.0992
Embodied world models
GigaWorld-0 0.5368 0.1552 0.6316 0.7596 0.6156 0.8591 0.1134
Genie Envisioner 0.2052 0.0679 0.8663 0.5284 0.2028 0.8544 0.0109
Vidar 0.5348 0.1928 0.7872 0.7592 0.5912 0.8826 0.0819
IRASim 0.5656 0.3639 0.9312 0.7788 0.6604 0.8849 0.0526
CtrlWorld 0.6212 0.4766 0.9300 0.7960 0.7272 0.8912 0.0210
Representation-guided models
TesserAct 0.5800 0.1396 0.7159 0.7920 0.6152 0.8783 0.0311
RoboMaster 0.5364 0.1158 0.8335 0.7588 0.5772 0.8761 0.0352
WoW 0.5564 0.2058 0.7283 0.7672 0.5692 0.8842 0.0434
Wan 2.2-based models
Wan 2.2 0.5184 0.1627 0.7768 0.7660 0.5376 0.8877 0.0512
Wan 2.2-AC 0.6672 0.4009 0.8657 0.8574 0.7394 0.8837 0.0089
Wan 2.2-AC + IMPACT (w/o ADS) 0.7546 0.3485 0.8974 0.9780 0.8472 0.8886 0.0143
Wan 2.2-AC + IMPACT 0.7516 0.3657 0.9026 0.9486 0.8560 0.8881 0.0559
Cosmos-based models
Cosmos-Predict 2.5 (text) 0.3872 0.0816 0.7051 0.7964 0.2664 0.7733 0.1418
Cosmos-Predict 2.5 (action) 0.5500 0.2945 0.8862 0.7644 0.5840 0.8879 0.0133
   + IMPACT 0.6360 0.2003 0.6557 0.9040 0.6360 0.8738 0.6260
Table 8: Full WorldArena results (2/2): the task-oriented metrics.\pdfdestname impact.tab:worldarena_full_task xyz

C.3 Additional Robot-Arm Cases

\pdfdest

name impact.sec:additional_robot_cases xyz In this section, we present more qualitative results of robot-arm manipulation on WorldArena.

[Uncaptioned image]

Prompt: Pick up the printed sneaker and place it on the blue mat.

Figure 5: Printed-sneaker placement comparison. Rows show WoW, CogVideoX, CtrlWorld, Wan 2.2-AC, and IMPACT (Ours) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout.\pdfdestname impact.fig:robot_case_0038 xyz
[Uncaptioned image]

Prompt: Pick up the brown-and-white bottle and place it on the blue mat.

Figure 6: Bottle placement comparison. Rows show WoW, CogVideoX, CtrlWorld, Wan 2.2-AC, and IMPACT (Ours) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout.\pdfdestname impact.fig:robot_case_0042 xyz
Refer to caption

Prompt: Pick up the red block and place it on the blue target.

Figure 7: Block placement comparison. Rows show WoW, CogVideoX, CtrlWorld, Wan 2.2-AC, and IMPACT (Ours) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout.\pdfdestname impact.fig:robot_case_0047 xyz
Refer to caption

Prompt: Grasp the lidded pot with both grippers and lift it.

Figure 8: Bimanual pot-lifting comparison. Rows show WoW, CogVideoX, CtrlWorld, Wan 2.2-AC, and IMPACT (Ours) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout.\pdfdestname impact.fig:robot_case_0059 xyz
Refer to caption

Prompt: Pick up the green block and stack it on the red block.

Figure 9: Block-stacking comparison. Rows show WoW, CogVideoX, CtrlWorld, Wan 2.2-AC, and IMPACT (Ours) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout.\pdfdestname impact.fig:robot_case_0080 xyz
Refer to caption

Prompt: Pick up the brown shoe and place it on the blue mat.

Figure 10: Shoe placement comparison. Rows show WoW, CogVideoX, CtrlWorld, Wan 2.2-AC, and IMPACT (Ours) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout.\pdfdestname impact.fig:robot_case_0108 xyz
Refer to caption

Prompt: Pick up one blue bowl and stack it inside the other blue bowl.

Figure 11: Bowl-stacking comparison. Rows show WoW, CogVideoX, CtrlWorld, Wan 2.2-AC, and IMPACT (Ours) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout.\pdfdestname impact.fig:robot_case_0143 xyz
Refer to caption

Prompt: Press the blue service bell.

Figure 12: Service-bell pressing comparison. Rows show WoW, CogVideoX, CtrlWorld, Wan 2.2-AC, and IMPACT (Ours) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout.\pdfdestname impact.fig:robot_case_0169 xyz
Refer to caption

Prompt: Place the toy hamburger and French fries in the tray.

Figure 13: Food-toy placement comparison. Rows show WoW, CogVideoX, CtrlWorld, Wan 2.2-AC, and IMPACT (Ours) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout.\pdfdestname impact.fig:robot_case_0220 xyz
Refer to caption

Prompt: Pick up the blue elephant toy from beside the black case.

Figure 14: Elephant-toy pickup comparison. Rows show HunyuanVideo-1.5, VACE, MimicMotion, Wan 2.2-AC, and IMPACT (Ours) under the same initial observation and task instruction; columns are uniformly sampled frames from each rollout.\pdfdestname impact.fig:hand_case_01 xyz

C.4 Additional Human-Hand Cases

\pdfdest

name impact.sec:additional_hand_cases xyz In this section, we present more qualitative results of human-hand manipulation on EgoDex.

[Uncaptioned image]

Prompt: Grasp the transparent container and lift it from the shelf.

Figure 15: Container-lifting comparison. Rows show HunyuanVideo-1.5, VACE, MimicMotion, Wan 2.2-AC, and IMPACT (Ours) under the same initial observation and task instruction; columns are uniformly sampled frames from each rollout.\pdfdestname impact.fig:hand_case_02 xyz
[Uncaptioned image]

Prompt: Grasp the white cylindrical object and lift it from the base.

Figure 16: Cylindrical-object lifting comparison. Rows show HunyuanVideo-1.5, VACE, MimicMotion, Wan 2.2-AC, and IMPACT (Ours) under the same initial observation and task instruction; columns are uniformly sampled frames from each rollout.\pdfdestname impact.fig:hand_case_03 xyz
Refer to caption

Prompt: Pick up the two white drawers and stack them together.

Figure 17: Drawer-stacking comparison. Rows show HunyuanVideo-1.5, VACE, MimicMotion, Wan 2.2-AC, and IMPACT (Ours) under the same initial observation and task instruction; columns are uniformly sampled frames from each rollout.\pdfdestname impact.fig:hand_case_04 xyz
Refer to caption

Prompt: Grasp and adjust the colorful block structure.

Figure 18: Block-structure adjustment comparison. Rows show HunyuanVideo-1.5, VACE, MimicMotion, Wan 2.2-AC, and IMPACT (Ours) under the same initial observation and task instruction; columns are uniformly sampled frames from each rollout.\pdfdestname impact.fig:hand_case_05 xyz
Refer to caption

Prompt: Pick up the red book from the table.

Figure 19: Book-pickup comparison. Rows show HunyuanVideo-1.5, VACE, MimicMotion, Wan 2.2-AC, and IMPACT (Ours) under the same initial observation and task instruction; columns are uniformly sampled frames from each rollout.\pdfdestname impact.fig:hand_case_06 xyz
Refer to caption

Prompt: Stack the orange bowl on the white bowl.

Figure 20: Bowl-stacking comparison. Rows show HunyuanVideo-1.5, VACE, MimicMotion, Wan 2.2-AC, and IMPACT (Ours) under the same initial observation and task instruction; columns are uniformly sampled frames from each rollout.\pdfdestname impact.fig:hand_case_07 xyz