arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2609.03774v1 [cs.AI] 03 Sep 2026

Rethinking World Models for Safety-Critical Embodied Systems

Kailang Ma1, Heye Huang1,đź–‚, Inhi Kim1, Kitae Jang1

1Cho Chun Shik Graduate School of Mobility, Korea Advanced Institute of Science and Technology, Daejeon 34051, South Korea.

đź–‚ Corresponding author. E-mail: heye.huang@kaist.ac.kr

1  Introduction

World-model research has progressed from compact latent dynamics for control (Ha and Schmidhuber, 2018) to generative, controllable, and interactive simulators of embodied environments. GAIA-2 (Russell et al., 2025) advances controllable multi-view synthesis; UniSim (Yang et al., 2024) enables interactive real-world simulation; and TrafficBots (Zhang et al., 2023), OccWorld (Zheng et al., 2024), and DriveDreamer4D (Zhao et al., 2025) extend multi-agent behavior modeling, 3D occupancy prediction, and 4D scene representation. These advances extend generative and predictive capabilities, but do not necessarily improve a model’s understanding of what matters for safety: even an accurate and visually realistic model may overlook evidence that should prompt a change in the current action.

Consider an unprotected left turn. An unobstructed crossing may be the most likely future, whereas a pedestrian emerging from behind an occluder, despite its lower probability, may determine whether acceleration is defensible. Waiting, creeping, and accelerating also alter how other road users respond, thereby changing both available information and recovery space. For safety-critical embodied systems, the key question is therefore not only what is likely to happen, but what could change the action, how alternative actions alter consequences, and whether accumulated experience should revise the current judgment. Objectives dominated by likelihood, reconstruction, or visual fidelity may suppress precisely this decision-relevant evidence.

We therefore organize world-model capability as a three-level cognitive progression. As shown in Fig. 1, predictive modeling asks what is most likely to happen; consequential modeling asks what could change the action; and reflective modeling asks whether the available evidence is sufficient to support action. These capabilities are cumulative rather than mutually exclusive. We define consequential futures as credible branches that alter the feasibility, acceptability, or recoverability of candidate actions.

Building on this view, we propose the Risk-Informed World Model (RIWM) as a research direction for safety-critical embodied systems. RIWM organizes world modeling around action consequences: decision-relevant evidence shapes representation, counterfactual reasoning, episodic memory, and runtime safety assurance, while execution feedback continuously revises that evidence. It treats physical, social, and operational consequences as distinct decision domains, with epistemic uncertainty qualifying the evidence used to assess them. Risk-informed modeling does not imply eliminating risk, achieving complete causal identification, or providing unconditional safety guarantees. Instead, consequences, epistemic uncertainty, and recoverability should materially influence what the model represents, reasons about, remembers, and when it should act, revise, sense, defer, or abstain. This direction extends efforts toward unified environmental understanding, decision making, and control in autonomous driving (J. Wang et al., 2021), while targeting safety-critical embodied systems more broadly.

2  Three structural mismatches

2.1  Likelihood-risk mismatch

The futures that matter most for action are not necessarily those that occur most often. As scenario dimensionality and interaction complexity increase, safety-critical events occupy a rapidly shrinking fraction of available data—the curse of rarity in autonomous-driving safety research (H. X. Liu and Feng, 2024). Safety validation must therefore prioritize informative and discriminative scenarios rather than rely on inefficient random exposure (Feng et al., 2023). As the unprotected-left-turn example shows, a lower-probability but consequential credible branch may govern the decision even when it is not the most likely future. Probability alone therefore cannot express a branch’s importance to the current action. A scalar probability-severity score similarly obscures important decision structure. Equivalent scores may correspond to different interaction responses, threatened requirements, or recovery margins, yet these distinctions determine whether the system should act, revise, sense, defer, or abstain.

This limitation becomes more consequential under epistemic uncertainty. Aleatoric uncertainty concerns irreducible variability or observation noise (Kendall and Gal, 2017), whereas epistemic uncertainty reflects limited knowledge about the model or data-generating process and may be reducible with additional evidence (HĂĽllermeier and Waegeman, 2021). Crucially, epistemic uncertainty may mean that the relevant branch has not been represented at all. Once occlusion topology, conflicting actors, or recovery space has been omitted or collapsed in the learned state, downstream risk scoring cannot reliably reconstruct it. Safety relevance must therefore shape representation and branch allocation rather than be appended after prediction.

Refer to caption
Fig. 1: Cognitive progression across predictive, consequential, and reflective world modeling, from a prediction-centric to a decision-centric orientation.

2.2  Prediction-intervention mismatch

An embodied agent does not merely predict the future; its actions help create it. At the left-turn intersection, waiting, creeping, and accelerating alter not only the ego trajectory but also others’ responses, available information, and recovery space. The future is therefore not a policy-independent sequence, but a branching structure jointly shaped by action, interaction, information change, and consequence.

Several levels of reasoning should remain distinct. Action-conditioned prediction estimates what may follow a candidate action; intervention-sensitive rollout models how responses vary with that action; counterfactual comparison contrasts the resulting action-response-consequence chains; and causal identification requires explicit causal assumptions (Pearl, 2009). Interactive prediction and multi-agent simulation model how futures vary with candidate actions (Z. Huang et al., 2023; Sun et al., 2022; Suo et al., 2021), but action conditioning within a predictive or simulated model does not by itself identify how the world would respond to an unchosen intervention. RIWM therefore requires reliable action-conditioned rollout and bounded counterfactual comparison, while making the assumptions behind stronger causal claims explicit.

2.3  Horizon-accumulation mismatch

A sequence of locally acceptable decisions can still accumulate into an unsafe outcome. A mildly assertive merge may remain collision-free while triggering a downstream braking cascade, whereas repeated conservative waiting may avoid immediate conflict yet produce deadlock or provoke less predictable behavior. Frame-wise safety therefore does not guarantee long-horizon recoverability.

This mismatch cannot be resolved simply by extending the prediction horizon or context window. Longer prediction looks farther ahead within a rollout, whereas evidence continuity preserves why an event mattered across successive observations, actions, and feedback. The system should therefore retain the consequential structure of experience—what action was taken, how the environment responded, what consequence followed, which assumptions failed, and whether recovery succeeded—rather than indiscriminate historical detail. Existing agent-memory and lifelong-learning research provides mechanisms for retaining and revising experience over time (Meng et al., 2025; Park et al., 2023), but such mechanisms do not by themselves preserve its safety relevance. RIWM therefore requires failures, near misses, and recoveries to be retained with their consequences, provenance, failed assumptions, and recovery outcomes so that they can revise current representation and future branching.

3  The risk-informed world model

The three structural mismatches require world models to organize evidence around consequences, intervention, uncertainty, and recovery. RIWM organizes consequences into physical (collision and recovery space), social (norms, interactions, and responses), and operational (progress, efficiency, information gain, and deadlock) domains. RIWM comprises four interdependent capabilities: decision-relevant representation preserves information that may change the action boundary; counterfactual reasoning compares responses and consequences under candidate actions; safety-critical episodic memory uses past failures, near misses, and recoveries to revise current judgment; and runtime safety assurance evaluates whether available evidence is sufficient to support action under explicit assumptions. Fig. 2 illustrates the closed-loop interaction between the model and the environment, integrating these capabilities with decision-relevant consequences, epistemic uncertainty, and execution feedback. Temporally, episodic memory looks backward, counterfactual reasoning unfolds forward, and runtime safety assurance operates concurrently with action, while epistemic uncertainty constrains the strength of evidence used to judge consequences.

Refer to caption
Fig. 2: Conceptual architecture of a risk-informed world model for safety-critical embodied systems.

3.1  Decision-relevant representation

Decision relevance concerns whether information changes the feasibility, acceptability, recoverability, ranking, or evidential support of candidate actions, rather than its visual salience or frequency. Decision-relevant representation consequently prioritizes relations that support action comparison. In driving, these include conflict regions, occlusion, right of way, social norms, escape space, uncertainty in others’ responses, and the model’s knowledge boundary.

Existing work provides partial but complementary ingredients for decision-relevant representation. K-Risk (H. Huang et al., 2026a) links structured driving trajectories and event-level metadata with extracted high-risk events and LLM-generated semantic annotations, while RiskNet (Q. Liu et al., 2026) represents directional interaction risk. However, neither component alone determines which evidence changes an action boundary across physical, social, and operational consequences. These consequence domains may coexist and need not be compressed into a single cost; instead, their distinct effects on action selection should remain identifiable. Epistemic uncertainty is not a consequence domain but qualifies the evidence supporting consequence assessment and action selection.

Two futures may have similar pixels and short-term trajectories yet differ decisively in recoverability when a guardrail and adjacent vehicle remove lateral escape space. Representation should thus be evaluated not only by image, occupancy, or trajectory accuracy, but also by whether changes in occlusion, right of way, escape space, and uncertainty revise counterfactual branches and action ranking. For safety-critical embodied systems, the representational objective should therefore shift from preserving everything accurately to preserving what may change the decision. Consequence, recovery margin, and evidence sufficiency should shape representation before execution, rather than being appended as a downstream risk score.

3.2  Counterfactual reasoning

Exhaustive simulation of all futures is neither feasible nor necessary. Rollout breadth, horizon, and fidelity should instead adapt to decision relevance. Familiar, weakly interactive situations with ample recovery margin may require only compact prediction, whereas severe occlusion, asymmetric responses, relevant failure memories, or narrowing recovery space should trigger broader, longer, or more detailed rollouts. Modeling resources should therefore be allocated dynamically rather than uniformly.

Under a fixed branch budget, counterfactual reasoning should preserve both the most likely baseline and credible alternatives that may change the safety boundary. Such alternatives should be supported by current observation, calibrated uncertainty, valid episodic memory, or bounded perturbation rather than arbitrary hazard generation. Generability is not credibility. Branch coverage should therefore make explicit which futures were considered, what evidence supports them, and which critical variables remain unknown.

Existing approaches provide partial foundations for such comparisons. Human-aware planning represents how ego actions may influence human responses (Sadigh et al., 2016), while LEAD (H. Huang et al., 2025) adapts game-theoretic decisions to changing interactions. These formulations support action-dependent comparison but do not by themselves establish that the resulting branches are causally identified or sufficiently complete for safety-critical decisions. RIWM should therefore concentrate finite computation on credible, decision-relevant alternatives and support a broader set of responses—including acting, revising, sensing, deferring, and abstaining—when evidence is insufficient.

3.3  Safety-critical episodic memory

Agent cognitive architectures distinguish memory stores and operations (Sumers et al., 2024), and generative agents demonstrate retrieval and reflection over recorded experience (Park et al., 2023). Lifelong robotic learning additionally studies how knowledge can be preserved and combined over time (Meng et al., 2025). RIWM builds on these foundations but requires more than a history indexed by visual similarity: a reusable safety episode should link the situation, action, others’ responses, consequence, recovery, assumptions, confidence, provenance, and time. Retention should prioritize failures, near misses, successful recoveries, and other experiences with high consequence, surprise, recovery, or transfer value.

Retrieval should emphasize interaction structure and action-response relations rather than appearance alone. Equally important, memory must remain revisable: changes in sensors, policies, traffic rules, or operating conditions may invalidate earlier experience. Memories should therefore support conflict tracking, revalidation, down-weighting, or forgetting under distribution shift. In RIWM, consequential experience is not merely training data but auditable evidence that can revise current interpretation, counterfactual branching, and recovery strategy at runtime.

3.4  Runtime safety assurance

World-model outputs should inform runtime safety assurance but remain distinct from constraint enforcement. MPC (Rawlings et al., 2020), CBFs (Ames et al., 2019), reachability analysis (Althoff et al., 2021), and runtime assurance (Hobbs et al., 2023) provide executable mechanisms for enforcing safety constraints. RIWM instead supplies decision-relevant information, such as conflicts, occlusions, interactions, recovery space, and evidence sufficiency, that these mechanisms can translate into constraints, action filters, or fallback conditions.

Probability calibration alone cannot resolve structural epistemic uncertainty arising from missing variables, omitted relations, or incorrect model structure. A hazard outside the represented hypothesis space cannot be recovered by better calibration. RIWM must therefore expose model boundaries and monitor whether critical assumptions remain valid at runtime. Safety conclusions should consequently remain conditional on assumptions about dynamics, state estimation, uncertainty bounds, and operating conditions. RIWM does not replace assurance mechanisms; it evaluates whether their premises are supported by current evidence and triggers revision, active sensing, deferral, or abstention when they are not. World-model output becomes safety-relevant only when its applicability and recovery conditions can be checked and translated into enforceable constraints.

4  Open challenges and outlook

Translating RIWM from a conceptual framework into a practical paradigm requires progress in how consequential evidence is defined, tested, retained, and connected to executable safety mechanisms.

4.1  Open challenges

What constitutes a decision-relevant consequence? Safety-critical decisions involve physical harm, social interaction, comfort, rule compliance, task completion, and recoverability. These dimensions cannot always be reduced to a single scalar risk score. A key challenge is to determine which consequential differences are sufficient to change an action and how heterogeneous forms of evidence and constraints should be combined.

When does action conditioning support counterfactual reasoning? Action-conditioned generation is not equivalent to reliable intervention reasoning. Observational data record outcomes only for executed actions, while responses to unchosen actions remain unobserved. Counterfactual claims therefore require explicit causal assumptions, uncertainty bounds, and distributional limits, supported where possible by interactive simulation, controlled experiments, natural experiments, or learned response models.

When should episodic memory be trusted? Past experience may become invalid under changes in sensing, policy, traffic rules, or operating conditions, whereas excessive forgetting may cause repeated failures. Memory provenance, confidence decay, conflict tracking, versioning, and revalidation are therefore central challenges for long-term safety memory.

How can learned consequences become enforceable constraints? RIWM may represent consequences in semantic or generative form, whereas runtime assurance typically operates on states, reachable sets, constraints, or verifiable temporal properties. Establishing a reliable interface between these representations while preserving uncertainty and applicability conditions remains a major challenge.

When is there enough evidence to act? Safety-critical systems must avoid both underestimating long-tail risk and becoming paralyzed by unbounded hazard enumeration. The key question is whether available evidence supports a recoverable decision. Computation and caution should therefore adapt to consequence, epistemic uncertainty, recovery margin, and information value, while allowing the system to act, revise, sense, defer, or abstain.

4.2  Outlook

Future world models will likely achieve longer horizons, stronger controllability, and greater visual consistency. Their value in safety-critical embodied systems, however, will depend less on generating ever more futures than on recognizing the consequential ones, especially when candidate actions induce different responses, evidence is incomplete, or recovery space is narrowing.

This perspective calls for a shift from prediction-centric to decision-centric world modeling. Risk should not be appended after prediction; it should shape what the model represents, reasons about, and remembers. Decision-Relevant Representation should preserve information that changes action boundaries; Counterfactual Reasoning should focus finite computation on credible alternatives; Safety-Critical Episodic Memory should allow relevant experience to revise current judgment; and Runtime Safety Assurance should connect these judgments to executable constraints.

In the longer term, safety-critical world modeling may progress from predicting what may happen, to identifying what could change the action, and ultimately to recognizing when the available evidence is insufficient to act. The next frontier may therefore lie not in imagining more futures, but in identifying which futures matter, revising judgment through relevant experience, and recognizing when to act, sense, defer, or abstain.

Acknowledgements

This research was conducted by KAIST as part of a joint research project under the Korea Institute of Science and Technology Information (KISTI) R&D program, “Development of the Next-Generation Integrated Wired/Wireless Communication Gateway (X-Gateway).”

References

  • Althoff et al. (2021) Althoff, M., Frehse, G., Girard, A., 2021. Set propagation techniques for reachability analysis. Annu Rev Control Robot Auton Syst, 4, 369–395.
  • Ames et al. (2019) Ames, A.D., Coogan, S., Egerstedt, M., Notomista, G., Sreenath, K., Tabuada, P., 2019. Control barrier functions: Theory and applications. In: Proceedings of the European Control Conference (ECC), 3420–3431. https://doi.org/10.23919/ECC.2019.8796030
  • Feng et al. (2023) Feng, S., Sun, H., Yan, X., Zhu, H., Zou, Z., Shen, S., et al., 2023. Dense reinforcement learning for safety validation of autonomous vehicles. Nature, 615, 620–627. https://doi.org/10.1038/s41586-023-05732-2
  • Ha and Schmidhuber (2018) Ha, D., Schmidhuber, J., 2018. World models. https://doi.org/10.48550/arXiv.1803.10122
  • Hobbs et al. (2023) Hobbs, K.L., Mote, M.L., Abate, M.C.L., Coogan, S.D., Feron, E.M., 2023. Runtime assurance for safety-critical systems: An introduction to safety filtering approaches for complex control systems. IEEE Control Syst Mag, 43(2), 28–65. https://doi.org/10.1109/MCS.2023.3234380
  • H. Huang et al. (2025) Huang, H., Liu, J., Zhang, B., Zhao, S., Li, B., Wang, J., 2025. LEAD: Learning-enhanced adaptive decision-making for autonomous driving in dynamic environments. IEEE Trans Intell Transp Syst, 26(5), 6142–6156. https://doi.org/10.1109/TITS.2025.3531293
  • H. Huang et al. (2026a) Huang, H., Li, J., Zhou, Z., Liang, P., Wu, M., Jang, K., et al., 2026a. A knowledge-augmented dataset of high-risk driving scenarios with LLM annotations for autonomous driving. https://doi.org/10.48550/arXiv.2607.07103
  • Z. Huang et al. (2023) Huang, Z., Liu, H., Lv, C., 2023. GameFormer: Game-theoretic modeling and learning of transformer-based interactive prediction and planning for autonomous driving. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 3880–3890. https://doi.org/10.1109/ICCV51070.2023.00361
  • HĂĽllermeier and Waegeman (2021) HĂĽllermeier, E., Waegeman, W., 2021. Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods. Mach Learn, 110, 457–506.
  • Kendall and Gal (2017) Kendall, A., Gal, Y., 2017. What uncertainties do we need in Bayesian deep learning for computer vision? Adv Neural Inf Process Syst, 30, 5574–5584.
  • H. X. Liu and Feng (2024) Liu, H.X., Feng, S., 2024. Curse of rarity for autonomous vehicles. Nat Commun, 15, 4808.
  • Q. Liu et al. (2026) Liu, Q., Huang, H., Zhao, S., Shi, L., Ahn, S., Li, X., 2026. RiskNet: Interaction-aware risk forecasting for autonomous driving in long-tail scenarios. Transport Res E-Log, 205, 104478. https://doi.org/10.1016/j.tre.2025.104478
  • Meng et al. (2025) Meng, Y., Bing, Z., Yao, X., Chen, K., Huang, K., Gao, Y., et al., 2025. Preserving and combining knowledge in robotic lifelong reinforcement learning. Nat Mach Intell, 7(2), 256–269. https://doi.org/10.1038/s42256-025-00983-2
  • Park et al. (2023) Park, J.S., O’Brien, J.C., Cai, C.J., Morris, M.R., Liang, P., Bernstein, M.S., 2023. Generative agents: Interactive simulacra of human behavior. In: Proceedings of the ACM Symposium on User Interface Software and Technology (UIST), Article 2, 1–22. https://doi.org/10.1145/3586183.3606763
  • Pearl (2009) Pearl, J., 2009. Causality: Models, Reasoning, and Inference. 2nd edn. Cambridge: Cambridge University Press.
  • Rawlings et al. (2020) Rawlings, J.B., Mayne, D.Q., Diehl, M., 2020. Model Predictive Control: Theory, Computation, and Design. 2nd edn. Santa Barbara, CA: Nob Hill Publishing.
  • Russell et al. (2025) Russell, L., Hu, A., Bertoni, L., Fedoseev, G., Shotton, J., Arani, E., et al., 2025. GAIA-2: A controllable multi-view generative world model for autonomous driving. https://doi.org/10.48550/arXiv.2503.20523
  • Sadigh et al. (2016) Sadigh, D., Sastry, S., Seshia, S.A., Dragan, A.D., 2016. Planning for autonomous cars that leverage effects on human actions. In: Proceedings of Robotics: Science and Systems (RSS). https://doi.org/10.15607/RSS.2016.XII.029
  • Sumers et al. (2024) Sumers, T.R., Yao, S., Narasimhan, K., Griffiths, T.L., 2024. Cognitive architectures for language agents. Trans Mach Learn Res, 2024.
  • Sun et al. (2022) Sun, Q., Huang, X., Gu, J., Williams, B.C., Zhao, H., 2022. M2I: From factored marginal trajectory prediction to interactive prediction. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 6533–6542.
  • Suo et al. (2021) Suo, S., Regalado, S., Casas, S., Urtasun, R., 2021. TrafficSim: Learning to simulate realistic multi-agent behaviors. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10395–10404. https://doi.org/10.1109/CVPR46437.2021.01026
  • J. Wang et al. (2021) Wang, J., Huang, H., Li, K., Li, J., 2021. Towards the unified principles for Level 5 autonomous vehicles. Engineering, 7(9), 1313–1325. https://doi.org/10.1016/j.eng.2020.10.018
  • Yang et al. (2024) Yang, M., Du, Y., Ghasemipour, K., Tompson, J., Schuurmans, D., Abbeel, P., 2024. Learning interactive real-world simulators. In: Proceedings of the International Conference on Learning Representations (ICLR).
  • Zhang et al. (2023) Zhang, Z., Liniger, A., Dai, D., Yu, F., Van Gool, L., 2023. TrafficBots: Towards world models for autonomous driving simulation and motion prediction. In: Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), 1522–1529.
  • Zhao et al. (2025) Zhao, G., Ni, C., Wang, X., Zhu, Z., Zhang, X., Wang, Y., et al., 2025. DriveDreamer4D: World models are effective data machines for 4D driving scene representation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 12015–12026. https://doi.org/10.1109/CVPR52734.2025.01122
  • Zheng et al. (2024) Zheng, W., Chen, W., Huang, Y., Zhang, B., Duan, Y., Lu, J., 2024. OccWorld: Learning a 3D occupancy world model for autonomous driving. In: Proceedings of the European Conference on Computer Vision (ECCV), 55–72. https://doi.org/10.1007/978-3-031-72624-8_4
[Uncaptioned image]

Kailang Ma received the B.S. degree in information security and the M.S. degree in cyberspace security from Beihang University, Beijing, China, in 2021 and 2024, respectively. He was an Algorithm Engineer at WeRide, an autonomous driving company, from 2024 to 2026. He is currently pursuing the Ph.D. degree with the Cho Chun Shik Graduate School of Mobility, Korea Advanced Institute of Science and Technology (KAIST), under the supervision of Prof. Heye Huang. His research interests include safe and trustworthy autonomy, generative AI for autonomous driving, world models, and risk-sensitive decision making.

[Uncaptioned image]

Heye Huang received the Ph.D. degree in mechanical engineering from Tsinghua University in 2023. She was a Postdoctoral Associate at MIT SMART center and UW–Madison, in 2025 and 2023, respectively. She is currently an Assistant Professor with the Cho Chun Shik Graduate School of Mobility, Korea Advanced Institute of Science and Technology. Her current research interests include safe and trustworthy autonomy, generative AI, risk-sensitive decision making, and human-centered AI for autonomous systems.

[Uncaptioned image]

Inhi Kim received the Ph.D. degree from the Department of Civil Engineering, The University of Queensland, in 2014. He was a Lecturer and later a Senior Lecturer with the Department of Civil Engineering, Monash University, Australia. He is currently an associate professor with the Cho Chun Shik Graduate School of Mobility, KAIST. His current research interests include traffic planning and operation, digital twins, traffic simulation, and the safety and efficiency evaluation of connected vehicles.

[Uncaptioned image]

Kitae Jang received the Ph.D. degree in civil and environmental engineering from the University of California, Berkeley, in 2011. He is currently a Professor with the Cho Chun Shik Graduate School of Mobility, KAIST. He also serves as the Dean of KAIST Academy and the Director of the Center for Excellence in Learning and Teaching and the Research Center for Eco-friendly and Smart Vehicles. His current research interests include transportation energy and environmental policy, traffic operations, intelligent transportation systems, and transportation systems for healthy and safe communities.