Towards a Belief-Based World Model for LLM Agents
Abstract
Large language models (LLMs) are being used as policies for autonomous decision-making and planning in many domains. Despite their strong reasoning capabilities, LLMs struggle with long-horizon tasks, especially under partial observability. World models are a promising way to enhance policy performance, both during training and inference. During inference, agents currently use world models to simulate the consequences of candidate actions before choosing an action, which can improve decision-making. However, we argue that simulation alone is an incomplete interface for decision-making under partial observability: simulation does not adequately capture uncertainty about the current state, which agents may need for accurate decision-making. We address this limitation with Belief-Based World Models (BB-WMs), which maintain a belief that LLMs can query to access information on what is known and uncertain about the current state. Before developing methods to learn accurate BB-WMs, this paper focuses on a more fundamental question: does exposing a world model’s belief directly to an LLM policy improve decision-making? Our results show that giving LLM agents access to beliefs improves task performance under partial observability, while remaining complementary to existing simulation-based world models. Code: https://github.com/skumar-ml/belief-world-models.
1 Introduction
Consider an autonomous system that plans and executes a series of actions to achieve some goal. It’s hypothesized that creating models of the world (or world models) can improve the learning of the system’s decision-making function, or policy. Instead of needing costly real-world interactions, a policy can be trained on diverse, imagined future trajectories from a world model (Hafner et al., 2020; Ha and Schmidhuber, 2018). World modeling training objectives can also be used alongside reinforcement learning (RL) to improve the policy itself (Yu et al., 2026; Schwarzer et al., 2021; Copet et al., 2025). Recent world modeling efforts have primarily gone towards improving the policy at training time (Bruce et al., 2024; Agarwal et al., 2025; Ren et al., 2025). Once trained, the policy does not use the world model within a control loop (i.e., model-free control). While this is a promising direction, it is unclear how well such policies generalize to novel situations at test time. To help with this, the policy may also have access to a world model at test-time (i.e., model-based control), where foresight from the world model enables test-time reasoning (e.g., planning) in difficult or novel situations. This direction has received comparatively less attention and is the focus of this work.
Recent model-based control methods (Maes et al., 2026; Zhou et al., 2025; Hao et al., 2023; Janner et al., 2022; Schrittwieser et al., 2020) allow the policy to interact with a world model through a simulation interface: given a representation of the current state and a candidate action (or action sequence), what future trajectory follows? We raise a broader question: is simulation the only interface that a world model should expose to a policy? While simulation may be sufficient for training model-free policies, we argue that it is inadequate for model-based control. Different types of actions require different information from a world model. Consider pragmatic actions, which primarily advance the agent toward its goal (e.g., picking up a mug in order to drink coffee). Simulation is well-suited to these actions because it provides foresight by evaluating the consequences of acting. In contrast, epistemic actions are information-gathering actions that primarily reduce uncertainty about the current state of the environment (Kaelbling et al., 1998; Kirsh and Maglio, 1994). In this setting, the policy benefits from a belief—what is known and what is uncertain—about the current state to reason about whether epistemic actions are needed. Simulation can only expose uncertainty indirectly through variation across predicted futures, which may conflate uncertainty about the present state with stochasticity in the environment’s state transition dynamics. Moreover, accurately recovering current-state uncertainty from simulation may require many samples and may still miss low-probability outcomes. Thus, we argue that the current-state belief should be exposed directly rather than reconstructed indirectly from future simulations (exemplified in Fig. 1).
| Setting | Simulation Interface | Belief | |
|---|---|---|---|
|
![]() |
|
![]() |
Recently, large language model (LLM) agents have emerged as general-purpose policies in open-ended environments, from software engineering to robotics (Yang et al., 2024; Drouin et al., 2024; Huang et al., 2022; Ahn et al., 2022). Given their general-purpose abilities, LLMs may be able to implicitly world model while planning actions. However, prior works have found that LLMs struggle with maintaining and updating beliefs in complex, long-horizon partially observable tasks (Rahmani, 2026; Jha et al., 2026; Zou et al., 2026). Recent world models for LLMs do not adequately alleviate this burden; they predominantly expose action-conditioned foresight while leaving belief estimation to the agent. The agent must therefore implicitly maintain uncertainty about the current state while simultaneously reasoning about action selection, which is likely a suboptimal division of responsibilities. This motivates a modular setting: the trained LLM performs general reasoning and action selection, while a separate world model performs current and future state belief estimation. The remaining question is how these independent components should communicate.
Separating belief estimation from action selection has precedence in decision-making under partial observability (Kaelbling et al., 1998) and has subsequently been explored in deep reinforcement learning (RL), where dedicated belief estimation can substantially improve decision-making (Hafner et al., 2020; Igl et al., 2018; Singh et al., 2021; Wang et al., 2023; Chen et al., 2022). However, these approaches typically train the policy around a particular belief representation, and the control loop—the manner in which the policy and world model interact—is determined by the algorithm designer. This largely avoids the need to design an explicit interface between world model and policy. The LLMs considered in this work present a different setting. The LLM may be a pretrained, frozen, or even proprietary general-purpose reasoner that cannot readily be retrained around an arbitrary belief representation. At the same time, the agent’s general reasoning abilities make a fixed control loop unnecessary: rather than specifying how the agent must use the world model, we can expose world model capabilities through an interpretable interface and allow the LLM itself to decide when they are useful. We therefore propose exposing belief in natural language, enabling an LLM to initiate queries about current or future states without being trained around a particular belief representation.
In short, we argue that the simulation interface of current world models for LLM agents should be augmented to explicitly expose beliefs over the current state. To achieve this, we introduce Belief-Based World Models (BB-WMs), which should maintain a belief over the current state, update it as new observations arrive, and propagate it forward during action-conditioned simulation. This affords an LLM agent two complementary interfaces: natural language queries ask about the world model’s current belief when reasoning about epistemic actions, while simulation evaluates the consequences of pragmatic actions. Before designing methods for learning scalable Belief-Based World Models, this paper focuses on a more fundamental question: if a world model represents a belief over the state, does exposing this belief to the LLM agent improve decision-making? To isolate this question, we intentionally design and study hand-crafted, benchmark-specific BB-WMs, abstracting away the separate challenge of learning accurate beliefs from experience.
2 Related Works
World Models for LLMs: There have been efforts to equip LLM agents with separate world models that provide foresight into the consequences of candidate actions. WMA (Chae et al., 2025) finds that out-of-the-box LLMs cannot accurately simulate the outcome of actions in web interaction tasks, so they learn a separate world model to provide action-conditioned foresight. WALL-E (Zhou et al., 2025) replaces generative world models with a neurosymbolic world model, which checks the validity of an LLM-proposed action before it acts; if the action is predicted to be invalid, the LLM is asked to replan. DreamPhase (Hamidi et al., 2026) learns a latent world model that generates hypothetical future observations from predicted future latent states, scores the generated futures with a learned value function, and distills feedback into natural language to condition a frozen LLM agent. These methods predominantly expose an action-conditioned simulation interface to the policy.
Interestingly, recent work by Qian et al. (Qian et al., 2026) shows that providing an LLM agent with a ground-truth simulator alone does not translate into improved decision-making at test-time; agents often fail to invoke simulation when useful, misuse its outputs, or even degrade when simulation is enforced. Their analysis suggests that a key challenge is in when an agent decides to query a world model and how it integrates the resulting information into its reasoning. These findings motivate our efforts to extend world models with belief modeling capabilities, allowing the LLM to benefit from a different type of information.
Belief Representation in LLM Agents: Some works have sought to improve LLM agents by maintaining estimates of the current environment state rather than relying solely on an agent’s ability to reason over the interaction history. Approaches such as Statler, QuBE, StateAct, and ABBEL (Yoneda et al., 2024; Kim et al., 2024; Rozanov and Rei, 2024; Lidayan et al., 2025) maintain compact textual or structured representations of task-relevant state that are updated as new observations become available. These representations emphasize a single estimate or summary of the current state rather than explicitly preserving uncertainty over many possible states. Some recent work has begun to model this uncertainty directly. BeliefMem (Liao et al., 2026) retains multiple candidate conclusions together with probabilities, allowing competing interpretations of past evidence to coexist and updating them as additional evidence is observed. Agent-BRACE (Singh et al., 2026) more directly adopts the POMDP notion of belief for LLM agents, representing the current state uncertainty and conditioning the policy on this structured belief. These works demonstrate growing interest in explicitly representing uncertainty about the current state and making this information available to the agent during decision making. However, these approaches are not world models. They focus on estimating and maintaining the environment’s current state but do not provide an action-conditioned simulation interface for reasoning about the future. BB-WMs seek to bridge simulation with belief estimation: the policy can directly access the world model’s belief when reasoning about uncertainty in the current state, while continuing to obtain foresight through simulation.
3 Belief-Based World Models
Preliminaries: An agent, at any timestep , receives an observation , takes action in the environment, and receives the subsequent observation . The agent selects actions in order to achieve some goal . A world model maintains some internal representation of the environment’s true latent state based on the agent’s interaction history. Most world models used for model-based control expose a similar interface to the policy. Specifically, the policy queries the world model with candidate actions (in practice, the policy may query with a sequence of actions). The world model predicts (or samples) the representation of the future state after taking the action: . The predicted internal state need not itself be exposed to the policy. We denote by the readout of the predicted future, which may take many different forms, including a latent representation, natural-language description, symbolic state, or generated observation. The policy uses these readouts to determine which action to execute in the real environment. Prior methods may differ in terms of their internal representation and the response given to the policy, but this simulation interface underlies many world model families, which we show in Fig. 2.
Belief-Based World Models: In a BB-WM (see bottom right of Fig. 2), the internal representation includes an explicit belief over the environment’s current latent state , with denoting any additional information. The BB-WM specifies an interface that allows the belief to be directly exposed to the policy. Specifically, given all past observations and actions , the belief is . Before executing , the prior representation of the next timestep is . After executing and observing the resulting , the posterior representation is . Then, the policy has access to two interfaces.
Simulation interface: For a candidate action , the world model predicts the resulting next state . The policy receives a policy-facing readout: .
Belief-query interface: Unlike current world modeling methods, with a BB-WM the policy can directly query the current posterior belief without supplying a hypothetical action. Specifically, the policy can ask some query over belief through the interface: .
Before taking an action, the policy may issue multiple queries across both interfaces, and the responses can be used as conditioning for the policy to decide the next action.
4 Method
Recall our research question: “if a world model represents a belief over the state, does exposing this belief to the LLM agent improve decision-making?" To isolate the effect of exposing the belief to the agent, we deliberately abstract away the separate challenge of learning an accurate BB-WM by studying simple text-based environments and leveraging prior knowledge about the game environment. This allows us to hand-specify the task-relevant state space, the prior belief prediction , and the posterior belief update , instead of needing to learn them from environment interactions. Because the hand-designed state space is semantic and interpretable, we can define a natural-language belief readout , allowing a pretrained LLM policy to query the BB-WM without finetuning. Our goal is not to propose hand-crafted world models as a scalable solution, but to establish whether explicit belief access is useful in the first place. We describe our benchmark-specific instantiations next.
4.1 ALFWorld
ALFWorld (Shridhar et al., 2021) is a text-based game where an agent completes household instructions in a simulated home. The agent must locate one or more target objects among a set of receptacles (cabinets, drawers, fridge, etc.), manipulate the objects (take, clean, heat, cool, etc), and place them at a specified receptacle. The environment is partially observed: objects are randomly initialized to specific receptacles, and receptacle contents are hidden until the agent navigates to them.
Prompt: An LLM agent is used as the policy. In the initial prompt, we provide environment-specific details, the agentic framework instructions, available actions, the WM query interface, and an in-context learning (ICL) example. The ICL example show the agent an end-to-end example of completing a task from the task-type and involves a WM belief query. Then, the goal task and initial observation are provided. Details and examples are in Appendix A.1.
State Space: The state space contains deterministic components and probabilistic components. Deterministic components include things like agent location and inventory. The only probabilistic component is a belief over object locations. It is represented as a categorical distribution over all receptacles in the environment. More details on the state space construction are in Appendix A.4.
Belief Update: Deterministic components are updated using rule-based logic, parsed from ALFWorld environment observations. The probabilistic component is seeded from a uniform prior over possible receptacles for each object, which is prior knowledge specified by ALFWorld’s game engine (see Appendix A.4.2 for more). Our belief update follows simple presence/absence renormalization: finding an object in a receptacle collapses the belief to a point mass; if not found, we zero that receptacle’s belief and renormalize the belief over the remaining unsearched receptacles.
Simulation interface: We use WALL-E (Zhou et al., 2025), which implements a rule-based mechanism to check if the contemplated action is valid (given the current static state) and responds to the agent in natural language (either valid or invalid with feedback). If WALL-E returns the action is valid, the agent executes it. Else, the agent re-plans the action conditioned on the feedback.
Note that WALL-E is an incomplete next-state predictor. It does not actually predict the complete next state, and it does not represent a belief. Regardless, we intentionally adopt it because it is simple and allows us to test the benefits of the belief-query interface in the BB-WM.
Belief-query interface: The world model’s belief is exposed to the agent through a query action. To access the probabilistic state, the agent can ask where is <object>, and the WM responds with a list of receptacles with non-zero probabilities (ranked by their belief). The agent can also access the static state through other queries. See Appendix A.5 for more details on the interface.
4.2 ScienceWorld
ScienceWorld (Wang et al., 2022) is a text-based game in which an agent performs elementary-science experiments in a fixed ten-room house. We evaluate on 24 test task types (e.g., changing states of matter, growing a plant, mixing chemicals). Relative to ALFWorld, ScienceWorld requires more common-sense reasoning from the agent. Similar to ALFWorld, uncertainty is limited to the location of randomly initialized task-relevant objects. However, we observe that the possible locations of task-relevant objects is far less varied compared to ALFWorld.
We follow ALFWorld’s prompt structure; details and examples are in Appendix B.1. The deterministic components of the world model’s state space are detailed in Appendix B.4.1. The probabilistic component operates similarly to what was done for ALFWorld (see Appendix B.4.2). Rule-based logic is used to update deterministic components of the state space. Similar to ALFWorld, belief over object locations is updated using presence/absence renormalization.
Simulation Interface: Since there is no publicly available implementation of WALL-E for ScienceWorld, we opt for an oracle version of WALL-E. We use the environment to check if the action is valid or invalid. If valid, the action is executed, and if invalid, the agent is given generic feedback and asked to retry (up to the same retry budget used in WALL-E). We refer to this as an oracle because we are guaranteed to always get the valid/invalid action prediction correct.
Belief-query interface: To access the probabilistic state, the agent can ask where is <object>, and the WM responds with a list of rooms ranked by their belief, and the container of the object if known. The agent can also query the deterministic state. See Appendix B.5 for more specifics.
4.3 BabyAI
BabyAI (Chevalier-Boisvert et al., 2019) is a text-based game in a grid world environment, where the agent uses simple navigation commands (e.g., go forward, turn left/right, pick/drop) to complete a pre-specified goal. We evaluate on 4 task types from the BALROG benchmark (Paglieri et al., 2025) (details in Appendix C). Relative to our other benchmarks, BabyAI requires spatial reasoning and lower-level planning, since the agent cannot issue semantic commands like “go to kitchen”. BabyAI adds partial observability through a field-of-view (FoV) mechanism, where only cells in the agent’s FoV are observed. We generally follow ALFWorld’s prompt structure, except there are no ICL examples provided (see Appendix C.1 for more). The deterministic and probabilistic components are detailed in Appendix C.3. Since there is no WALL-E implementation for BabyAI, we create our own (detailed in Appendix C.5). To access the probabilistic state, the agent can ask for an object’s location or a map of the grid. The agent can also query deterministic components. See Appendix C.4 for more on the query interface.
5 Experiments
Our goal is to study whether exposing a belief over the current state to an LLM agent improves its decision-making. We therefore perform a controlled study in which the LLM agent is frozen and variants differ only in the information exposed by the world model. Our objective is not to maximize benchmark performance, but to isolate the relative benefit of belief access. For this reason, we also do not compare to concurrent methods for belief tracking (referenced in Sec. 2).
Methods Evaluated: On each benchmark, we evaluate three LLMs (Llama-3.1-8B-Instruct (Grattafiori et al., 2024), Qwen3-14B (Yang et al., 2025), and Sonnet 4.6) on two agentic frameworks (ReAct (Yao et al., 2023) and ReflAct (Kim et al., 2025)). For each agent, we evaluate several different world models. In Belief, we allow the agent to only ask queries about the current belief state (no simulation). In WALL-E, we pair the agent with WALL-E’s action-conditioned next-state predictor (no belief state, just a valid/invalid action prediction). In BB-WM, we compose the belief modeling with WALL-E, allowing the agent to both do simulation and ask questions on the current belief state.
Evaluation Budget: In ALFWorld/BabyAI, the agent can take up to 30/64 environment actions to achieve the task, following. ScienceWorld varies the step budget per task family. In both benchmarks, we terminate execution early if the agent seems to be derailed; specifically if it issues 10 consecutive invalid actions or 10 consecutive world model queries.
Metrics: An effective agent should complete the most number of tasks in the fewest number of steps. We use reward-within-budget to measure both. The average reward obtain so far is reported as a function of the fraction of steps taken out of the total action budget. ALFWorld’s and BabyAI’s return is binary (0/1) depending on task success and ScienceWorld’s reward is a score from [0,1], allowing partial rewards in intermediate steps for reaching certain milestones.
5.1 Results
We report results on the benchmarks in Fig. 3. We make the following observations:
(1) Adding BB-WM helps the base agent in all cases, both in performance and efficiency. In general, BB-WM is superior to the other variants (WALL-E and Belief).
(2) Enabling belief queries (represented by Belief) and simulation queries (represented by WALL-E) have a complementary effect: they address different shortcomings in the base agent. For example, on ALFWorld the Llama ReAct agent enjoys performance gains under each method individually, which compound when they are combined in the BB-WM. In other cases, WALL-E does not improve the base agent’s performance (observe, Qwen3 ReflAct on ScienceWorld), but when integrated into the BB-WM, the agent outperforms the Belief case.
(3) The LLM benefits from beliefs if the task is hard relative to the LLM’s capabilities. Sonnet, by itself mostly saturates ALFWorld and ScienceWorld. Nevertheless, we notice consistent, small efficiency gains over the base agent when incorporating beliefs. On BabyAI, however, Sonnet has not saturated the benchmark, and adding WALL-E does not make much of an impact, likely because Sonnet does not tend to perform invalid actions. Adding beliefs (either with Belief or BB-WM) bring substantial improvements in terms of efficiency and performance. For easier tasks, an LLM of sufficient scale may be able to implicitly track state (even with uncertainty) and rarely performs invalid actions, which constrains the benefits of any world modeling efforts; for more difficult tasks, an LLM benefits from access to beliefs, and the benefits tend to compound when adding foresight through simulation-based mechanisms.
5.2 Memory vs. Belief
Our world model maintains deterministic information extracted from past observations, together with a belief over uncertain components of the current state. The deterministic component can be viewed as within-episode memory: it records information acquired earlier in the trajectory. Since prior work has shown that explicit memory can improve LLM agents (Yoneda et al., 2024; Hu et al., 2025), a natural alternative explanation is that the gains from Belief arise primarily from externalizing this memory rather than from representing uncertainty.
To evaluate these effects, we introduce a Memory-only ablation on ALFWorld and ScienceWorld for the Llama and Qwen agents. Memory retains the same deterministic state as Belief, but removes the belief over unobserved object locations. Accordingly, queries about the observed state are unchanged, while a where is <object> query for an unobserved object only reports that the object has not yet been observed. Thus, Memory provides access to what the agent has already seen from the trajectory, whereas Belief additionally represents what remains possible about the unobserved state.
Results are shown in Fig. 4. Memory provides gains over the base agent in some settings, re-affirming that state tracking is useful. However, Belief consistently improves over Memory. These results indicate that the gains from belief access cannot be explained by deterministic memory alone: a belief over uncertain components of the state provides additional decision-relevant information.
5.3 Non-Uniform Priors
In our main experiments, the belief over an object’s initial location is uniform over its valid receptacles, reflecting the initialization logic of the underlying benchmarks. In this setting, knowing the belief support—which receptacles remain possible—captures most of the useful information, since these locations are equally likely. In more realistic partially observed environments, however, some states may be substantially more likely than others. We therefore ask whether an LLM agent benefits from access to the probabilities of a belief, beyond simply knowing its support.
To study this question, we modify ALFWorld’s object-initialization process. For each object type, we define a Zipf (or zeta) distribution () over its valid receptacles, which we use to sample the task-relevant object’s initial location. We use a Zipf distribution because it introduces a simple heavy-tailed ranking over otherwise-valid locations: a small number of receptacles are substantially more likely, while all valid receptacles retain non-zero probability. The BB-WM is given the corresponding prior and updates it as observations eliminate possible locations. We additionally introduce Belief-NoProb to isolate the value of explicit probability information. This variant is identical to Belief, except that the world model omits numerical probabilities when answering location queries and randomizes the location order in the response to remove ordinal information. Thus, the agent still knows which receptacles are possible, but does not receive the probability mass assigned to each one. Results are shown in Fig. 5.
(1) Non-uniform initialization makes the base agent less effective. Although the tasks, environments, and set of valid object locations are unchanged, altering the initialization distribution reduces the performance of the base agent. Providing access to the belief largely recovers the performance observed under the original initialization process (in Sec. 5.1).
(2) Probability information improves search efficiency. Belief and Belief-NoProb reach similar performance by the end of the action budget, but Belief achieves substantially higher reward earlier in the trajectory. Thus, the benefit of belief access is not limited to identifying which states remain possible: the agent uses the probability assigned to those states to prioritize more likely locations and solve tasks using fewer actions.
5.4 Qualitative Analysis
BB-WMs help the agent generalize to challenging test-time situations. We see this with an ALFWorld task that the Sonnet ReAct agent failed with WALL-E but completed with BB-WM. The task is to heat some egg and put in garbagecan. The main difficulty in this task is that the egg gets randomly initialized to the garbagecan, which likely represents a novel setting for Sonnet (in the real world, eggs are not commonly found in the garbagecan, especially those that need to be heated). In the WALL-E trajectory (see Appendix D), the agent exhausts the action budget by checking many receptacles (fridge, countertops, shelves, cabinets, drawers, stoveburners). Critically, the ALFWorld game engine will never initialize an egg into the violet-colored receptacles, but WALL-E cannot expose this world knowledge to the LLM agent. When coupled with the BB-WM (see Appendix D), the agent starts by querying for possible egg locations, and the BB-WM, due to its prior, only responds with valid initial locations. The agent checks the more sensible receptacles first (fridge, countertop, microwave), then re-queries the BB-WM for an updated receptacle list, and finally checks the garbagecan to find the egg, heat it, and complete the task.
6 Conclusion
This work shows that belief access improves LLM agent decision-making and often compounds with simulation-based world modeling, motivating Belief-Based World Models (BB-WMs), where current-state belief and future-state prediction provide distinct and complementary information to an agent. This work has two important limitations. First, our BB-WMs are hand-crafted and benchmark-specific. Thus, our results establish that accurate beliefs are useful when exposed to an LLM policy but do not address how such beliefs should be learned in realistic, open-ended settings. Second, our simulation interface is based on WALL-E-style action-validity prediction. This is a relatively limited form of simulation and does not propagate the current belief forward into a belief over hypothetical future states. Both limitations are important directions for future work, and our results indicate that addressing them could yield meaningful benefits in agent decision-making under partial observability.
AI use statement
In this work, we used generative AI tools for creating or modifying scientific figures or images, creating or editing software code, summarizing or analyzing existing literature, sourcing/searching for information, editing a research paper to improve readability, and identifying relevant literature.
We have not used generative AI tools to help develop theoretical models or conceptual frameworks, propose or refine hypotheses, design or provide feedback on research methodology or experiments, implement methods, support qualitative and thematic data analysis, interpret results. The following are not applicable to this work: generate synthetic data sets, formulate mathematical claims, provide critical ingredients for proving mathematical claims, assist in the writing of proofs, assist with translation, clean and reformat dataset.
We have reviewed all AI-assisted work. Figures and images were checked for correctness and readability. AI-generated code was tested for correctness. Sources raised in literature reviews were manually reviewed and read to obtain a more complete understanding. We take responsibility for the final content of this work, including text, claims or artifacts produced with the aid of generative AI.
References
- Cosmos world foundation model platform for physical ai. ArXiv abs/2501.03575. External Links: Link Cited by: §1.
- Do as i can, not as i say: grounding language in robotic affordances. In Conference on Robot Learning, External Links: Link Cited by: §1.
- Genie: generative interactive environments. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. Cited by: §1, Figure 2.
- Web agents with world models: learning and leveraging environment dynamics in web navigation. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §2.
- Flow-based recurrent belief state learning for POMDPs. In Proceedings of the 39th International Conference on Machine Learning, K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato (Eds.), Proceedings of Machine Learning Research, Vol. 162, pp. 3444–3468. External Links: Link Cited by: §1.
- BabyAI: first steps towards grounded language learning with a human in the loop. In International Conference on Learning Representations, External Links: Link Cited by: §4.3.
- CWM: an open-weights llm for research on code generation with world models. ArXiv abs/2510.02387. External Links: Link Cited by: §1.
- WorkArena: how capable are web agents at solving common knowledge work tasks?. External Links: 2403.07718 Cited by: §1.
- The llama 3 herd of models. External Links: 2407.21783, Link Cited by: §5.
- Is your llm secretly a world model of the internet? model-based planning for web agents. Transactions on Machine Learning Research. Cited by: Figure 2.
- Recurrent world models facilitate policy evolution. In Advances in Neural Information Processing Systems 31, pp. 2451–2463. Note: https://worldmodels.github.io External Links: Link Cited by: §1.
- Dream to control: learning behaviors by latent imagination. In International Conference on Learning Representations, External Links: Link Cited by: §1, §1, Figure 2.
- Learning latent dynamics for planning from pixels. arXiv preprint arXiv:1811.04551. Cited by: Figure 2.
- DreamPhase: offline imagination and uncertainty-guided planning for large-language-model agents. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §2.
- Reasoning with language model is planning with world model. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 8154–8173. External Links: Link, Document Cited by: §1, Figure 2.
- GAIA-1: a generative world model for autonomous driving. ArXiv abs/2309.17080. External Links: Link Cited by: Figure 2.
- HiAgent: hierarchical working memory management for solving long-horizon agent tasks with large language model. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 32779–32798. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §5.2.
- Language models as zero-shot planners: extracting actionable knowledge for embodied agents. In Proceedings of the 39th International Conference on Machine Learning, K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato (Eds.), Proceedings of Machine Learning Research, Vol. 162, pp. 9118–9147. External Links: Link Cited by: §1.
- Deep variational reinforcement learning for POMDPs. In Proceedings of the 35th International Conference on Machine Learning, J. Dy and A. Krause (Eds.), Proceedings of Machine Learning Research, Vol. 80, pp. 2117–2126. External Links: Link Cited by: §1.
- Planning with diffusion for flexible behavior synthesis. In International Conference on Machine Learning, Cited by: §1.
- Think locally, explain globally: graph-guided llm investigations via local reasoning and belief propagation. ArXiv abs/2601.17915. External Links: Link Cited by: §1.
- Planning and acting in partially observable stochastic domains. Artificial Intelligence 101 (1-2), pp. 99–134. Cited by: §1, §1.
- ReflAct: world-grounded decision making in LLM agents via goal-state reflection. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 33433–33465. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §5.
- QuBE: question-based belief enhancement for agentic LLM reasoning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 21403–21423. External Links: Link, Document Cited by: §2.
- On distinguishing epistemic from pragmatic action. Cognitive Science 18 (4), pp. 513–549. External Links: Document Cited by: §1.
- Belief memory: agent memory under partial observability. ArXiv abs/2605.05583. External Links: Link Cited by: §2.
- ABBEL: learning natural-language belief states for memory-efficient interaction. External Links: Link Cited by: §2.
- LeWorldModel: stable end-to-end joint-embedding predictive architecture from pixels. ArXiv abs/2603.19312. External Links: Link Cited by: §1, Figure 2.
- BALROG: benchmarking agentic LLM and VLM reasoning on games. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §4.3.
- Current agents fail to leverage world model as tool for foresight. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 13686–13723. External Links: Link, Document, ISBN 979-8-89176-390-6 Cited by: §2.
- Debugging code world models. ArXiv abs/2602.07672. External Links: Link Cited by: §1.
- Cosmos-drive-dreams: scalable synthetic driving data generation with world foundation models. ArXiv abs/2506.09042. External Links: Link Cited by: §1.
- StateAct: state tracking and reasoning for acting and planning with large language models. External Links: 2410.02810, Link Cited by: §2.
- Mastering atari, go, chess and shogi by planning with a learned model. Nature 588 (7839), pp. 604–609. Cited by: §1.
- Data-efficient reinforcement learning with self-predictive representations. In ICLR, External Links: Link Cited by: §1.
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning. In Proceedings of the International Conference on Learning Representations (ICLR), External Links: Link Cited by: §4.1.
- Structured world belief for reinforcement learning in pomdp. In Proceedings of the 38th International Conference on Machine Learning, M. Meila and T. Zhang (Eds.), Proceedings of Machine Learning Research, Vol. 139, pp. 9744–9755. External Links: Link Cited by: §1.
- Agent-brace: decoupling beliefs from actions in long-horizon tasks via verbalized state uncertainty. External Links: 2605.11436, Link Cited by: §2.
- Learning belief representations for partially observable deep RL. In Proceedings of the 40th International Conference on Machine Learning, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett (Eds.), Proceedings of Machine Learning Research, Vol. 202, pp. 35970–35988. External Links: Link Cited by: §1.
- ScienceWorld: is your agent smarter than a 5th grader?. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y. Goldberg, Z. Kozareva, and Y. Zhang (Eds.), Abu Dhabi, United Arab Emirates, pp. 11279–11298. External Links: Link, Document Cited by: §4.2.
- MPO: boosting LLM agents with meta plan optimization. In Findings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 3914–3935. External Links: Link, Document, ISBN 979-8-89176-335-7 Cited by: §A.3.
- Qwen3 technical report. External Links: 2505.09388, Link Cited by: §5.
- SWE-agent: agent-computer interfaces enable automated software engineering. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §1.
- ReAct: synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR), Cited by: §5.
- Statler: state-maintaining language models for embodied reasoning. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Cited by: §2, §5.2.
- Reinforcement world model learning for llm-based agents. ArXiv abs/2602.05842. External Links: Link Cited by: §1.
- DINO-wm: world models on pre-trained visual features enable zero-shot planning. External Links: 2411.04983, Link Cited by: Figure 2.
- WALL-e 2.0: world alignment by neurosymbolic learning improves world model-based llm agents. arXiv preprint arXiv:2504.15785. Cited by: §1, Figure 2, §2, §4.1.
- Reducing belief deviation in reinforcement learning for active reasoning of LLM agents. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §1.
Appendix A BB-WM Instantiation for ALFWorld
A.1 General Agent Prompt
We provide the agent prompt used for ALFWorld below. We separate the prompt into benchmark instructions, agent instructions, the available action space, the in-context example, and the task-specific instruction. For brevity, we truncate the in-context example, and add color styling to help the reader navigate the text.
Benchmark Instructions
Agent Instructions
We specifically show ReflAct’s agent instructions here.
Actions
In-Context Example
Task Instruction
A.2 World Model Prompt Instructions
WALL-E World Model Instructions
Belief-Based World Model Instructions
Below are the instructions used for the Belief variant. For BB-WM, we compose both the WALL-E and these instructions.
A.3 In-Context Examples
We use per-task in-context examples, following (Xiong et al., 2025). They are publicly available in this repo: https://github.com/WeiminXiong/MPO
For the Belief and BB-WM variants, we alter the in-context example to show the agent using the world model to query for the location of the task’s target object. Nothing else in the example is changed (following action selection or agentic reasoning thoughts).
A.4 State Space
A.4.1 Deterministic State Space
The deterministic component of the world-model state maintains the task goal, receptacle and object attributes observed so far, and the agent’s current location and inventory. We represent the state at timestep as
with the following components.
Goal.
The task goal is fixed for an episode and consists of
where act is one of {pick_and_place, clean, heat, cool, look, picktwo}. The transform and destination fields are optional depending on the task type.
Receptacles.
contains one state entry for each receptacle , with all receptacle identifiers known at the beginning of the episode:
Here, open records the receptacle’s open/closed state, searched indicates whether its contents have been observed, and contents is the set of observed object identifiers contained in it.
Objects.
contains entries for object instances that have been observed in the environment:
An object’s location is a receptacle identifier, inventory, or unknown. The remaining Boolean attributes record whether the object has been cleaned, heated, or cooled.
Agent.
The agent state consists of its current location and inventory , where is the current receptacle/location and is the set of object identifiers currently held by the agent.
A.4.2 Probabilistic State Space
The probabilistic component represents uncertainty over the locations of task-relevant objects. It contains only categorical distributions over receptacle locations
Let
denote the set of receptacle instances in the scene, all of which are known at the beginning of the episode. We represent the belief over the target object’s location as
where
is the probability that the target object is located at receptacle . There is no additional “unknown” or undiscovered-location category; probability mass is distributed only over receptacles in .
Target Belief.
The belief maintains a categorical distribution over receptacle identifiers, together with an optional concrete object identifier and observed location:
Before the target object is observed, instance_id and located_at are undefined. Once observed, they identify the concrete object instance and its known location, and the distribution becomes degenerate at that receptacle.
Only the task-relevant target object type receives a persistent probabilistic belief. For tasks requiring multiple objects of the same type, this representation is extended with a separate target belief for each required instance.
Prior.
The initial belief is derived from ALFWorld’s object–receptacle placement constraints. Let
The prior is uniform over compatible receptacle instances present in the scene:
If no placement information is available for the object type, or no compatible receptacle type appears in the scene, we instead initialize a uniform distribution over all .
A.5 Belief-Query Interface
The policy accesses the world-model state through an explicit query interface. Queries use the same action channel as environment actions: instead of issuing an environment action, the agent outputs
The world model answers the query from its current state and returns the result as the next Observation:. Queries do not modify the environment or consume an environment step. The interface exposes four query types.
query where is {obj}.
This query exposes the location belief described in Appendix A.4.2. For the task-relevant target object, the world model ranks unsearched receptacles with positive belief mass under and returns the most likely locations and their probabilities, e.g.,
Locations are returned in decreasing probability order until their cumulative probability reaches at least ; any remaining probability mass is summarized as and other receptacles (X%).
For non-target object types, which do not have a persistent belief in the probabilistic state, the world model instead constructs the corresponding placement prior over currently unsearched receptacles at query time. If a specific observed object instance is requested, the response is obtained directly from the deterministic object state, reporting whether the object is held, located at a particular receptacle, or has not yet been located.
query what is in {recep}.
This query accesses the deterministic receptacle state. If the receptacle has been searched, the world model returns its observed contents; otherwise, it reports that the receptacle has not yet been searched. For example,
query searched {recep}.
This query returns whether the specified receptacle has already been searched, using the searched field of the deterministic state:
or
query state.
This query returns a compact summary of the deterministic state described in Appendix A.4.1, including the agent’s current location and inventory, observed objects, and searched receptacles:
Appendix B BB-WM Instantiation for ScienceWorld
Our ScienceWorld benchmark uses all available environment simplifications, detailed on the repo: https://github.com/allenai/scienceworld
B.1 General Agent Prompt
We provide the agent prompt used for ScienceWorld below. We separate the prompt into benchmark instructions, agent instructions, the available action space, the in-context example, and the task-specific instruction. We truncate the in-context example for brevity and add color styling to help the reader navigate the text.
Benchmark Instructions
Agent Instructions
Actions
In-Context Example
Task Instruction
B.2 World Model Prompt Instructions
WALL-E Oracle World Model Instructions
Belief-Based World Model Instructions
B.3 In-Context Examples
Following our process for ALFWorld, we create per-task in-context examples. We generate in-context examples by using the environments gold path generator. We distill the gold path action sequence by dropping redundant actions, and we verify on a fresh run that the distilled action sequence succeeds in solving the task. Reasoning thoughts (for use in ReAct and ReflAct) frameworks are written by an agent.
B.4 State Space
B.4.1 Deterministic State Space
The deterministic component of the ScienceWorld state maintains the task goal, room contents, object nesting, and the agent’s current location, inventory, and focused object. All quantities in this component represent observed or task-specified facts. We represent the deterministic state at timestep as
with the following components.
Goal.
The task goal is fixed for an episode and contains
where target_type identifies the task-relevant object (e.g., orange juice or aluminum foil), and named_location optionally specifies a room in which the task description explicitly states that the target is located.
Rooms.
ScienceWorld uses a fixed set of ten rooms,
For each room , the deterministic state maintains
where searched indicates whether the room has been observed via the look around action, and contents is the set of object referents observed in that room.
Unlike ALFWorld, ScienceWorld objects are represented directly by potentially multi-word referents such as orange juice, red apple, or metal pot, rather than numbered object identifiers.
Object Nesting.
We record observed containment relationships through
where is the set of objects observed inside or on container . An inverse object-location map
records the room and, when applicable, container associated with an observed object.
For held containers, the state additionally maintains their known contents:
Agent.
The agent state consists of its current room , inventory , and focused object :
Here, is the agent’s current room, is the set of object referents currently held, and is the object most recently selected by a focus on action.
B.4.2 Probabilistic State Space
The probabilistic component of the ScienceWorld world model represents uncertainty over the room containing an object. Let denote the fixed set of ten ScienceWorld rooms. For a tracked object type , we maintain a categorical belief
where
is the probability that an instance of object type is located in room .
Each object belief additionally records a resolved location,
where dist is the categorical distribution over rooms and located_at is initially undefined. Once the object is observed, located_at records its room; if the object is held by the agent, it takes the value inventory. The target object specified by the task is tracked from the beginning of the episode. Additional object types may be instantiated as needed and subsequently follow the same belief representation and update rules.
Prior.
The initial belief is derived from ScienceWorld’s object–room placement constraints. Let
denote the set of rooms in which object type may occur. In the absence of additional task information, the prior is uniform over these candidate rooms:
If no placement information is available for an object type, the prior is uniform over all rooms in .
Task descriptions may provide additional location information. If the task explicitly specifies the target object’s room , the target belief is initialized deterministically:
For other objects whose locations are stated in the task description, we place probability on the stated room and distribute the remaining uniformly over the object’s other candidate rooms.
B.5 Belief-Query Interface
The policy accesses the ScienceWorld world-model state through an explicit query interface. Queries use the same action channel as environment actions: instead of issuing an environment action, the agent outputs
The world model answers from its current state and returns the result as the next Observation:. Queries do not modify the ScienceWorld environment or consume an environment step. The interface exposes four query types.
query where is OBJ.
This query exposes the room-level location belief described in Appendix B.4.2. For an unresolved object with a maintained belief , the world model ranks rooms with positive belief mass that have not yet been searched and returns the most likely locations and their probabilities, e.g.,
Rooms are returned in decreasing probability order until their cumulative probability reaches at least ; any remaining mass is summarized as and other rooms (X%).
Once an object has been observed, the query may instead resolve its location from the deterministic state described in Appendix B.4.1. In particular, if the object has been observed inside a container, the response includes both the room and container:
If the object is held by the agent, its location is reported as inventory with probability . If all candidate rooms have been searched without finding the object, the world model reports that no candidate rooms remain.
query what is in X.
This query accesses the deterministic room and container contents maintained by the world model. If is a searched room, the query returns the objects observed in that room. If is an observed container, it instead returns the objects recorded inside that container. For example,
If the requested room or container has not yet been observed, the world model reports that it has not been searched.
query searched ROOM.
This query reports whether a look around observation has been obtained for the specified room, corresponding to the searched field in the deterministic room state:
or
query state.
This query returns a compact summary of the deterministic agent state described in Appendix B.4.1, including the current room, inventory, focused object, and searched rooms:
When the agent holds a container, its known contents are included in the inventory summary, e.g., metal pot (containing orange juice).
Appendix C BB-WM Instantiation for BabyAI
We evaluate on the following task families from BALROG: goto, pickup, pick_up_seq_go_to, putnext. We omit open, as it requires a different environment configuration than the other four tasks.
C.1 General Agent Prompt
We provide the agent prompt used for BabyAI below. We separate the prompt into benchmark instructions, agent instructions, the available action space, tips, and the task-specific instruction. There is no in-context example.
Benchmark Instructions
Agent Instructions
We specifically show ReflAct’s agent instructions here.
Actions
Tips
Task Instruction
C.2 World Model Prompt Instructions
WALL-E World Model Instructions
When WALL-E is used alone, the prompt also includes the following tip:
Belief-Based World Model Instructions
Below are the instructions used for the Belief variant. For BB-WM, we compose both the WALL-E and these instructions. In that composition the belief paragraph begins “The same world model also tracks a 6x6 room prior and where unseen objects are likely to be.”
C.3 State Space
C.3.1 Deterministic State Space
The deterministic component is a partial map built from the text observations. The world model does not use MiniGrid’s coordinates. It keeps its own grid, fixed at the start of the episode: the cell where the agent begins is , and the direction the agent faces at initialization is (start-forward). Later turns and steps are recorded in that same grid. An observation such as “a red ball 1 step forward” is converted from the agent’s current position and heading into one cell of this grid.
The map starts with only that origin cell. A cell is added when an observation names it, or when a successful go forward moves the agent onto it. A line “a wall steps forward” adds a wall cells along the current heading and marks the cells in between as empty. An object line adds that object at the named offset. Cells that have neither been described nor stepped on are left out of the map. We represent the state at timestep as
with the following components.
Goal.
The mission is fixed for an episode and consists of
where act is one of {goto, pickup, putnext, seq}. Each referent is a color together with a type in {key, ball, box}. seq_order is then or after for pick-up-then-go-to missions, and is otherwise absent.
Pose.
The agent pose in this grid is
with heading named start-forward, start-right, start-back, and start-left, in clockwise order from the initial facing. turn left and turn right update directly. go forward updates only when the new observation differs from the previous one and the cell in front is not already known to be blocked. An unchanged observation is treated as a failed step, so the agent stays in place.
Cells.
maps coordinates in this grid to the cells recorded so far. A cell has
where kind is one of {empty, wall, object}. visited is true for cells the agent has occupied, including the origin.
Objects.
is indexed by color and type, such as green key; there are no instance id’s. It records every object that appears in the observations, and it’s location (a cell in the grid or inventory).
Inventory.
is the carried object, a color–type pair parsed from observations (e.g., “You carry …”), or empty.
C.3.2 Probabilistic State Space
The probabilistic component is a uniform placement prior over BabyAI’s walkable room. Since the starting orientation of the agent is not known, the room cannot be localized, or pinned, until at least two adjacent walls are observed.
Let be a mission referent. We represent its location belief as
where the atoms are UNSEEN before localization and the remaining candidate cells afterward. A sequential mission keeps one belief for each referent. Referents that share a color and type share one belief. Objects that are not in the mission have no persistent belief; queries about them use .
Target Belief.
Each belief stores the categorical distributiontogether with an optional located atom:
Before the referent is observed, located_at is undefined. After being observed, location becomes that cell, and on pickup it becomes inventory.
Prior.
Before the room is localized, the belief is a point mass on the unseen atom,
Once the room is pinned, that mass is replaced by a uniform distribution over interior cells the generator can still occupy. Let be the interior cells that are not the start cell, every cell within Manhattan distance of the start, the agent’s current cell, and any cell already known to be empty, a wall, or occupied by a different object. Then
Later observations remove ruled-out cells and renormalize. The first sighting replaces the prior with a point mass, as above.
C.4 Belief-Query Interface
The policy accesses the world-model state through an explicit query interface. Queries use the same action channel as environment actions: instead of issuing an environment action, the agent outputs
The world model answers from its current state and returns the result as the next Observation:. Queries do not modify the environment or consume an environment step. Responses are egocentric (“1 step forward”, “2 steps back and 1 step right”, “front”). The interface exposes six query types.
query where is {obj}.
If the object is carried, the answer is taken from . If it has been observed, the answer is the last-known atom in , rendered relative to the current pose, e.g.,
If it has not been seen and the room is not yet localized, the world model asks for a wall ahead and a wall to one side:
Once the room is localized, the answer is the placement prior of Appendix C.3.2: the number of remaining cells, the nearest cells, and which side of the agent still holds mass, e.g.,
query map.
This query prints the pinned room. "?" marks a cell that still has prior mass, "." a cell that has been ruled out, and a letter marks a seen object. Before both wall axes are known, the answer states that the room is not localized yet.
query what is at {place}.
This query reads at the cell corresponding to the egocentric place. If that cell has been written, the world model returns its fact; otherwise it reports whether the cell is still a candidate, ruled out, or not yet observed. For example,
query searched {place}.
A place counts as searched when its cell is present in :
or
query what have I seen.
This query lists referent keys in with their last-known places, e.g.,
query state.
This query returns the inventory and the same map as query map:
C.5 WALL-E Implementation
Our hand-crafted implementation of WALL-E checks a proposed action against the environment’s state before it is executed. If the action would fail, WALL-E does not step the environment. The agent instead receives some feedback and asked to reconsider the action. Any case not handled below is left to the environment, which if invalid, returns Nothing happens. That is not a valid action. and does count as a step.
go forward.
The step is rejected when the cell in front blocks movement. On these tasks that cell is a wall or an object.
pick up.
Pickup uses the cell in front of the agent. It is rejected when that cell is empty, when it cannot be carried, or when the agent is already holding something.
drop.
Drop is rejected when the agent is holding nothing, or when the cell in front is already occupied.
toggle.
Toggle is rejected only when the cell in front is empty. A toggle aimed at a wall or an object is executed.
Movement aliases.
go back, go backward, and go left or go right (including move, walk, and step, with or without a step count) are not MiniGrid actions. WALL-E rejects them and names the legal verbs. The same template is used for left and right, with that direction in place of backward.
Appendix D Example ALFWorld Trajectories
We provide an example ALFWorld trajectory (task #10 from the unseen split) from Sonnet+ReAct using WALL-E and BB-WM. For readability, we color-code agent thoughts, actions, and environment observations, and we highlight in yellow critical parts of the trajectory.
WALL-E
We provide an example ALFWorld trajectory (task #10 from the unseen split) from Sonnet+ReAct using WALL-E. For readability, we color-code agent thoughts, actions, and environment observations, and we highlight in yellow critical parts of the trajectory, which show why the agent misses searching the garbagecan for the egg.
The trajectory terminates because the agent hits the task’s action budget.
BB-WM
Appendix E Example ScienceWorld Trajectories
We provide an example ScienceWorld trajectory (grow plant task, #93 from the unseen split) from Qwen3-14B+ReflAct using WALL-E Oracle and BB-WM. For readability, we color-code agent thoughts, actions, and environment observations, and we highlight in yellow critical parts of the trajectory.
WALL-E Oracle
BB-WM

