TreeSeeker: Tree-Structured Trial, Error, and Return in Deep Search
Abstract
Deep search requires agents to answer complex questions through multi-step web search, browsing, evidence comparison, and synthesis. A central challenge is deciding how to search when several directions look plausible but only some will later lead to reliable evidence. If an agent greedily follows the current best-looking direction, it may keep extending a weak continuation. If it explores without discipline, it may waste budget on disconnected trials. We propose TreeSeeker, an inference-time framework for controlled trial-and-error in deep search. TreeSeeker organizes search as branch-and-return search over tree-structured states, where each branch is a tentative direction for a sub-goal. At each round, TreeSearch reads all sub-goal trees, identifies active goals, and uses textual UCB signals of value, uncertainty, and risk to select among exploiting a promising branch, exploring an uncertain alternative, or pruning an unproductive continuation and returning to an earlier branch point. TreeMem supports this control loop by keeping evidence, uncertainty, conflicts, progress, and failure cues attached to the branches that produced them, so trial outcomes can guide later decisions. Experiments on XBench-DeepSearch, BrowseComp, and BrowseComp-ZH show that TreeSeeker consistently outperforms strong open-source baselines, suggesting that explicit branch-and-return control complements stronger reasoning and tool execution.
1 Introduction
Deep search is a long-horizon information-seeking task where an agent answers complex questions by interacting with web information sources over multiple steps. Unlike closed-book question answering or single-step retrieval, deep search requires the agent to plan subgoals (Nie et al., 2026; Qin et al., 2026), issue search queries, browse webpages (Wu et al., 2025; Li et al., 2025c), inspect evidence (Chen et al., 2026a; Ye et al., 2025), resolve conflicts, and synthesize a final answer. LLM-based agents are a natural fit for this setting because they can combine reasoning with tool use, as shown by ReAct-style agents (Yao et al., 2023) that interleave reasoning and actions. As the search horizon grows, however, the main challenge is not only choosing the next action, but deciding when to try a new direction, continue a useful one, or return from a failed one.
Recent work has improved deep-search agents in several ways. Agentic training and post-training (Team et al., 2025b; Lu et al., 2025; Li et al., 2025a; Wu et al., 2025) help models plan, search, browse, and synthesize evidence over long trajectories. Context and memory methods reduce the burden of long histories by compressing past interactions (Wu et al., 2026) or reconstructing an evolving workspace across research rounds (Chen et al., 2026a; Ye et al., 2025). Inference-time execution methods improve efficiency by decomposing tasks into structured subgoals and executing independent parts in parallel (Nie et al., 2026; Qin et al., 2026). These methods make agents stronger and more efficient, but they do not fully address how an agent should control search when multiple directions look plausible. The agent still needs to decide which direction to continue, which alternative to try, and when to stop a weak continuation and return to an earlier branch point.
This control problem appears because early search directions are often uncertain. A webpage, source family, query formulation, or intermediate hypothesis may look useful at first but later lead to weak, conflicting, or incomplete evidence. If the agent greedily follows the current best-looking direction, it may keep extending a weak continuation. If it tries alternatives without a clear rule, it may waste budget on disconnected attempts. Effective deep search therefore needs controlled trial-and-error, where the agent can continue useful directions, test uncertain alternatives, and return from unproductive continuations under a limited budget.
We propose TreeSeeker, an inference-time framework that organizes deep search as branch-and-return search. Each branch represents a search direction, such as a query, source family, or hypothesis, and each tree maintains the branches for a sub-goal. At each decision round, TreeSeeker reads all sub-goal trees, identifies dependency-ready unresolved goals, and outputs one operation for each active goal in a single pass. Its controller, TreeSearch, uses an operation-level textual UCB rule to allocate search budget across exploiting a promising branch, exploring an uncertain alternative, or pruning an unhelpful continuation and returning to an earlier branch point. The rule compares three ordinal semantic signals, value, uncertainty, and risk, and extends the explore–exploit principle with a risk-aware return mechanism. TreeMem supports this control loop by keeping evidence, uncertainty, conflicts, progress, and failure cues attached to the branches that produced them. Together, TreeSearch and TreeMem allow the agent to compare tentative directions, continue useful ones, test uncertain alternatives, and recover from unhelpful search paths without repeatedly extending a single linear trajectory.
We evaluate TreeSeeker on XBench-DeepSearch (Chen et al., 2025), BrowseComp (Wei et al., 2025), and BrowseComp-ZH (Zhou et al., 2025). TreeSeeker achieves 56.3 on XBench-DS, 47.0 on BrowseComp, and 43.0 on BrowseComp-ZH, achieving the best performance among the evaluated open-source baselines. Ablations on XBench-DS show that removing textual UCB signals and disabling the branch-and-return operations reduce performance by 4.3 and 8.3 points, respectively, confirming that both operation-level scoring and branch-and-return control are necessary.
In summary, our contributions are:
- •
We propose TreeSeeker, a branch-and-return framework that organizes deep search over tree-structured states. At each decision round, it reads all sub-goal trees and decides one operation for each active sub-goal.
- •
We introduce TreeSearch and TreeMem as the two core components of TreeSeeker. TreeSearch performs operation-level textual UCB control using value, uncertainty, and risk signals, while TreeMem keeps evidence, uncertainty, conflicts, progress, and failure cues attached to their corresponding branches.
- •
We evaluate TreeSeeker on XBench-DeepSearch, BrowseComp, and BrowseComp-ZH, with ablations and cost analysis to study the effects of branch-and-return control, textual UCB-style decision making, and structured branch memory.
2 Related Work
2.1 Single-Path deep search Agents and the Premature Commitment Problem
A dominant paradigm in deep-search agent design is the sequential reason-act-observe loop, exemplified by ReAct (Yao et al., 2023), where the agent maintains a single evolving trajectory of plans, tool calls, and observations. IterResearch (Chen et al., 2026a), WebSailor (Li et al., 2025a), WebDancer (Wu et al., 2025), and WebtThinker (Li et al., 2025c) strengthen this paradigm through workspace reconstruction, improved web reasoning, or long-horizon training and interaction scaling. Flash-Searcher (Qin et al., 2026) and ParallelResearch (Nie et al., 2026) further introduce DAG-structured subgoals and parallel execution to improve throughput and coverage. Search More, Think Less (Chen et al., 2026c) takes a complementary direction by replacing deep per-step reasoning with wider evidence acquisition to improve efficiency and generalization. However, these systems still do not provide evidence-driven control for reallocating budget across alternative paths within a goal.
This limitation matters under early-stage uncertainty, a single-path or fixed-schedule agent may continue building on weak evidence rather than redirecting budget based on what has been learned mid-search. TreeSeeker differs by treating each candidate path as a persistent semantic branch state and using TreeSearch to decide which branch to deepen, explore, or prune based on accumulated evidence and failure signals.
2.2 Context Management and Flat History
Long-horizon agents also face a context bottleneck, motivating summarization and workspace-reconstruction methods. ReSum (Wu et al., 2026) periodically compresses exploration history, IterResearch (Chen et al., 2026a) reconstructs an evolving workspace across research rounds, MemAgent (Yu et al., 2026) rebuilds task state on demand, and AgentFold (Ye et al., 2025) condenses and reorganizes search history to prevent context overload.
These methods address the context bottleneck effectively, but they organize history as a single evolving state. When multiple tentative search directions are mixed into a flat context or a single reconstructed workspace, it becomes difficult to distinguish which attempt was useful, which failed, and which alternative branches remain worth pursuing. TreeSeeker uses a different role for summarization: rather than compressing history into one state, TreeMem keeps summarized evidence attached to the branch that produced it. The resulting branch-local states are not memory devices for continued single-path reasoning, but decision objects that TreeSearch can compare, deepen, or prune.
2.3 Trial-and-Error, UCB, and Tree Search for Language Agents
The explore–exploit tradeoff formalized by UCB (Auer et al., 2002) provides a principled basis for allocating effort among competing alternatives under uncertainty. In the language agent setting, tree-search methods have applied this intuition to reasoning and planning: Language Agent Tree Search (LATS) (Zhou et al., 2024) unifies reasoning, acting, and planning in an MCTS-style framework, and follow-up works such as Plan-MCTS (Zhang et al., 2026b), ExACT (Yu et al., 2025), and uncertainty-aware web-agent frameworks (Zhang et al., 2026a) apply tree-based exploration to web navigation or interactive decision settings. In retrieval-augmented generation, MCTS-RAG (Hu et al., 2025) structures retrieval as tree search, Air-RAG (Feng et al., 2025) interleaves retrieval with reasoning via expansion and simulation, and ReARTeR (Sun et al., 2025) uses MCTS-guided process rewards for preference optimization.
These methods operate over well-defined action spaces such as query reformulations, document selections, or reasoning steps, where branch quality can be estimated with scalar rollout rewards or retrieval scores. Deep search presents a different challenge: the objects being compared are partially synthesized semantic evidence states, not discrete actions, and quality signals include partial answers, source conflicts, unresolved constraints, and failure cues rather than numeric rewards. TreeSeeker extends the UCB intuition to this setting through a textual scoring rule that estimates branch value, residual uncertainty, and search risk from semantic branch states, and couples it with a return mechanism that prunes unproductive continuations and redirects search to earlier branch points. This is broadly related to test-time hypothesis exploration in inductive reasoning (Chen et al., 2026b), but our setting differs in that hypotheses are embedded in open-ended web-search branches whose evidence must be actively gathered and revised.
3 Method
3.1 Branch-and-Return Search
We present TreeSeeker, an inference-time framework that organizes deep search as branch-and-return search. As shown in Figure 1, given a root question, TreeSeeker first decomposes it into sub-goals, since complex tasks often require resolving multiple pieces of evidence before synthesizing the final answer. Each sub-goal is associated with a search tree. The first-level branches are candidate paths for solving the sub-goal, such as different queries, sources, or hypotheses, while deeper nodes record the actions and observations produced along each path. Instead of following one linear search trajectory, TreeSeeker keeps these possible directions separated in the trees. The agent can continue a useful branch, try an uncertain alternative, or prune an unhelpful continuation and return to an earlier branch point for revision. In this way, trial-and-error is not repeated independent runs, but a structured process over search branches.
TreeSeeker has two components. TreeSearch decides how the search should move through the trees. At each decision round, it first reads all sub-goal trees in one shot, identifies dependency-ready unresolved goals, and outputs one operation for each active goal in a single pass. Here, operation denotes the TreeSearch decision (Explore/Exploit/Prune). Action denotes the tool action used to execute the selected operation. TreeMem stores compact, failure-aware records for each branch, including evidence, uncertainty, conflicts, progress, and failure cues. By keeping different search directions separated, these records allow TreeSearch to compare branches, identify useful or failed attempts, and decide whether to explore, exploit, or prune.
Figure 1 shows the overall loop. TreeSearch reads the current TreeMem states and outputs one decision for each active sub-goal tree (Explore/Exploit/Prune). After tool execution or pruning, TreeMem updates the corresponding active branch states. This loop lets the agent compare search directions, continue useful ones, test alternatives, and recover from unhelpful paths.
Together, TreeMem and TreeSearch turn trial-and-error into branch-and-return search. We use Textual UCB because trial-and-error needs an explicit control principle. It should not greedily follow the current best-looking branch. It should also not explore alternatives without discipline. A literal path-level UCB formulation would score every path-action pair per sub-goal. That creates a large combinatorial space and turns deep-search control into fine-grained path ranking. We therefore decide at the operation level. For each dependency-ready sub-goal, TreeSearch compares Exploit/Explore/Prune as budget-allocation choices and then grounds the selected operation on concrete path target(s). This preserves the explore and exploit logic, allocates budget across exploit/explore/return in a principled way, and avoids exhaustive path-action enumeration.
3.2 TreeMem
Tree-structured search state.
TreeMem defines the state interface that TreeSearch reads before each decision round. For each sub-goal , TreeMem stores a tree . The root stores the goal state, including the sub-goal summary and current result candidates. The first-level nodes are candidate paths for solving the sub-goal, and each path node stores a branch state, including evidence, uncertainty, progress, and failure cues. Deeper nodes store the recent trace, including the latest tool calls and returned observations.
Figure 2 illustrates this three-level structure. TreeMem does not keep every trial in full. Long-term history is summarized into goal and branch states. Recent traces keep only the latest raw interactions. Leaf traces store interactions at leaf nodes. Pruned continuations are compressed into short failure cues. This gives TreeSearch enough structured information to compare branches and decide whether to explore, exploit, or prune without replaying the full interaction history.
3.3 TreeSearch
TreeSearch realizes branch-and-return search as an explicit inference-time control loop. At each decision round, TreeSearch reads all sub-goal trees, selects dependency-ready unresolved goals, assigns one operation to each selected tree in a single pass, and updates active trees accordingly. For each selected goal, the decision can be Explore, Exploit, or Prune. This lets the controller advance multiple sub-goals in parallel while keeping each tree update explicit and local.
At decision round , TreeSearch first reads all sub-goal trees in one shot.
| (1) |
It then constructs an actionable frontier over trees that are ready to search but not solved yet.
| (2) | ||||
Here, is a compact view of tree , including its active branches and their current states. Let be the indices of trees in . TreeSearch then outputs one decision for every tree in in a single pass. Trees outside are not executed in the current round and remain unchanged.
Textual UCB operation selection.
Textual UCB gives TreeSearch a simple rule for controlled trial-and-error. It avoids both greedy continuation along the current best-looking branch and undisciplined exploration. It also gives a decision logic for allocating budget across exploit, explore, and return. For this reason, TreeSearch performs operation selection rather than branch ranking. For each active goal , it compares Exploit, Explore, and Prune, then binds each operation to concrete target(s) inside the tree. This operation binding compresses the action space from unconstrained path-action combinations to compact operation-conditioned candidate sets. Concretely, we apply two stage filtering before textual scoring. The first stage keeps DAG-ready unresolved goals. The second stage truncates operation-conditioned targets at the path level. This yields a bounded candidate budget with . Because branch quality is semantically expressed (partial evidence, source reliability, unresolved conflicts, progress, and failure cues), we use textual rather than numeric UCB.
TreeSearch then applies a textual UCB-style rule defined at the operation level over TreeMem states. For each candidate operation with binding and induced state , it estimates three ordinal semantic signals.
| (3) |
Here, (Value) measures expected progress, (Uncertainty) measures expected information gain, and (Risk) measures the chance of committing budget to a misleading continuation.
To make the decision actionable across rounds, we map these ordinal signals to discrete levels (Low, Medium, High), and use a parameter-free score.
| (4) |
where . TreeSearch selects the operation with the largest . In this way, textual UCB preserves the explore-exploit principle while adapting it to semantically rich branch states and adding a risk-aware return mechanism for branch-and-return search.
Figure 3 illustrates this decision process with a concrete pruning-and-revision case. For the Xinjiang sub-goal, TreeSearch identifies the cross-check continuation as high-risk because it mixes prefecture-level and county-level units, prunes it into a compact failure cue, and returns to the earlier branch point for revision.
Trial-and-error operations.
TreeSearch outputs one operation for each active sub-goal tree in the current round. Together, these operations implement branch-and-return by either expanding a search direction, continuing it, or returning from it.
| (5) | ||||
Here, is the active set of dependency-ready unresolved goals, is the selected target in tree , and is the selected operation. For Explore and Exploit, is the branch to extend. For Prune, includes the weak continuation to stop and the earlier branch point to return to.
- •
Exploit continues a promising branch by extending its current path with a new action. For example, if an official-source branch has found a reliable government page, TreeSearch may continue it by extracting the candidate prefecture list.
- •
Explore tests an uncertain alternative by opening a new continuation or sibling branch from the selected branch point. For example, if a map-based branch may reveal missing border units but is not yet verified, TreeSearch may open a new continuation to inspect administrative border maps.
- •
Prune stops an unhelpful continuation and uses return to move back to an earlier branch point. TreeMem keeps a compact failure cue at the pruned continuation and keeps the return point available for later revision. For example, if a cross-check continuation repeatedly mixes prefecture-level and county-level units, TreeSearch prunes it and returns to the earlier branch point instead of extending the same weak path.
3.4 TreeSearch and TreeMem Feedback Loop
After TreeSearch outputs the decision set , each active tree () is updated independently in the current round. For Explore and Exploit, the selected branch is passed to the action generator, which executes branch-tagged tool actions. The resulting observation is written back to the corresponding branch in TreeMem. For Prune, TreeMem records a failure cue and preserves the earlier branch point for future revision.
The update for each selected tree is written as
| (6) |
where is either the returned observation or the pruning signal. Trees outside remain unchanged. The updated collection of trees is then used in the next decision round.
Together, these operations implement structured trial-and-error, where Explore creates trials, Exploit deepens useful trials, and Prune returns from failed ones.
The main text abstracts this process as a TreeSearch–TreeMem feedback loop; the concrete inference procedure and prompt templates are provided in Appendices A and E.
4 Experiments
4.1 Experimental Setup
Benchmarks.
We evaluate TreeSeeker on three public deep-search benchmarks: XBench-DeepSearch (Chen et al., 2025), BrowseComp (Wei et al., 2025), and BrowseComp-ZH (Zhou et al., 2025). Together, they cover long-horizon web search, compositional browsing, and Chinese-language browsing scenarios. Due to resource constraints, we randomly sample instances from BrowseComp and instances from BrowseComp-ZH as test subsets and report results on these sampled subsets. Additional benchmark details are provided in Appendix B.
Baselines.
We report results for representative open-source deep-search systems, including Tongyi DeepSearch (Team et al., 2025b), IterResearch (Chen et al., 2026a), Flash-Searcher (Qin et al., 2026), and LATS (Zhou et al., 2024), as well as reported proprietary systems such as OpenAI DeepResearch (OpenAI, 2025) and Gemini DeepResearch (Google, 2025), among other systems listed in Table 1.
Implementation Details.
In this paper, gpt-5.2 refers to gpt-5.2-20251211, and gpt-4.1 refers to gpt-4.1-20250414. We use gpt-5.2 as the default backbone, and additionally evaluate both Flash-Searcher and TreeSeeker with gpt-4.1 under the same search and browsing tools. Web search uses the Bing Search API v7 (Microsoft, 2026), and page access/parsing uses Firecrawl (Firecrawl, 2025); further implementation details are provided in Appendix B.
4.2 Main Results
| Methods | XB-DS | BC | BC-ZH |
| Closed-source deep search Systems/Models | |||
| Gemini-DR* | 53 | 37.8 | – |
| Perplexity Deep Research* | – | 22.0 | 22.6 |
| Claude-4-Sonnet-Thinking* | 53.0 | 14.7 | 30.8 |
| Claude-4.5-Sonnet* | 66.0 | 19.6 | 40.8 |
| Gemini-2.5-Pro* | 56.0 | 9.9 | 32.2 |
| OpenAI GPT-5* | 30.0 | 19.8 | 34.3 |
| OpenAI o1* | – | 9.9 | 29.1 |
| OpenAI o3* | 68.0 | 55.0 | 59.0 |
| OpenAI DeepResearch* | 66.7 | 51.5 | 42.9 |
| Grok3 DeepResearch* | 50+ | 12.9 | – |
| Doubao DeepResearch* | 50+ | – | 26.0 |
| Open-source deep search Systems/Models | |||
| DeepSeek-R1* | 32.7 | 2.0 | 23.2 |
| Qwen3-235B-A22B-Instruct-2507* | 45.5 | 8.0 | 23.0 |
| Kimi K2* | 54.0 | 11.0 | 22.0 |
| Search-o1-32B* | 25.0 | 2.8 | 17.9 |
| WebDancer-32B* | 39.0 | 3.8 | 18.0 |
| WebSailor-32B* | 53.3 | 10.5 | 25.5 |
| LATS (gpt-5.2) | 31.0 | 16.0 | 25.7 |
| Tongyi-DeepSearch-30B-A3B | 45.0 | 33.3 | 33.0 |
| IterResearch (gpt-5.2) | 44.0 | 35.3 | 34.0 |
| Flash-Searcher (gpt-4.1) | 21.3 | 5.7 | 17.7 |
| Flash-Searcher (gpt-5.2) | 50.7 | 43.0 | 40.3 |
| Ours | |||
| TreeSeeker (gpt-4.1) | 23.0 | 7.7 | 20.3 |
| TreeSeeker (gpt-5.2) | 56.3 | 47.0 | 43.0 |
Using gpt-5.2, TreeSeeker achieves 56.3 on XBench-DS, outperforming Flash-Searcher, IterResearch, and Tongyi-DeepSearch by 5.6, 12.3, and 11.3 points, respectively. On BrowseComp and BrowseComp-ZH, it reaches 47.0 and 43.0, ranking first among the evaluated open-source baselines. Under a shared gpt-4.1 backend, TreeSeeker also remains ahead of Flash-Searcher, with gains of 1.7, 2.0, and 2.6 points across the three benchmarks, suggesting that the advantage of TreeSearch–TreeMem is not tied to a single backend model. The strongest controlled comparison is with Flash-Searcher, since both systems use tree-structured search but differ in how candidate paths are controlled after intermediate evidence appears. The consistent gains over Flash-Searcher under both backbones suggest that explicitly maintaining branch states and selecting among continuation, exploration, and pruning decisions improves the effectiveness of the search process itself. Taken together, these results show that TreeSeeker generalizes well across different deep-search benchmarks, including both English and Chinese settings, and achieves the best performance among the evaluated open-source baselines. Beyond task accuracy, we also analyze the token usage and tool-call cost of each system in Appendix C.
4.3 Cumulative Success over Action Steps
Figure 4 shows cumulative success rate versus action steps on the XBench-DS queries shared by all runs, reported as the per-step mean over three runs with min–max envelopes.
The two methods behave similarly in the first several action steps, but their trajectories diverge around step , after which TreeSeeker remains consistently ahead. The gap is visible in both the main growth region (roughly action steps –) and near the end of the budget, where TreeSeeker reaches a higher final cumulative success rate. We attribute this gap to evidence-driven branch selection rather than fixed-schedule path execution. Flash-Searcher executes candidate paths within a goal under a fixed DAG schedule, so later observations have limited effect on whether a path should be continued, revised, or abandoned. In contrast, TreeSeeker keeps branch-local evidence, uncertainty, conflicts, and failure cues in TreeMem, and TreeSearch compares Exploit, Explore, and Prune directly over these branch states using textual UCB signals.
4.4 Ablation Study
To isolate the contribution of each core component, we conduct ablation experiments on the XBench-DS benchmark. We consider three ablated variants: (1) w/o Textual UCB, which removes the operation-level value, uncertainty, and risk signals from the TreeSearch decision prompt, so the controller selects operations without the textual UCB-style comparison defined in the method section; (2) w/o Explore & Prune, which disables the Explore and Prune operations, so TreeSearch retains only Exploit: it can continue existing branches, but cannot open new continuations or sibling branches, and cannot stop high-risk continuations and return to earlier branch points; (3) w/o Leaf Trace in TreeMem, which removes the retained latest raw leaf trace from TreeMem, forcing the agent to rely only on summarized goal and branch states and losing the short-term continuation anchor for each path.
| Variant | avg |
| TreeSeeker (full) | 56.3 |
| TreeSeeker w/o Textual UCB | 52.0 |
| TreeSeeker w/o Explore & Prune | 48.0 |
| TreeSeeker w/o Leaf Trace in TreeMem | 51.3 |
Table 2 reports the results. Removing any component degrades performance, showing that textual UCB scoring, branch-and-return operations, and the TreeMem design make complementary contributions. Removing textual UCB lowers XBench-DS performance from 56.3 to 52.0 (4.3), disabling Explore and Prune causes the largest drop to 48.0 (8.3), and removing the retained leaf trace in TreeMem lowers performance to 51.3 (5.0). These results indicate that semantic operation scoring, branch-and-return control, and structured branch memory make complementary contributions to effective long-horizon search.
4.5 Operation Decision Analysis
Table 3 reports the empirical frequencies of branch operations on XBench-DS for both the full controller and the w/o Textual UCB ablation.
| Exploit | Explore | Prune | |
| TreeSeeker | |||
| w/o Textual UCB |
With textual UCB guidance, TreeSeeker keeps Exploit and Explore relatively balanced ( vs. ), while using Prune sparingly (). Removing textual UCB shifts the controller toward substantially more Explore decisions () and fewer Exploit and Prune decisions, suggesting that the semantic control signal helps allocate budget more effectively across deepening, branching, and correction. Detailed discussion of these tradeoffs is provided in Appendix D.
5 Conclusion
We presented TreeSeeker, an inference-time framework that structures deep search as branch-and-return search over tree-structured states. The key insight is that early-stage decision uncertainty in deep search is best addressed not by a stronger single-path agent, but by maintaining multiple tentative search directions as explicit decision objects and repeatedly allocating budget across them. TreeSeeker realizes this through two tightly coupled components: TreeSearch, which applies an operation-level textual UCB-style rule to decide whether to exploit a promising branch, explore an uncertain alternative, or prune an unproductive continuation and return to an earlier branch point; and TreeMem, which keeps evidence, uncertainty, conflicts, progress, and failure cues attached to the branches that produced them. Experiments on XBench-DeepSearch, BrowseComp, and BrowseComp-ZH show that TreeSeeker consistently outperforms strong open-source baselines, suggesting that explicit branch-and-return control is a valuable complement to stronger reasoning and tool execution in long-horizon deep search.
Limitations
TreeSeeker has several limitations. First, our evaluation is limited to text-based deep search benchmarks. We do not consider multimodal deep research settings because TreeSeeker currently does not integrate multimodal tools such as image or video understanding. Extending the framework to multimodal evidence sources is left for future work.
Second, branch-and-return control introduces additional inference cost. TreeSearch requires controller decisions based on value, uncertainty, and risk signals, and TreeMem periodically summarizes branch-local states. Although our cost analysis shows that TreeSeeker remains below Flash-Searcher in total tokens and tool calls, the additional controller and memory updates may still be a concern in latency or budget-sensitive deployments.
Third, TreeSeeker relies on external web search and browsing results, which can be noisy, incomplete, outdated, or biased. TreeSearch and TreeMem help compare evidence, preserve failure cues, and prune unproductive continuations, but they do not guarantee that all retrieved sources are reliable or that all conflicts are fully resolved. In high-stakes settings, outputs should therefore be checked against trusted sources or human expert review.
References
- Auer et al. (2002) Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer. 2002. Finite-time analysis of the multiarmed bandit problem. Mach. Learn., 47(2–3):235–256.
- Chen et al. (2026a) Guoxin Chen, Zile Qiao, Xuanzhong Chen, Donglei Yu, Haotian Xu, Xin Zhao, Ruihua Song, Wenbiao Yin, Huifeng Yin, Liwen Zhang, Kuan Li, Minpeng Liao, Yong Jiang, Pengjun Xie, Fei Huang, and Jingren Zhou. 2026a. Iterresearch: Rethinking long-horizon agents with interaction scaling. In The Fourteenth International Conference on Learning Representations.
- Chen et al. (2025) Kaiyuan Chen, Yixin Ren, Yang Liu, Xiaobo Hu, Haotong Tian, Tianbao Xie, Fangfu Liu, Haoye Zhang, Hongzhang Liu, Yuan Gong, Chen Sun, Han Hou, Hui Yang, James Pan, Jianan Lou, Jiayi Mao, Jizheng Liu, Jinpeng Li, Kangyi Liu, and 14 others. 2025. xbench: Tracking agents productivity scaling with profession-aligned real-world evaluations. Preprint, arXiv:2506.13651.
- Chen et al. (2026b) Kedi Chen, Dezhao Ruan, Yuhao Dan, Yaoting Wang, Siyu Yan, Xuecheng Wu, Yinqi Zhang, Qin Chen, Jie Zhou, Liang He, Biqing Qi, Linyang Li, Qipeng Guo, Xiaoming Shi, and Wei Zhang. 2026b. A survey of inductive reasoning for large language models. Preprint, arXiv:2510.10182.
- Chen et al. (2026c) Qianben Chen, Tianrui Qin, King Zhu, Qiexiang Wang, Chengjun Yu, Shu Xu, Jiaqi Wu, Jiayu Zhang, Xinpeng Liu, Xin Gui, Jingyi Cao, Piaohong Wang, Dingfeng Shi, He Zhu, Tiannan Wang, Yuqing Wang, Maojia Song, Tianyu Zheng, Ge Zhang, and 5 others. 2026c. Search more, think less: Rethinking long-horizon agentic search for efficiency and generalization. Preprint, arXiv:2602.22675.
- Feng et al. (2025) Wenfeng Feng, Chuzhan Hao, Yuewei Zhang, Guochao Jiang, Jingyi Song, and Hao Wang. 2025. Airrag: Autonomous strategic planning and reasoning steer retrieval augmented generation. Preprint, arXiv:2501.10053.
- Firecrawl (2025) Firecrawl. 2025. Firecrawl: Search, scrape, and clean the web for ai agents. Accessed: 2026-04-24.
- Google (2025) Google. 2025. Gemini deep research — your personal research assistant. https://gemini.google/overview/deep-research/. Accessed: 2025-12-29.
- Guo et al. (2025) Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, and 175 others. 2025. Deepseek-r1 incentivizes reasoning in llms through reinforcement learning. Nature, 645(8081):633–638.
- Hu et al. (2025) Yunhai Hu, Yilun Zhao, Chen Zhao, and Arman Cohan. 2025. Mcts-rag: Enhancing retrieval-augmented generation with monte carlo tree search. Preprint, arXiv:2503.20757.
- Li et al. (2025a) Kuan Li, Zhongwang Zhang, Huifeng Yin, Liwen Zhang, Litu Ou, Jialong Wu, Wenbiao Yin, Baixuan Li, Zhengwei Tao, Xinyu Wang, Weizhou Shen, Junkai Zhang, Dingchu Zhang, Xixi Wu, Yong Jiang, Ming Yan, Pengjun Xie, Fei Huang, and Jingren Zhou. 2025a. Websailor: Navigating super-human reasoning for web agent. Preprint, arXiv:2507.02592.
- Li et al. (2025b) Xiaoxi Li, Guanting Dong, Jiajie Jin, Yuyao Zhang, Yujia Zhou, Yutao Zhu, Peitian Zhang, and Zhicheng Dou. 2025b. Search-o1: Agentic search-enhanced large reasoning models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 5420–5438, Suzhou, China. Association for Computational Linguistics.
- Li et al. (2025c) Xiaoxi Li, Jiajie Jin, Guanting Dong, Hongjin Qian, Yongkang Wu, Ji-Rong Wen, Yutao Zhu, and Zhicheng Dou. 2025c. Webthinker: Empowering large reasoning models with deep research capability. Preprint, arXiv:2504.21776.
- Lu et al. (2025) Rui Lu, Zhenyu Hou, Zihan Wang, Hanchen Zhang, Xiao Liu, Yujiang Li, Shi Feng, Jie Tang, and Yuxiao Dong. 2025. Deepdive: Advancing deep search agents with knowledge graphs and multi-turn rl. Preprint, arXiv:2509.10446.
- Microsoft (2026) Microsoft. 2026. Grounding with bing. Web page. Accessed: 2026-05-01.
- Nie et al. (2026) Lunyiu Nie, Nedim Lipka, Ryan A. Rossi, and Swarat Chaudhuri. 2026. Efficient tree-structured deep research with adaptive resource allocation. In ICLR 2026 Workshop on Lifelong Agents: Learning, Aligning, Evolving.
- OpenAI (2025) OpenAI. 2025. Introducing deep research. https://openai.com/index/introducing-deep-research/. Published: 2025-02-02. Accessed: 2026-05-01.
- Qin et al. (2026) Tianrui Qin, Qianben Chen, Sinuo Wang, He Xing, King Zhu, He Zhu, Dingfeng Shi, Xinxin Liu, Ge Zhang, Jiaheng Liu, Xitong Gao, Yuchen Eleanor Jiang, and Wangchunshu Zhou. 2026. Flash-searcher: Fast and effective web agents via DAG-based parallel execution. In The Fourteenth International Conference on Learning Representations.
- Sun et al. (2025) Zhongxiang Sun, Qipeng Wang, Weijie Yu, Xiaoxue Zang, Kai Zheng, Jun Xu, Xiao Zhang, Yang Song, and Han Li. 2025. Rearter: Retrieval-augmented reasoning with trustworthy process rewarding. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’25, page 1251–1261, New York, NY, USA. Association for Computing Machinery.
- Team et al. (2025a) Kimi Team, Yifan Bai, Yiping Bao, Guanduo Chen, Jiahao Chen, Ningxin Chen, Ruijue Chen, Yanru Chen, Yuankun Chen, Yutian Chen, and 1 others. 2025a. Kimi k2: Open agentic intelligence. arXiv preprint arXiv:2507.20534.
- Team et al. (2025b) Tongyi DeepResearch Team, Baixuan Li, Bo Zhang, Dingchu Zhang, Fei Huang, Guangyu Li, Guoxin Chen, Huifeng Yin, Jialong Wu, Jingren Zhou, Kuan Li, Liangcai Su, Litu Ou, Liwen Zhang, Pengjun Xie, Rui Ye, Wenbiao Yin, Xinmiao Yu, Xinyu Wang, and 38 others. 2025b. Tongyi deepresearch technical report. Preprint, arXiv:2510.24701.
- Wei et al. (2025) Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, Jeffrey Han, Isa Fulford, Hyung Won Chung, Alex Tachard Passos, William Fedus, and Amelia Glaese. 2025. Browsecomp: A simple yet challenging benchmark for browsing agents. Preprint, arXiv:2504.12516.
- Wu et al. (2025) Jialong Wu, Baixuan Li, Runnan Fang, Wenbiao Yin, Liwen Zhang, Zhengwei Tao, Dingchu Zhang, Zekun Xi, Gang Fu, Yong Jiang, Pengjun Xie, Fei Huang, and Jingren Zhou. 2025. Webdancer: Towards autonomous information seeking agency. Preprint, arXiv:2505.22648.
- Wu et al. (2026) Xixi Wu, Kuan Li, Yida Zhao, Liwen Zhang, Litu Ou, Huifeng Yin, Zhongwang Zhang, Xinmiao Yu, Dingchu Zhang, Yong Jiang, Pengjun Xie, Fei Huang, Minhao Cheng, Shuai Wang, Hong Cheng, and Jingren Zhou. 2026. Resum: Unlocking long-horizon search intelligence via context summarization. Preprint, arXiv:2509.13313.
- Yang et al. (2025) An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, and 41 others. 2025. Qwen3 technical report. Preprint, arXiv:2505.09388.
- Yao et al. (2023) Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations.
- Ye et al. (2025) Rui Ye, Zhongwang Zhang, Kuan Li, Huifeng Yin, Zhengwei Tao, Yida Zhao, Liangcai Su, Liwen Zhang, Zile Qiao, Xinyu Wang, and 1 others. 2025. Agentfold: Long-horizon web agents with proactive context management. arXiv preprint arXiv:2510.24699.
- Yu et al. (2026) Hongli Yu, Tinghong Chen, Jiangtao Feng, Jiangjie Chen, Weinan Dai, Qiying Yu, Ya-Qin Zhang, Wei-Ying Ma, Jingjing Liu, Mingxuan Wang, and Hao Zhou. 2026. Memagent: Reshaping long-context LLM with multi-conv RL-based memory agent. In The Fourteenth International Conference on Learning Representations.
- Yu et al. (2025) Xiao Yu, Baolin Peng, Vineeth Vajipey, Hao Cheng, Michel Galley, Jianfeng Gao, and Zhou Yu. 2025. ExACT: Teaching AI agents to explore with reflective-MCTS and exploratory learning. In The Thirteenth International Conference on Learning Representations.
- Zhang et al. (2026a) Lingfeng Zhang, Yongan Sun, Jinpeng Hu, Hui Ma, Yang Ying, Kuien Liu, Zenglin Shi, and Meng Wang. 2026a. Webuncertainty: Dual-level uncertainty driven planning and reasoning for autonomous web agent. Preprint, arXiv:2604.17821.
- Zhang et al. (2026b) Weiming Zhang, Jihong Wang, Jiamu Zhou, Qingyao Li, Xinbei Ma, Congmin Zheng, Xingyu Lou, Weiwen Liu, Zhuosheng Zhang, Jun Wang, Yong Yu, and Weinan Zhang. 2026b. Plan-mcts: Plan exploration for action exploitation in web navigation. Preprint, arXiv:2602.14083.
- Zhou et al. (2024) Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang, and Yu-Xiong Wang. 2024. Language agent tree search unifies reasoning, acting, and planning in language models. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org.
- Zhou et al. (2025) Peilin Zhou, Bruce Leon, Xiang Ying, Can Zhang, Yifan Shao, Qichen Ye, Dading Chong, Zhiling Jin, Chenxuan Xie, Meng Cao, Yuxin Gu, Sixin Hong, Jing Ren, Jian Chen, Chao Liu, and Yining Hua. 2025. Browsecomp-zh: Benchmarking web browsing ability of large language models in chinese. Preprint, arXiv:2504.19314.
Appendix A TreeSeeker Inference Procedure and Implementation Details
The main text describes TreeSeeker as a TreeSearch–TreeMem loop over dependency-ready goal trees. TreeSearch constructs an active frontier of dependency-ready unresolved goal trees, selects operation-level decisions according to the textual UCB rule, and binds each selected operation to concrete branch target(s). TreeMem stores the summarized goal and branch states used by later decisions, while a short-term overlay keeps recent unsummarized records and retained leaf traces. Algorithm 1 specifies the concrete execution procedure used in our implementation, including planning initialization, frontier construction, operation-level decision making, leaf-trace retention, periodic summarization, and final answer generation.
Planning and TreeMem initialization.
The inference procedure starts with a planning agent. Given the root query , the planner decomposes the task into sub-goals , constructs a dependency DAG , and proposes candidate paths for each goal . This follows the formulation in the main text: the DAG controls which goals are dependency-ready, while candidate paths define alternative search directions under each goal. TreeMem initializes a goal-local tree for each goal, with candidate paths represented as major branches. In the implementation, denotes the summarized TreeMem states, while stores recent unsummarized records and retained leaf traces as a short-term overlay.
Frontier construction and textual UCB decision.
At each iteration, TreeSearch constructs an active frontier from the summarized TreeMem states, the recent-record overlay, and the dependency DAG. A goal tree is eligible only when its dependencies are satisfied and the goal has not yet been resolved. TreeSearch then applies the operation-level textual UCB rule described in the main text: for each active goal, it first binds each candidate operation (Exploit, Explore, or Prune) to concrete branch target(s), then scores the bound operations using value, uncertainty, and risk signals. The resulting decision set contains one bound operation-level decision for each selected active goal tree. The recent records and retained leaf traces provide short-term context, so the controller can use fresh evidence before it is fully compressed into branch summaries.
Decision execution and leaf-trace update.
For Explore and Exploit, the execution agent carries out the selected branch action using the available execution tools, including web search, page crawling, and a Python sandbox for code execution. The resulting action-observation pair is appended to the recent records of the selected path and stored as the latest leaf trace at that path’s leaf node, replacing the previous continuation anchor. For Prune, no external browsing action is required. TreeMem records a compact failure cue for the weak continuation and preserves the earlier branch point as the return target. Thus, each path keeps a compact summarized state for long-term decision making and a latest leaf trace or return point for short-term continuation.
Periodic summary updates.
TreeMem is summarized periodically rather than fully rewritten after every action. In our implementation, the summary interval is decision rounds. At each summary step, the summary agent folds recent unsummarized operation records, including action-observation pairs and pruning cues, into the corresponding goal and path states. The branch summaries are updated with revised evidence, provenance, candidate answers, unresolved constraints, progress signals, conflicts, and failure cues. After consolidation, older unsummarized records may be compressed or removed to prevent TreeMem from degenerating into a flat growing history. However, the latest leaf trace of each active path is retained. This leaf trace is included in the summary update, but it remains explicitly available as the continuation anchor for future TreeSearch decisions and execution.
Termination and answer generation.
The loop terminates when AnswerFound determines that sufficient evidence has been collected, or when the global step budget is exhausted. If unsummarized records remain at termination, they are summarized once more before final synthesis. The final answer is then generated from the updated TreeMem state, including surviving candidate answers, supporting evidence, confidence signals, and unresolved conflict markers when applicable.
Appendix B Experimental Details
Benchmarks.
We evaluate TreeSeeker on three publicly available benchmarks that require long-horizon web search, multi-source evidence integration, and complex information synthesis:
- •
XBench-DeepSearch (Chen et al., 2025) is a benchmark designed to evaluate deep-search capabilities in realistic information-seeking scenarios. Its tasks require systems to decompose underspecified questions, perform multiple rounds of search refinement, integrate heterogeneous evidence from different web sources, and synthesize a final answer.
- •
BrowseComp (Wei et al., 2025) is a browsing-based compositional search benchmark released by OpenAI. It focuses on hard-to-find, entangled information that cannot usually be answered by a single retrieval step. Solving these questions requires formulating effective search queries, navigating search results, inspecting web pages, extracting relevant facts, and synthesizing evidence across multiple sources. Due to resource constraints, we randomly sample BrowseComp instances as a test subset and report results on this sampled subset.
- •
BrowseComp-ZH (Zhou et al., 2025) extends the BrowseComp setting to Chinese web environments and information sources. It evaluates whether a deep-search system can handle non-English browsing, cross-page reasoning, and culturally or linguistically specific evidence aggregation. We include this benchmark to test cross-lingual and cross-cultural generalization beyond English web search. Due to resource constraints, we randomly sample BrowseComp-ZH instances as a test subset and report results on this sampled subset.
Together, these benchmarks provide a broad evaluation landscape for testing whether TreeSeeker can allocate search budget effectively across English and Chinese browsing tasks, short- and long-horizon evidence gathering, and both single-benchmark and cross-benchmark generalization settings.
Baselines.
We compare TreeSeeker with representative systems that cover both open-source research agents and reported proprietary deep-search products. The main goal of these comparisons is to distinguish gains from the proposed branch-and-return controller from gains that may come from stronger backend models or larger search budgets.
- •
Open-source deep-search systems. Tongyi DeepSearch (Team et al., 2025b) is an end-to-end deep-search agent based on its own post-trained Tongyi-A30B-A3B model variant. IterResearch (Chen et al., 2026a) improves long-horizon search through iterative workspace reconstruction and interaction scaling. Flash-Searcher (Qin et al., 2026) uses a DAG-style parallel search framework and is the most direct structural baseline for evaluating whether explicit branch-level control improves over fixed-schedule parallel search. LATS (Zhou et al., 2024) integrates reasoning, acting, and planning in an MCTS-style framework, providing a tree-search baseline for agentic reasoning.
- •
Reported open-source model systems. We also include reported results for open-source model or agent systems such as DeepSeek-R1 (Guo et al., 2025), Qwen3 (Yang et al., 2025), K2 (Team et al., 2025a), Search-o1 (Li et al., 2025b), WebDancer (Wu et al., 2025), and WebSailor (Li et al., 2025a) when benchmark numbers are available. These results help contextualize the performance of TreeSeeker against broader open-source progress in reasoning and web-agent systems.
- •
Proprietary deep-search systems and models. We report available results from commercial systems and closed-source models, including OpenAI DeepResearch (OpenAI, 2025), Gemini DeepResearch (Google, 2025), and other systems listed in Table 1. These comparisons are not controlled for model scale or product-specific infrastructure, but they provide useful reference points for understanding how far an open research framework can approach strong proprietary deep-search systems.
Implementation Details.
For reimplemented open-source systems, we keep the backend model and tool interface aligned whenever the original system design permits, so that differences are mainly attributable to the search-control framework rather than to different web access tools. Our default backend for reimplemented systems and TreeSeeker is gpt-5.2-20251211; Tongyi DeepSearch is evaluated with its released Tongyi-A30B-A3B model variant. To separate framework-level gains from backend-model effects, we additionally run controlled gpt-4.1-20250414 variants of Flash-Searcher and TreeSeeker under the same search and browsing stack. The tool interface consists of web search through the Bing Search API v7 (Microsoft, 2026), URL visiting/page parsing through Firecrawl (Firecrawl, 2025), and a Python sandbox for code execution when needed. For each Bing search query, we use the default setting that returns the top search results. For open-source systems with executable implementations, we report averages over three independent runs when applicable; for proprietary systems and model-only entries, we use the publicly reported numbers indicated in Table 1.
License and Terms of Use.
All benchmarks and baseline implementations used in this work are publicly available research artifacts. We use them only for research and evaluation purposes, following their original licenses and terms of use. When reusing or adapting baseline code, we preserve the corresponding license notices and attributions. We do not redistribute any restricted third-party datasets or artifacts beyond what is permitted by their original licenses.
Appendix C Cost Analysis
We report the computational cost of TreeSeeker to clarify the efficiency tradeoff introduced by branch-and-return control. Our method is not designed to minimize token usage alone: TreeSearch adds a controller decision before execution, and TreeMem periodically summarizes branch-local states so that later decisions can compare, continue, or prune search directions. This additional control cost is the price paid for structured trial-and-error. Nevertheless, the overall cost remains within a practical range for deep-search systems. As shown in Table 4, TreeSeeker is not the cheapest system in our comparison: it uses more total tokens and tool calls than lighter baselines such as IterResearch, Tongyi-DeepResearch, and LATS. However, compared with Flash-Searcher, the strongest reimplemented open-source baseline in our main results, TreeSeeker still achieves better XBench-DS performance while using fewer total tokens and fewer tool calls.
| Methods | Input (K) | Output (K) | Total (K) | ToolCalls |
| Baseline | ||||
| IterResearch | 342.2 | 66.5 | 408.7 | 18.47 |
| Flash-Searcher | 1445.0 | 29.6 | 1474.6 | 90.20 |
| Tongyi-DeepSearch | 916.2 | 12.1 | 928.4 | 20.71 |
| LATS | 1344.9 | 19.6 | 1364.5 | 25.49 |
| Ours | ||||
| TreeSeeker | 1337.6 | 78.2 | 1415.8 | 71.92 |
These results indicate that the improvement of TreeSeeker is not obtained by simply scaling up inference cost beyond existing strong systems. Compared with Flash-Searcher, TreeSeeker reduces total token usage from 1474.6K to 1415.8K and reduces tool calls from 90.20 to 71.92, while improving XBench-DS performance from 50.7 to 56.3. At the same time, its cost remains higher than lighter baselines such as IterResearch and Tongyi-DeepResearch, and it also exceeds LATS in both total tokens and tool calls, reflecting the additional controller and memory updates required for branch-level search control. We view this as a reasonable efficiency–accuracy tradeoff: TreeSeeker does not minimize raw inference cost, but it converts budget into stronger task performance more effectively than the fixed-schedule Flash-Searcher baseline.
Appendix D Operation Decision Analysis Details
The operation frequencies in Table 3 help clarify the role of textual UCB as a balancing signal rather than a simple trigger for more exploration. With UCB guidance, TreeSeeker keeps Exploit and Explore at comparable levels, indicating that the controller continues promising branches often enough to consolidate evidence while still opening alternatives to reduce premature commitment. In contrast, the w/o Textual UCB variant explores much more frequently and exploits much less frequently, suggesting that without explicit value–uncertainty–risk signals, the controller is more likely to keep broadening the search frontier instead of committing budget to evidence chains that have already become promising.
This distinction should not be interpreted as saying that one raw operation distribution is intrinsically optimal. More Explore is not automatically better, because excessive exploration can leave promising evidence under-developed; more Exploit is also not automatically better, because over-commitment can lock the agent onto a locally plausible but incomplete candidate. The second case study in Appendix F illustrates this tradeoff: TreeSeeker first uses exploration to avoid prematurely locking onto distractor activities, and then shifts to exploitation once Nuo Opera becomes the strongest candidate, concentrating budget on the most discriminative evidence. Thus, the useful behavior is not maximizing either Explore or Exploit, but deciding when to trade off breadth for depth.
The modest but non-trivial Prune rate () further suggests that UCB guidance also supports risk-aware correction. Compared with w/o Textual UCB, the full controller prunes more often, but pruning remains sparse, indicating that TreeSearch does not aggressively discard branches whenever value estimates fluctuate. Instead, it invokes Prune when the value–uncertainty–risk profile suggests that a continuation is confidently unproductive, allowing the agent to redirect budget while preserving potentially recoverable branches. Together with the ablation results, these statistics support the intended role of textual UCB: it provides a semantic control signal for allocating budget among exploiting strong leads, exploring uncertain alternatives, and pruning risky continuations.
Appendix E Prompts
E.1 Initial Planning
The following prompt initializes the DAG-structured plan by decomposing the root question into goals and alternative paths.
E.2 Textual UCB Decision
The following prompt is used by the TreeSearch controller to select goals, paths, and decision modes from the current TreeMem state. In the implementation prompt, the mode name “backtrack” corresponds to the Prune operation described in the main text.
E.3 Summarize
The following prompt is used by the summary agent to consolidate execution trajectories into progress summaries and structured TreeMem updates.