arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2609.02074v1 [cs.AI] 02 Sep 2026

CHIME: Credit-Aware Hierarchical Memory Evolution for
Long-Horizon Agentic Planning

Yongshi Ye    Tian Lan    Feihu Jiang    Muyang Ye    Bin Zhu    Qianghuai Jia    Longyue Wang\corresponding    Zhao Xu    Weihua Luo    Xiaodong Shi\corresponding
Abstract

Planning is a central capability that enables agents to decompose complex long-horizon tasks into manageable steps. Test-time search and training-based methods improve planning but incur high inference costs or require expensive training data. Self-evolving memory instead accumulates reusable experience from agent interaction outcomes into an external memory bank, so planning capability keeps improving at inference time without parameter updates. However, existing self-evolving memory methods share an inherent credit assignment problem: they rely on final task outcomes as feedback, but such outcomes conflate plan quality with execution errors and environmental factors, so the accumulated planning experience is often biased and noisy. To address this problem, we propose Credit-Aware HIerarchical Memory Evolution (CHIME), a self-evolving memory framework that maintains a separate planning bank and execution bank and follows an attribute-before-memorize principle: CHIME first attributes each task outcome to the plan, the execution, both, or neither, and then updates only the corresponding memory bank. Extensive experiments on four long-horizon agent benchmarks show that CHIME consistently outperforms state-of-the-art training-based and self-evolving memory baselines. Further analyses reveal several interesting findings. For example, CHIME accumulates effective memory with far fewer items. In addition, the learned memory values faithfully reflect downstream utility: high-quality planning memories are more valuable than execution memories. Finally, the accumulated memory effectively transfers across backbone models. Code will be released at https://github.com/ATH-MaaS/Marco-DeepResearch.

1Institute of Artificial Intelligence, Xiamen University

2Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage

of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism

3Zhejiang University

4Alibaba Group

1 Introduction

Long-horizon tasks pose a defining challenge for agentic systems: agents must coordinate decisions across interdependent steps while sustaining adherence to task constraints throughout execution Xie et al. (2024); Wang et al. (2024); Erdogan et al. (2025). Agentic planning addresses these requirements by decomposing the overall objective into modular, independently executable subgoals Wang et al. (2024); Shen et al. (2023); Zhang et al. (2026b); Liu et al. (2026) and establishing the dependencies and constraints among them Kim et al. (2023); Guo, Kingston, and Kavraki (2024); Lan et al. (2026), making it a central capability for solving long-horizon tasks.

This has motivated prior work to improve agents’ planning capability. Test-time search Yao et al. (2023a); Hao et al. (2023); Zhou et al. (2024); Yu et al. (2026); Team et al. (2026a) improves per-task performance by exploring multiple candidate plans or trajectories, but incurs high inference costs and does not retain reusable experience. To retain and reuse planning experience across tasks, subsequent work has pursued two directions: (1) Planner model training Erdogan et al. (2025); Si et al. (2026); Liu et al. (2026); Team et al. (2026b) internalizes planning capability into model parameters, but requires costly data collection and post-training, since identifying high-quality plans relies on extensive downstream executions; (2) Self-evolving memory Chen et al. (2026); Ouyang et al. (2026) instead accumulates experience in an external, training-free memory bank, offering a more efficient and scalable alternative. We therefore focus on improving planning within the self-evolving memory paradigm in this paper.

Refer to caption
Figure 1: Vanilla self-evolving agents vs. CHIME. Vanilla self-evolving agents write experience directly to shared memory. CHIME attributes task outcomes through a credit attribution gate before updating planning or execution memory.

However, existing self-evolving memory methods share a fundamental limitation for agentic planning: they equate downstream task outcomes with plan quality. The same outcome can arise from different causes: success may result from a sound plan or from executor recovery from a flawed one, while failure may stem from a flawed plan, incorrect execution, or environmental issues. Using such outcomes directly as memory feedback therefore leads to a fundamental credit assignment failure. On the one hand, once biased experience is written into the memory bank, it is repeatedly retrieved and reused across tasks, allowing local misjudgments to accumulate into systematic planning errors. On the other hand, execution- or environment-level experience may be mistaken for planning memory; retrieving might low-level details then misguides high-level planning decisions instead of providing planning guidance.

To address this, we propose Credit-Aware HIerarchical Memory Evolution (CHIME), a self-evolving agentic planning framework that replaces outcome-based memory updates Ye et al. (2026) with an attribute-before-memorize process. Specifically, CHIME comprises three components: (1) Hierarchical Memory Bank explicitly separates planning and execution memory into a planning bank and an execution bank, retrieving stage-specific memories through similarity retrieval and value-based reranking to separately guide the agent’s planning and execution; (2) Credit Attribution Gate then reflects on the task, plan, execution, and retrieved memories to attribute reusable feedback to planning, execution, both stages, or neither, before any memory is written; it also generates stage-specific experience and identifies misleading memories; and (3) Credit-Aware Memory Evolution updates the values of retrieved memories using confidence-weighted feedback, while filtering, merging, or inserting new experience only into the attributed bank. In this way, CHIME evolves each bank only with memory attributed to that stage, preventing biased signals from persisting and execution-level memory from contaminating the planning bank.

Experiments on four long-horizon benchmarks with two backbone models show that CHIME consistently outperforms training-based and self-evolving memory baselines, improving the eval average over the strongest baseline by 2.96% and 3.68% on the two backbones, respectively. Further analyses reveal several interesting findings: (1) CHIME achieves the highest accuracy while retaining only 129 memories, versus 3,585 for the strong baseline; (2) the learned memory values faithfully reflect downstream utility, and planning memories bring more than twice the accuracy gain of execution ones (21.7%→\to50.8% vs. 23.7%→\to35.4%); (3) the Credit Attribution Gate is reliable, with repeated attributions agreeing in up to 97.9% of cases and over half of the failures rescued by the generated experience; (4) replacing model’s self-reflection credit with a stronger credit model yields a further gain of up to 3.12%; and (5) the accumulated memory effectively transfers across backbones, outperforming A-MapReduce’s transferred memory by up to 4.68%. These findings confirm that attribute-before-memorize is the key to CHIME’s effectiveness.

2 Related Work

Self-evolving memory.

Self-evolving memory methods treat an external memory bank as the agent’s evolvable parameters, distilling interaction trajectories into reusable experience or skills Yu et al. (2025); Wang et al. (2025); Zhang et al. (2026a). Unlike training-based methods, they keep model parameters frozen and update only the memory bank, avoiding costly trajectory collection and post-training Li et al. (2025); Wu et al. (2025). For example, ReasoningBank Ouyang et al. (2026) distills generalizable reasoning strategies from self-judged successful and failed trajectories, and UMEM Ye et al. (2026) jointly optimizes memory extraction and management with reinforcement learning.

Agentic planning.

Agentic planning decomposes complex objectives into structured steps and coordinates their execution. It underlies a broad range of agent systems Yao et al. (2023b); Erdogan et al. (2025); Liu et al. (2026); Lan et al. (2026). Long-horizon tasks often involve many interdependent subtasks. Without an explicit plan to organize them, agents easily lose track of progress, a failure known as the lost-in-the-middle problem Lan et al. (2026); Liu et al. (2026). As recent benchmarks introduce increasingly challenging tasks, improving planning capability has become a central research problem for agent systems Wong et al. (2025); Lan et al. (2025); Yang et al. (2025).

Improving agentic planning.

Existing efforts fall into three paradigms: (1) Test-time search explores multiple candidate plans or trajectories at inference time and selects among them with value estimates or environmental feedback Yao et al. (2023a); Hao et al. (2023); Zhou et al. (2024). For example, MiroThinker-H1 Team et al. (2026b) and WebAnchor Yu et al. (2026) select high-quality plans via rejection sampling; (2) Planner model training internalizes planning capability into model parameters Erdogan et al. (2025); Si et al. (2026); Yu et al. (2026). For example, TodoEvolve Liu et al. (2026) constructs a modular design space of planning architectures and trains a meta-planner with reinforcement learning to synthesize task-specific planning systems. However, such methods require costly trajectory collection and post-training, and the resulting planner is expensive to update continually Liu et al. (2026); and (3) Self-evolving memory offers a training-free and continual alternative, but few works evolve memory specifically for agentic planning Kagaya et al. (2024). A representative is A-MapReduce Chen et al. (2026), which evolves structured hints from past executions to improve task decomposition and result aggregation in long-horizon agentic search. However, all three paradigms evaluate agent’s plans by final task outcomes, implicitly equating outcome with plan quality. This equivalence does not hold, as outcomes are also shaped by execution details and environmental factors.

3 Methodology of CHIME

3.1 Task Formulation of Self-Evolving Agents

We consider a self-evolving agent that solves a stream of tasks. The agent’s policy parameters remain frozen, and adaptation is carried by an external memory bank ℳt\mathcal{M}_{t} that evolves across episodes. At episode tt, the agent receives a task xtx_{t} and retrieves task-relevant experience ℳtret\mathcal{M}_{t}^{\mathrm{ret}} from ℳt\mathcal{M}_{t}. Conditioned on xtx_{t} and ℳtret\mathcal{M}_{t}^{\mathrm{ret}}, the frozen policy produces a trajectory τt\tau_{t} and receives an outcome sts_{t} from the environment and the evaluator. A memory update operator then distills the episode into reusable experience:

ℳt+1=𝒰⁡(ℳt,xt,τt,st).\mathcal{M}_{t+1}=\mathcal{U}\left(\mathcal{M}_{t},x_{t},\tau_{t},s_{t}\right). (1)

During evaluation, the memory bank is frozen, so that performance reflects experience accumulated before evaluation. Existing self-evolving memory methods mainly use the final outcome sts_{t} as the supervision signal for memory updates. For agentic planning, this signal is biased: the outcome depends on both the plan and the execution process, so it does not faithfully reflect plan quality.

3.2 Overview of CHIME

Refer to caption
Figure 2: Overview of CHIME. Left: the Hierarchical Memory Bank maintains the planning and execution memory banks. For each task, similarity retrieval and value-based reranking select stage-specific memories to guide the planner and the executor. Right: at each episode, the Credit Attribution Gate reflects on the task, plan, execution, and retrieved memories, attributes reusable feedback to planning, execution, both, or neither, and generates stage-specific experience. Feedback loop: Credit-Aware Memory Evolution updates the memory values, and filters, merges, or inserts new experience into the attributed bank.

We propose CHIME (Figure 2), an attribute-before-memorize self-evolving framework for agentic planning. It attributes each task outcome to planning, execution, or external factors, and updates only the corresponding memory bank, which reduces biased feedback and cross-stage memory contamination. CHIME consists of three components: a Hierarchical Memory Bank that separates planning and execution experience, a Credit Attribution Gate that attributes each outcome to its responsible stage, and Credit-Aware Memory Evolution that updates the banks from the attributed credit.

Hierarchical Memory Bank.

Unlike existing self-evolving methods Ye et al. (2026), CHIME maintains hierarchical memory banks

ℳt=(ℳplan,t,ℳexec,t),\mathcal{M}_{t}=\left(\mathcal{M}_{\mathrm{plan},t},\mathcal{M}_{\mathrm{exec},t}\right), (2)

which comprises a planning bank and an execution bank, indexed by ℓ∈{plan,exec}\ell\in\{\mathrm{plan},\mathrm{exec}\}. The planning bank stores strategic experience (e.g., subtask decomposition, dependency modeling, and constraint incorporation), whereas the execution bank stores operational experience (e.g., tool selection and invocation). Each bank consists of memory items

mi=(ci,ei,vi,ni),m_{i}=(c_{i},e_{i},v_{i},n_{i}), (3)

where cic_{i} describes the tasks to which the memory applies, serving as the key for memory retrieval and merging; eie_{i} is the reusable experience memory distilled from past episodes; vi∈[−1,1]v_{i}\in[-1,1] is a value score estimated from reuse feedback, which guides reranking (Eq. 5) and pruning; and nin_{i} is the reuse count, which calibrates the value score (Eq. 3.2). At episode tt, CHIME receives a long-horizon task xtx_{t} and constructs a planning query qplan,tq_{\mathrm{plan},t} to retrieve ℳplan,tret\mathcal{M}_{\mathrm{plan},t}^{\mathrm{ret}} from ℳplan,t\mathcal{M}_{\mathrm{plan},t}. The planner then generates a plan ptp_{t} from xtx_{t} and ℳplan,tret\mathcal{M}_{\mathrm{plan},t}^{\mathrm{ret}}. Next, CHIME constructs an execution query qexec,tq_{\mathrm{exec},t} from (xt,pt)(x_{t},p_{t}) and retrieves ℳexec,tret\mathcal{M}_{\mathrm{exec},t}^{\mathrm{ret}} from ℳexec,t\mathcal{M}_{\mathrm{exec},t}. The executor uses ptp_{t} and ℳexec,tret\mathcal{M}_{\mathrm{exec},t}^{\mathrm{ret}} to produce a trajectory τt\tau_{t}, and the environment returns a binary outcome st∈{0,1}s_{t}\in\{0,1\} indicating task success. Finally, the credit attribution gate produces a diagnosis gtg_{t} from the plan, trajectory, outcome, and retrieved memories, which drives the transition from ℳt\mathcal{M}_{t} to ℳt+1\mathcal{M}_{t+1}. One CHIME episode is summarized as

ℳt→ℳplan,tret→pt→ℳexec,tret→τt→st→gt→ℳt+1.\mathcal{M}_{t}\!\rightarrow\!\mathcal{M}_{\mathrm{plan},t}^{\mathrm{ret}}\!\rightarrow\!p_{t}\!\rightarrow\!\mathcal{M}_{\mathrm{exec},t}^{\mathrm{ret}}\!\rightarrow\!\tau_{t}\!\rightarrow\!s_{t}\!\rightarrow\!g_{t}\!\rightarrow\!\mathcal{M}_{t+1}. (4)

These steps correspond to the left, right, and feedback-loop parts of Figure 2.

Hierarchical Memory Retrieval.

Retrieval operates independently at each stage ℓ∈{plan,exec}\ell\in\{\mathrm{plan},\mathrm{exec}\}. Given the stage-specific query qℓ,tq_{\ell,t}, CHIME retrieves the Top-NN memories from the corresponding bank ℳℓ,t\mathcal{M}_{\ell,t} according to sim⁡(qℓ,t,ci)\mathrm{sim}(q_{\ell,t},c_{i}), yielding the candidate set 𝒞ℓ,t\mathcal{C}_{\ell,t}. For each candidate memory mi∈𝒞ℓ,tm_{i}\in\mathcal{C}_{\ell,t}, CHIME adjusts its value score according to its reuse count: v~i=vi​nini+λ\widetilde{v}_{i}=v_{i}\frac{n_{i}}{n_{i}+\lambda}, where nin_{i} is the reuse count of mim_{i}, and λ\lambda is a smoothing constant. The factor shrinks the value estimates of rarely reused memories toward zero, so that they fall back to similarity-only ranking. Semantic similarity does not indicate whether a memory actually helps, so CHIME reranks the candidates by their credited value and retrieves the Top-KK:

ℳℓ,tret=TopKmi∈𝒞ℓ,t​sim​(qℓ,t,ci)​[1+clip⁡(α​v~i,−β,β)],\mathcal{M}_{\ell,t}^{\mathrm{ret}}=\underset{m_{i}\in\mathcal{C}_{\ell,t}}{\mathrm{TopK}}\;\mathrm{sim}(q_{\ell,t},c_{i})\left[1+\mathrm{clip}\left(\alpha\widetilde{v}_{i},-\beta,\beta\right)\right], (5)

where α\alpha is the value weight, and β\beta bounds the value effect so that similarity remains the dominant ranking factor. As a result, memories validated as helpful in past episodes are preferred, whereas misleading ones are suppressed. The retrieved memories guide the corresponding planning or execution stage.

Credit Attribution Gate.

After planning and execution with the retrieved memories, the agent obtains the task outcome sts_{t}. The credit attribution gate then attributes this outcome to planning, execution, or external factors. CHIME implements the gate as a structured self-reflection of the frozen policy: the model is prompted to review the task, plan, trajectory, outcome, and retrieved memories, assess plan sufficiency, execution correctness, and the effects of external factors, and produce a diagnosis

gt=(at,{(cℓ,tnew,eℓ,tnew)}ℓ,γt,ℳtmis),g_{t}=\left(a_{t},\{(c_{\ell,t}^{\mathrm{new}},e_{\ell,t}^{\mathrm{new}})\}_{\ell},\gamma_{t},\mathcal{M}_{t}^{\mathrm{mis}}\right), (6)

where the attribution label ata_{t} takes one of four values: planning (e.g., an inadequate plan), execution (e.g., incorrect execution of a sufficient plan), both (e.g., distinct issues at both stages), and none (e.g., external causes, insufficient evidence, or no reusable experience); in successful episodes, it instead identifies the stage credited with the success; (cℓ,tnew,eℓ,tnew)(c_{\ell,t}^{\mathrm{new}},e_{\ell,t}^{\mathrm{new}}) is the task scenario and experience generated for each stage ℓ∈{plan,exec}\ell\in\{\mathrm{plan},\mathrm{exec}\}; γt∈[0,1]\gamma_{t}\in[0,1] is the attribution confidence; and ℳtmis\mathcal{M}_{t}^{\mathrm{mis}} contains retrieved memories identified as misleading. These outputs determine the subsequent memory value and content updates.

Credit-Aware Memory Evolution.

CHIME then uses the diagnosis gtg_{t} to evolve the memory bank in two steps, both directed at the attributed stages: (1)  Memory value evolution: CHIME evaluates each retrieved memory mim_{i} with a value feedback δi,t\delta_{i,t}: the memory is rewarded when its stage is credited for a successful outcome, penalized when it is identified as misleading, and unchanged otherwise. The attribution label determines the memory stages to update:

ℒ⁡(at)={{plan},at=planning,{exec},at=execution,{plan,exec},at=both,∅,at=none.\mathcal{L}(a_{t})=\begin{cases}\{\mathrm{plan}\},&a_{t}=\mathrm{planning},\\ \{\mathrm{exec}\},&a_{t}=\mathrm{execution},\\ \{\mathrm{plan},\mathrm{exec}\},&a_{t}=\mathrm{both},\\ \emptyset,&a_{t}=\mathrm{none}.\end{cases} (7)

The feedback for memory mim_{i} in stage ℓ⁡(i)\ell(i) is then

δi,t={−γt,mi∈ℳtmis,+γt,st=1​ and ​ℓ​(i)∈ℒ⁡(at),0,otherwise,\delta_{i,t}=\begin{cases}-\gamma_{t},&m_{i}\in\mathcal{M}_{t}^{\mathrm{mis}},\\ +\gamma_{t},&s_{t}=1\text{ and }\ell(i)\in\mathcal{L}(a_{t}),\\ 0,&\text{otherwise},\end{cases} (8)

The feedback magnitude is the attribution confidence γt\gamma_{t}. CHIME updates the value score and reuse count using an online average:

vi←clip⁡(ni​vi+δi,tni+1,−1,1),ni←ni+1.v_{i}\leftarrow\mathrm{clip}\!\left(\frac{n_{i}v_{i}+\delta_{i,t}}{n_{i}+1},-1,1\right),\qquad n_{i}\leftarrow n_{i}+1. (9)

The updated viv_{i} affects subsequent retrieval through Eq. 5. To maintain the quality of the memory bank, CHIME removes a memory when it has been reused sufficiently often yet remains low-valued, i.e., ni≥nminn_{i}\geq n_{\min} and vi<θprunev_{i}<\theta_{\mathrm{prune}}. (2) Memory content evolution: For each attributed stage ℓ∈ℒ⁡(at)\ell\in\mathcal{L}(a_{t}), CHIME updates the corresponding bank with the new experience in one of three ways: merging it into a sufficiently similar memory, inserting it as a new memory item, or directly discarding it.

{merge ​eℓ,tnew​ into ​mi∗,if accepted and similar,insert ​(cℓ,tnew,eℓ,tnew,0,0),if accepted but not similar,discard the experience,otherwise,\begin{cases}\text{merge }e_{\ell,t}^{\mathrm{new}}\text{ into }m_{i^{*}},&\text{if accepted and similar},\\ \text{insert }(c_{\ell,t}^{\mathrm{new}},e_{\ell,t}^{\mathrm{new}},0,0),&\text{if accepted but not similar},\\ \text{discard the experience},&\text{otherwise},\end{cases} (10)

where mi∗m_{i^{*}} is the most similar memory in the corresponding memory bank, with similarity defined as sim⁡(cℓ,tnew,ci∗)>θmerge\mathrm{sim}(c_{\ell,t}^{\mathrm{new}},c_{i^{*}})>\theta_{\mathrm{merge}}; acceptance requires γt≥θconf\gamma_{t}\geq\theta_{\mathrm{conf}} and a concrete task scenario with actionable content. A newly inserted memory starts with a neutral value and zero reuse count, so its value is estimated from future reuse rather than inherited from the gate. Applying this update only to the attributed stages yields ℳt+1\mathcal{M}_{t+1} and limits cross-stage memory contamination.

4 Experimental Setup

τ2\tau^{2}-Bench VitaBench BrowseComp-ZH BFCL-v4 Average
Method Train Eval Train Eval Train Eval Train Eval Train Eval
Qwen3.5-Flash
Qwen3.5-Flash No-Plan 21.68 22.32 9.40 8.89 20.20 21.25 71.33 70.86 30.65 30.83
WebAnchor (Yu et al., 2026) 24.77 23.45 14.76 11.94 17.68 21.98 69.36 70.26 31.64 31.91
TodoEvolve (Liu et al., 2026) 20.13 21.28 10.60 6.11 20.88 20.51 69.75 70.20 30.34 29.53
A-MapReduce (Chen et al., 2026) 19.81 20.55 12.02 6.67 18.35 18.31 73.84 72.45 31.01 29.50
CHIME (Our Method) 26.25 25.92 16.19 15.83 23.74 23.81 73.90 73.91 35.02 34.87
DeepSeek-V4-Flash
DeepSeek-V4-Flash No-Plan 31.42 32.86 16.31 11.11 26.94 23.81 63.75 64.15 34.61 32.98
WebAnchor Yu et al. (2026) 31.30 29.61 22.62 17.50 29.46 24.54 58.27 60.88 35.41 33.13
TodoEvolve Liu et al. (2026) 28.45 26.18 18.57 15.56 26.26 22.34 56.29 58.32 32.39 30.60
A-MapReduce Chen et al. (2026) 30.55 30.95 21.07 15.83 27.61 24.54 69.18 69.98 37.10 35.33
CHIME (Our Method) 33.15 36.28 22.98 21.11 29.97 31.50 65.89 67.15 38.00 39.01
Table 1: Results (%) averaged over three random instance orderings (Avg@3) across two backbone models and four benchmarks. Average is computed across benchmarks within each split. Bold indicates the best result in each column under each backbone.
τ2\tau^{2}-Bench VitaBench BrowseComp-ZH BFCL-v4 Average
Method Train Eval Train Eval Train Eval Train Eval Train Eval
CHIME 26.25 25.92 16.19 15.83 23.74 23.81 73.90 73.91 35.02 34.87
Hierarchical Memory Design
   w/o Plan Memory 22.96 23.11 15.95 12.22 20.71 20.51 69.52 70.60 32.29↓2.7 31.61↓3.3
   w/o Execution Memory 24.38 25.14 15.47 12.50 20.88 9.16 72.93 73.42 33.42↓1.6 30.06↓4.8
   w/o All Memory 24.05 23.67 13.22 12.50 19.70 19.78 69.52 70.99 31.62↓3.4 31.74↓3.1
   w/o Hierarchical Bank 24.87 22.76 15.60 11.11 21.04 19.41 72.63 71.79 33.54↓1.5 31.27↓3.6
Credit Attribution and Memory Evolution
   w/o Attribution Gate 21.21 22.28 17.02 12.50 20.20 22.71 73.40 72.23 32.96↓2.1 32.43↓2.4
   w/o Credit Rerank 24.32 24.75 16.31 10.83 21.72 23.44 73.75 72.28 34.03↓1.0 32.83↓2.0
   w/o Memory Evolution 25.61 24.32 16.79 8.61 20.71 22.34 73.79 72.05 34.23↓0.8 31.83↓3.0
Table 2: Ablation study results (%) with Qwen3.5-Flash, averaged over three random orderings (Avg@3). Average is computed across benchmarks within each split; downward values indicate drops from CHIME. Best results are bolded.

Baselines.

We compare CHIME with four baselines, ordered by increasing proximity to our setting: (1) No-Plan, a standard LLM agent without explicit planning or persistent memory; (2) WebAnchor, which improves a frozen policy through test-time search over sampled candidate plans (Yu et al., 2026; Team et al., 2026b); (3) TodoEvolve, which trains a meta-planner to generate task-specific planning architectures but accumulates no experience (Liu et al., 2026); and (4) A-MapReduce, the closest baseline to ours, which distills reusable insights from trajectories but updates its memory directly from task outcomes (Chen et al., 2026). All methods use the same benchmark runtimes, tools, and verifiers.

Benchmarks.

We evaluate CHIME on four long-horizon agent benchmarks: (1) τ2\tau^{2}-bench, multi-turn customer-service interactions (Barres et al., 2025); (2) VitaBench, multi-turn life-service tasks with changing user intent and large tool spaces (He et al., 2025); (3) BrowseComp-ZH, long-horizon InfoSeeking (Zhou et al., 2025); and (4) BFCL-v4, function calling, using Live and long-horizon Agentic groups (Patil et al., 2025).

Backbones.

We evaluate CHIME with two foundation model backbones: Qwen3.5-Flash model (35B total, 3B active) (Qwen Team, 2026a) and DeepSeek-V4-Flash model (284B total, 13B active) (DeepSeek-AI et al., 2026). The same backbone serves all roles in CHIME, including the planner, the executor, and the Credit Attribution Gate.

Test-Time Evaluation Protocol.

Following previous works (Ye et al., 2026; Ouyang et al., 2026), we process each benchmark as a sequential task stream and split it 7:3 into a train split and an eval split along the stream order. For benchmarks with multiple domains, we truncate each domain to the smallest domain size and split each domain independently, so that all domains contribute equally. The train split measures accumulation: memory-based methods update their memory from outcomes on this stream, and train performance reflects the utility of the accumulated experience for subsequent tasks. The eval split measures generalization: we freeze the memory banks accumulated on the train split and solve each eval task by directly retrieving the accumulated memory, without any further memory update, so eval performance reflects transfer of the accumulated memory to unseen tasks rather than further adaptation. To reduce variance from sample ordering, we run the protocol under three random orderings and report Avg@3 performance.

Implementation Details.

All evaluation-side LLM roles use Qwen3.6-Flash (Qwen Team, 2026b) under the official benchmark protocols (Barres et al., 2025; He et al., 2025; Zhou et al., 2025). Memory and τ2\tau^{2}-bench knowledge retrieval use Qwen3-Embedding-0.6B (Zhang et al., 2025). For retrieval and reranking, we set (N,K)=(20,3)(N,K)=(20,3), λ=3.0\lambda=3.0, α=1.0\alpha=1.0, and β=0.2\beta=0.2. For memory evolution, we use θconf=0.5\theta_{\mathrm{conf}}=0.5, θmerge=0.88\theta_{\mathrm{merge}}=0.88, nmin=3n_{\min}=3, and θprune=−0.1\theta_{\mathrm{prune}}=-0.1. All LLM calls use temperature 0.00.0, except WebAnchor plan generation (0.60.6). Further settings are provided in Appendix.

5 Main Experimental Results

We analyze the main results (Table 1) in three parts: memory accumulation on the train split, memory transfer to unseen tasks on the eval split, and ablations of CHIME’s key designs.

Memory Accumulation (Train Split).

CHIME achieves the highest average with both backbones (35.02% on Qwen3.5-Flash and 38.00% on DeepSeek-V4-Flash), outperforming the strongest baseline by 3.38% and 0.90%. These gains show that CHIME’s accumulated memory guides subsequent tasks more effectively than baseline memory.

Memory Transfer (Eval Split).

With memory banks frozen, CHIME achieves the highest eval average with both backbones: it outperforms the strongest baseline by 2.96% on Qwen3.5-Flash and 3.68% on DeepSeek-V4-Flash. Compared with A-MapReduce, the closest baseline to ours, CHIME leads by 5.37% on Qwen3.5-Flash and 3.68% on DeepSeek-V4-Flash. These gains show that attributed memory transfers better than outcome-based memory.

Ablation Studies.

We group the ablations into two parts: hierarchical memory design, and credit attribution and memory evolution. Table 2 reports the results with Qwen3.5-Flash: (1) Hierarchical memory design: On the eval split, removing planning memory, execution memory, or both reduces the average performance by 3.26%, 4.81%, and 3.13%, respectively. Moreover, collapsing the two banks into a shared one (w/o Hierarchical Bank) reduces it by 3.60%. These results show that CHIME’s gains stem from stage-specific experience and its explicit separation, rather than from memory accumulation alone. (2) Credit attribution and memory evolution: Removing the Credit Attribution Gate (w/o Attribution Gate) reverts memory updates to direct outcome feedback and reduces the eval average by 2.44%. Besides, disabling value-based reranking (w/o Credit Rerank) and credit-aware memory evolution (w/o Memory Evolution) reduces it by 2.04% and 3.04%, respectively. These results indicate that attribution decides which stage receives feedback, while value-based reranking and credit-aware evolution decide which memory is reused and retained.

6 Analysis

This section investigates the sources of CHIME’s effectiveness through six research questions: memory quality (RQ1), evolution efficiency (RQ2), memory value utility (RQ3), credit attribution gate reliability (RQ4), gains from stronger credit models (RQ5), and cross-backbone transfer (RQ6).

RQ1: Does CHIME Accumulate High-Quality Memory?

We use GLM-5.2 as an external judge to evaluate every memory in the Qwen3.5-Flash train-split banks. Each memory receives three binary judgments: Correctness (valid stage-specific guidance), Non-Redundancy (no duplicated experience), and Generality (applicability beyond the source episode). Table 3 shows that CHIME achieves the highest pass rate on all three dimensions, with 81.04% of its memories passing all of them. Removing the Gate floods the bank with noisy outcome-level traces, dropping non-redundancy and generality below 31%; A-MapReduce’s insights reach only 36.08% correctness, as they are induced without credited trajectories. These results show that credit-aware construction yields valid, distinct, and reusable memories, supporting the attribute-before-memorize principle.

Method Corr. Non-R. Gen. All
CHIME 81.20 93.60 90.13 81.04
A-MapReduce 36.08 81.65 75.32 36.08
w/o Attribution Gate 37.09 25.77 30.53 21.38
Table 3: Memory quality on the Qwen3.5-Flash train split under GLM-5.2 evaluation (%). Corr., Non-Red., and Gen. denote correctness, non-redundancy, and generality, respectively; All denotes the fraction passing all three criteria.
Figure 3: Memory evolution efficiency on τ2\tau^{2}-bench. Bars show memory-bank size, and lines show task accuracy at four accumulation checkpoints.

RQ2: Does Stage-Attributed Evolution Improve Memory Efficiency?

We track accuracy against memory-bank size on τ2\tau^{2}-bench along the train stream. We freeze the memory bank at four accumulation checkpoints (25%, 50%, 75%, and 100% progress of the stream) for measurement. We compare CHIME with two outcome-based baselines: Outcome-based Evolution (w/o Attribution Gate) Ye et al. (2026) and A-MapReduce, which update memory directly from final outcomes. As shown in Figure 3, CHIME improves accuracy as accumulation proceeds while keeping the smallest bank. At the end of the stream, CHIME retains only 129 memories, compared with 3,585 for A-MapReduce, yet achieves the highest accuracy. In contrast, both counterparts keep growing their banks, but their accuracy peaks at the 75% checkpoint and declines at full accumulation: without attribution, misleading memories are written into the bank and accumulate, which increasingly disturbs retrieval. These results show that effective self-evolving memory depends on the quality of the accumulated memories, not their quantity.

Figure 4: Downstream task accuracy by memory value level. Bars report accuracy for planning and execution memories in the low-, mid-, and high-value groups.

RQ3: Do Learned Memory Values Reflect Downstream Utility?

We partition the retrieved planning and execution memories into low-, mid-, and high-value groups by evenly splitting their value range, and report the average task accuracy of each group. Figure 4 shows a monotonic increase at both stages: from low to high value, accuracy rises from 21.7% to 50.8% for planning memories and from 23.7% to 35.4% for execution memories. This correlation suggests that the learned values faithfully reflect memory quality, indicating that the Credit Attribution Gate attributes credit effectively. Notably, higher-value planning memories bring larger accuracy gains than their execution counterparts (29.1% vs. 11.7% between the low- and high-value groups), whereas low-value planning memories are even less useful than low-value execution ones (21.7% vs. 23.7%), underscoring the importance of planning in long-horizon tasks. Together with the value-reranking ablation in Table 2, these results support selecting reusable experience by learned values.

RQ4: Is the Credit Attribution Gate Reliable?

We assess the reliability of the Credit Attribution Gate on Qwen3.5-Flash along two dimensions: Stability, the agreement between the original attribution and the majority label of three samples; and Rescued, the fraction of failed episodes whose attributed failure no longer persists when the task is rerun with only the generated memory. As shown in Table 4, planning and execution labels reach an agreement between 87.8% and 97.9% on both benchmarks, and between 56.3% and 67.6% of the failures are rescued, indicating that the Gate assigns repeatable credit and that the generated experience effectively corrects the attributed failure.

Benchmark Stage Stability Rescued
VitaBench planning 91.7 67.6
execution 97.9 60.2
BFCL-v4 planning 87.8 64.7
execution 92.9 56.3
Table 4: Gate attribution reliability (%) on VitaBench and BFCL-v4. Stability: agreement between the original attribution and the majority label of three samples. Rescued: fraction of failed episodes whose attributed failure no longer occurs when the task is rerun with only the generated memory.

RQ5: Does Stronger Attribution Further Improve Performance?

The Credit Attribution Gate relies on the backbone’s self-reflection capability. We replace it with a stronger model, Qwen3.7-Flash, for both credit attribution and experience generation, and evaluate the resulting memory on the eval split. Table 5 shows consistent gains on all three benchmarks, from 2.25% to 3.12%. These gains indicate that more accurate attribution yields more effective experience, so CHIME benefits directly from stronger credit models.

Benchmark Self-Reflect Qwen3.7-Flash Δ\Delta
τ2\tau^{2}-Bench 23.80 26.92 +3.12
VitaBench 12.50 15.00 +2.50
BFCL-v4 72.45 74.70 +2.25
Table 5: Eval accuracy (%). Qwen3.5-Flash model is replaced by stronger Qwen3.7-Flash for credit attribution gate. Results are from a single run rather than Avg@3.

RQ6: Does CHIME’s Memory Transfer Across Backbones?

We copy the memory accumulated by a source backbone on a benchmark’s train split into a target backbone’s frozen eval run. As shown in Table 6, w/o Mem. denotes the target running without memory, and Self denotes the target with its self-accumulated memory. CHIME’s transferred memory outperforms A-MapReduce’s in all four settings by between 2.25% and 4.68%, and closely approaches the upper bound: matching or exceeding it on the Qwen3.5-Flash target model, and staying within 1.72% on BFCL-v4 for the DeepSeek-V4-Flash target model. These experimental results reveal that attributed memory thus encodes reusable guidance rather than backbone-specific patterns.

Benchmark w/o Mem. A-Map. CHIME Self
DeepSeek-V4-Flash →\to Qwen3.5-Flash
τ2\tau^{2}-Bench 24.45 21.72 26.40 25.92
BFCL-v4 70.86 71.66 73.91 73.91
Qwen3.5-Flash →\to DeepSeek-V4-Flash
τ2\tau^{2}-Bench 30.04 28.22 30.82 36.28
BFCL-v4 60.13 61.85 65.43 67.15
Table 6: Cross-model memory transfer on the eval split (%). Each group header denotes the source →\to target backbones.

7 Conclusion

In this work, we presented CHIME, an self-evolving agentic planning framework based on our proposed attribute-before-memorize principle, which separates and updates planning and execution memory bank through stage-level credit attribution. Experiments on four long-horizon agent benchmarks show that CHIME consistently outperforms strong training-based and self-evolving memory baselines. Further analyses validate CHIME along six dimensions: memory quality, evolution efficiency, memory value utility, credit attribution gate reliability, credit-model scaling, and cross-backbone transfer. These results suggest that reliable memory evolution depends not only on what agents remember, but also on whether each experience is attributed to the decision stage where it can provide effective guidance.

References

  • Barres et al. (2025) Barres, V.; Dong, H.; Ray, S.; Si, X.; and Narasimhan, K. 2025. τ2\tau^{2}-Bench: Evaluating Conversational Agents in a Dual-Control Environment. arXiv:2506.07982.
  • Chen et al. (2026) Chen, M.; Zhang, G.; Chang, H.; Guo, Y.; and Zhou, S. 2026. A-MapReduce: Executing Wide Search via Agentic MapReduce. arXiv:2602.01331.
  • DeepSeek-AI et al. (2026) DeepSeek-AI; et al. 2026. DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence. arXiv:2606.19348.
  • Erdogan et al. (2025) Erdogan, L. E.; Lee, N.; Kim, S.; Moon, S.; Furuta, H.; Anumanchipalli, G.; Keutzer, K.; and Gholami, A. 2025. Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks. arXiv:2503.09572.
  • Guo, Kingston, and Kavraki (2024) Guo, W.; Kingston, Z.; and Kavraki, L. E. 2024. CaStL: Constraints as Specifications through LLM Translation for Long-Horizon Task and Motion Planning. arXiv:2410.22225.
  • Hao et al. (2023) Hao, S.; Gu, Y.; Ma, H.; Hong, J. J.; Wang, Z.; Wang, D. Z.; and Hu, Z. 2023. Reasoning with Language Model is Planning with World Model. arXiv:2305.14992.
  • He et al. (2025) He, W.; Sun, Y.; Hao, H.; Hao, X.; Xia, Z.; Gu, Q.; Han, C.; Zhao, D.; Su, H.; Zhang, K.; Gao, M.; Su, X.; Cai, X.; Cai, X.; Yang, Y.; and Zhao, Y. 2025. VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications. arXiv:2509.26490.
  • Kagaya et al. (2024) Kagaya, T.; Yuan, T. J.; Lou, Y.; Karlekar, J.; Pranata, S.; Kinose, A.; Oguri, K.; Wick, F.; and You, Y. 2024. RAP: Retrieval-Augmented Planning with Contextual Memory for Multimodal LLM Agents. arXiv:2402.03610.
  • Kim et al. (2023) Kim, S.; Moon, S.; Tabrizi, R.; Lee, N.; Mahoney, M. W.; Keutzer, K.; and Gholami, A. 2023. An LLM Compiler for Parallel Function Calling. arXiv:2312.04511.
  • Lan et al. (2026) Lan, T.; Henry, F.; Zhu, B.; Jia, Q.; Ren, J.; PU, Q.; Li, H.; Wang, L.; Xu, Z.; and Luo, W. 2026. Table-as-Search: Agentic Information Seeking is Table Completion. In Liakata, M.; Moreira, V. P.; Zhang, J.; and Jurgens, D., eds., Findings of the Association for Computational Linguistics: ACL 2026, 15088–15107. San Diego, California, United States: Association for Computational Linguistics. ISBN 979-8-89176-395-1.
  • Lan et al. (2025) Lan, T.; Zhu, B.; Jia, Q.; Ren, J.; Li, H.; Wang, L.; Xu, Z.; Luo, W.; and Zhang, K. 2025. DeepWideSearch: Benchmarking Depth and Width in Agentic Information Seeking. arXiv:2510.20168.
  • Li et al. (2025) Li, X.; Jiao, W.; Jin, J.; Dong, G.; Jin, J.; Wang, Y.; Wang, H.; Zhu, Y.; Wen, J.-R.; Lu, Y.; and Dou, Z. 2025. DeepAgent: A General Reasoning Agent with Scalable Toolsets. arXiv:2510.21618.
  • Liu et al. (2026) Liu, J.; Jiang, Y.; Zhang, G.; Zhang, Z.; Chang, H.; Yin, Z.; Ren, Q.; and Yan, J. 2026. TodoEvolve: Learning to Architect Agent Planning Systems. arXiv:2602.07839.
  • Ouyang et al. (2026) Ouyang, S.; Yan, J.; Hsu, I.-H.; Chen, Y.; Jiang, K.; Wang, Z.; Han, R.; Le, L. T.; Daruki, S.; Tang, X.; Tirumalashetty, V.; Lee, G.; Rofouei, M.; Lin, H.; Han, J.; Lee, C.-Y.; and Pfister, T. 2026. ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory. arXiv:2509.25140.
  • Patil et al. (2025) Patil, S. G.; Mao, H.; Cheng-Jie Ji, C.; Yan, F.; Suresh, V.; Stoica, I.; and E. Gonzalez, J. 2025. The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models. In Forty-second International Conference on Machine Learning.
  • Qwen Team (2026a) Qwen Team. 2026a. Qwen3.5: Towards Native Multimodal Agents.
  • Qwen Team (2026b) Qwen Team. 2026b. Qwen3.6-35B-A3B: Agentic Coding Power, Now Open to All.
  • Shen et al. (2023) Shen, Y.; Song, K.; Tan, X.; Li, D.; Lu, W.; and Zhuang, Y. 2023. HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in HuggingFace. In Advances in Neural Information Processing Systems.
  • Si et al. (2026) Si, S.; Zhao, H.; Luo, K.; Chen, G.; Qi, F.; Zhang, M.; Chang, B.; and Sun, M. 2026. A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Tasks. arXiv:2510.05608.
  • Team et al. (2026a) Team, C.; et al. 2026a. MiMo-V2-Flash Technical Report. arXiv:2601.02780.
  • Team et al. (2026b) Team, M.; Bai, S.; Bing, L.; Lei, L.; Li, R.; Li, X.; Lin, X.; Min, E.; Su, L.; Wang, B.; Wang, L.; Wang, L.; Wang, S.; Wang, X.; Zhang, Y.; Zhang, Z.; Chen, G.; Chen, L.; Cheng, Z.; Deng, Y.; Huang, Z.; Ng, D.; Ni, J.; Ren, Q.; Tang, X.; Wang, B. L.; Wang, H.; Wang, N.; Wei, C.; Wu, Q.; Xia, J.; Xiao, Y.; Xu, H.; Xu, X.; Xue, C.; Yang, Z.; Yang, Z.; Ye, F.; Ye, H.; Yu, J.; Zhang, C.; Zhang, W.; Zhao, H.; and Zhu, P. 2026b. MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification. arXiv:2603.15726.
  • Wang et al. (2025) Wang, Y.; Takanobu, R.; Liang, Z.; Mao, Y.; Hu, Y.; McAuley, J.; and Wu, X. 2025. Mem-α\alpha: Learning Memory Construction via Reinforcement Learning. arXiv:2509.25911.
  • Wang et al. (2024) Wang, Z.; Cai, S.; Chen, G.; Liu, A.; Ma, X.; and Liang, Y. 2024. Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents. arXiv:2302.01560.
  • Wong et al. (2025) Wong, R.; Wang, J.; Zhao, J.; Chen, L.; Gao, Y.; Zhang, L.; Zhou, X.; Wang, Z.; Xiang, K.; Zhang, G.; Huang, W.; Wang, Y.; and Wang, K. 2025. WideSearch: Benchmarking Agentic Broad Info-Seeking. arXiv:2508.07999.
  • Wu et al. (2025) Wu, R.; Wang, X.; Mei, J.; Cai, P.; Fu, D.; Yang, C.; Wen, L.; Yang, X.; Shen, Y.; Wang, Y.; and Shi, B. 2025. EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle. arXiv:2510.16079.
  • Xie et al. (2024) Xie, J.; Zhang, K.; Chen, J.; Zhu, T.; Lou, R.; Tian, Y.; Xiao, Y.; and Su, Y. 2024. TravelPlanner: A Benchmark for Real-World Planning with Language Agents. arXiv:2402.01622.
  • Yang et al. (2025) Yang, Y.; Lan, T.; Jia, Q.; Zhu, L.; Jiang, H.; Zhu, H.; Wang, L.; Luo, W.; and Zhang, K. 2025. HSCodeComp: A Realistic and Expert-level Benchmark for Deep Search Agents in Hierarchical Rule Application. arXiv:2510.19631.
  • Yao et al. (2023a) Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T. L.; Cao, Y.; and Narasimhan, K. 2023a. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv:2305.10601.
  • Yao et al. (2023b) Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; and Cao, Y. 2023b. ReAct: Synergizing Reasoning and Acting in Language Models. arXiv:2210.03629.
  • Ye et al. (2026) Ye, Y.; Jiang, H.; Jiang, F.; Lan, T.; Du, Y.; Fu, B.; Shi, X.; Jia, Q.; Wang, L.; and Luo, W. 2026. UMEM: Unified Memory Extraction and Management Framework for Generalizable Memory. arXiv:2602.10652.
  • Yu et al. (2025) Yu, H.; Chen, T.; Feng, J.; Chen, J.; Dai, W.; Yu, Q.; Zhang, Y.-Q.; Ma, W.-Y.; Liu, J.; Wang, M.; and Zhou, H. 2025. MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent. arXiv:2507.02259.
  • Yu et al. (2026) Yu, X.; Zhang, L.; Feng, X.; Jiang, Y.; Qin, B.; Xie, P.; and Zhou, J. 2026. WebAnchor: Anchoring Agent Planning to Stabilize Long-Horizon Web Reasoning. arXiv:2601.03164.
  • Zhang et al. (2026a) Zhang, S.; Wang, J.; Zhou, R.; Liao, J.; Feng, Y.; Li, Z.; Zheng, Y.; Zhang, W.; Wen, Y.; Li, Z.; Xiong, F.; Qi, Y.; Tang, B.; and Wen, M. 2026a. MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory. arXiv:2601.03192.
  • Zhang et al. (2026b) Zhang, Y.; Jiang, S.; Li, R.; Tu, J.; Su, Y.; Deng, L.; Guo, X.; Lv, C.; and Lin, J. 2026b. DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints. arXiv:2601.18137.
  • Zhang et al. (2025) Zhang, Y.; Li, M.; Long, D.; Zhang, X.; Lin, H.; Yang, B.; Xie, P.; Yang, A.; Liu, D.; Lin, J.; Huang, F.; and Zhou, J. 2025. Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models. arXiv:2506.05176.
  • Zhou et al. (2024) Zhou, A.; Yan, K.; Shlapentokh-Rothman, M.; Wang, H.; and Wang, Y.-X. 2024. Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models. arXiv:2310.04406.
  • Zhou et al. (2025) Zhou, P.; Leon, B.; Ying, X.; Zhang, C.; Shao, Y.; Ye, Q.; Chong, D.; Jin, Z.; Xie, C.; Cao, M.; Gu, Y.; Hong, S.; Ren, J.; Chen, J.; Liu, C.; and Hua, Y. 2025. BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese. arXiv:2504.19314.

Appendix A Implementation Details

Component Setting Value
Models and decoding Agent backbone qwen3.5-flash or deepseek-v4-flash
Evaluation-side LLM qwen3.6-flash
Default decoding Temperature 0.00.0; thinking disabled
WebAnchor plan generation Four candidates; temperature 0.60.6
Maximum input/output length 245,000245{,}000 / 16,38416{,}384 tokens
Request policy 2,4002{,}400-s timeout; at most 15 retries; 50-s base retry interval
Retrieval and reranking Embedding configuration Qwen3-Embedding-0.6B; batch size 16; maximum length 16,38416{,}384
Candidate/retrieved set sizes (N,K)=(20,3)(N,K)=(20,3) per memory layer
Candidate similarity threshold 0.30.3
Reranking parameters λ=3.0\lambda=3.0, α=1.0\alpha=1.0, and β=0.2\beta=0.2
Near-tie relevance margin 0.030.03
Memory evolution Update confidence threshold θconf=0.5\theta_{\mathrm{conf}}=0.5
Experience-addition threshold 0.10.1
Initial memory state (vi,ni)=(0,0)(v_{i},n_{i})=(0,0)
Merge threshold θmerge=0.88\theta_{\mathrm{merge}}=0.88
Pruning thresholds nmin=3n_{\min}=3 and θprune=−0.1\theta_{\mathrm{prune}}=-0.1
Memory capacity Unlimited
Evaluation protocol Train/eval split 7:37{:}3 within each benchmark domain
Selected memory checkpoint Epoch 2 for τ2\tau^{2}-Bench, VitaBench, and BFCL-v4; epoch 5 for BrowseComp-ZH
τ2\tau^{2}-Bench User-simulator LLM qwen3.6-flash
Interaction controls 100 steps; at most 10 errors; seed 42
Subprocess timeout 3,6003{,}600 s
Knowledge retrieval Local embedding retrieval; Top-88
VitaBench User-simulator LLM qwen3.6-flash
Agent/user configuration llm_agent/user_simulator
Benchmark language English
Interaction controls 100 steps; at most 10 errors; seed 42
BrowseComp-ZH Search framework smolagents
Search-agent backbone qwen3.5-flash or deepseek-v4-flash, matching the evaluated backbone
Search configuration Top-1010; at most 10 search steps
Search/page-fetch timeout 120 s each
Page-length limit 40,00040{,}000 characters
Answer judge Deterministic qwen3.6-flash
BFCL-v4 Evaluator qwen3.6-flash checker
Search configuration Top-1010; at most 10 tool/search steps
Shared execution Episode concurrency 16 for τ2\tau^{2}-Bench, BrowseComp-ZH, and BFCL-v4; 8 for VitaBench
Table 7: Implementation parameters for the main experiments. Benchmark-specific settings are shared across methods.

CHIME keeps the agent backbone frozen throughout, updating only the external planning and execution memory banks during accumulation and freezing them during transfer. These updates use only interaction trajectories and environment feedback, not reference answers. For consistent comparison, all methods share the same benchmark setup and evaluation-side models, and each baseline uses the best-performing configuration reported in its paper. All results are averaged over three random instance orderings (Avg@3).

We evaluate CHIME with Qwen3.5-Flash (Qwen Team, 2026a) and Deepseek-V4-Flash (DeepSeek-AI et al., 2026); in each setting, the same backbone handles planning, execution, and self-reflection in the Credit Attribution Gate. All evaluation-side LLM roles use Qwen3.6-Flash (Qwen Team, 2026b), while Qwen3-Embedding-0.6B (Zhang et al., 2025) serves as the retriever. Table 7 lists the complete implementation settings.

Appendix B Failure Analysis

Our analysis reveals two practical boundaries of CHIME.

Agent Capability.

CHIME augments the agent with memory guidance but leaves the backbone unchanged. As a result, even useful memory may not improve the outcome when the agent cannot follow the guidance or execute the required actions correctly. In the gate reliability analysis, rerunning failed episodes with the generated memory resolves 56.3–67.6% of attributed failures, but not all of them. The benefit of CHIME therefore remains bounded by the underlying capabilities of the agent.

Memory Transferability.

Our cross-backbone transfer results show that memory can transfer across backbones when the benchmark remains unchanged. Tasks within the same benchmark generally share the conditions under which a memory applies, including the available tools, valid actions, and success criteria. In contrast, we observe weaker transfer across benchmarks, where semantic similarity does not ensure that these conditions remain unchanged. CHIME’s memory is therefore most transferable across tasks with similar scenarios and operating conditions, while transfer across different benchmarks remains limited.

Appendix C Case Studies

We include three representative cases to illustrate how the memory bank is used and updated. The first case shows a planner-side benefit, where retrieved memories help decompose a complex request before execution. The second case shows an executor-side benefit, where retrieved memories help bind exact tool arguments. The third case shows the attribution gate itself: a failed trajectory is routed to the planner layer because the executor followed the plan and the reusable error lies in the planning abstraction.

Case Study: Planner Memory for Multi-Status Retail Requests Benchmark: τ2\tau^{2} Bench
Scores: Ours: 1.01.0  ∣\mid  AMapReduce-Mem: 0.00.0  ∣\mid  NoPlan: 0.00.0
 
Task Input:
The user wants to return the bookshelf and jigsaw puzzle received in the same delivered order; return the backpack from another delivered order that also contains a vacuum cleaner; modify a pending order by changing its shipping address to the user’s default Chicago address and changing the item color to red; and obtain the tracking number of a cancelled order. The user reveals information gradually and refers to different orders by item names rather than order IDs.
Retrieved Planner Memories: • Item-status mapping. For multi-item retail requests, explicitly map each item to its order ID and status before choosing an action. Pending orders should use modify_pending_order_items, while delivered orders should use return_delivered_order_items. • Cancelled-order handling. Before planning to reverse or operate on a cancelled order, first verify whether a supported tool exists; do not assume cancelled state changes are reversible. • One-shot pending modification. For pending-order modifications, collect all requested item/address changes into one call because the order may not allow incremental follow-up modifications. Ours Trajectory: • Step 1: Authenticate Lucas Brown and retrieve his order history. • Step 2: Separate the request into four status-specific subproblems: delivered-order return for bookshelf+jigsaw, delivered-order return for backpack, pending-order modification for address+color, and cancelled-order tracking lookup. • Step 3: Return the complete bookshelf+jigsaw item set from order ##W6239298 and return the backpack from order ##W9218746. • Step 4: Modify pending order ##W4860251 once, bundling the Chicago address update and red-item replacement into the same pending-order modification. • Step 5: Retrieve and report the tracking number 286422338955 for the cancelled order ##W1154986. Prediction:
Ours reports that it processed the two delivered-order returns, modified the pending order’s address and color, and provided the tracking number 286422338955 for the cancelled order. The official Tau2 verifier marks the episode correct.
Baseline Contrast:
AMapReduce-Mem produces a plausible plan, but its final answer reports the tracking number of the returned bookshelf/jigsaw order rather than treating the cancelled order as a distinct status. NoPlan also performs several local actions but fails the official verifier. The difference is that Ours uses planner memory to form the correct order-status decomposition before execution.
Figure 5: Planner memory helps the agent decompose a single natural-language request into the correct status-conditioned retail actions. The selected memories are useful before tool execution because they determine which objects should be grouped together and which tool family applies to each group.
Case Study: Executor Memory for Account-Level Net Credit Benchmark: τ2\tau^{2} Bench
Scores: Ours: 1.01.0  ∣\mid  AMapReduce-Mem: 0.00.0  ∣\mid  NoPlan: 0.00.0
 
Task Input:
Kim Junho has three checking accounts with nine ATM-fee discrepancies. The agent must verify the user, retrieve transaction histories for the Blue, Green, and Light Green accounts, cross-reference account-specific ATM-fee rules, identify both overcharges and missing fees, and apply the correct net credits: Blue Account $9.50, Green Account $9.00, and Light Green Account $1.50.
Retrieved Executor Memory:
When applying net credits to multiple bank accounts with mixed overcharges and undercharges, sum all overcharge refunds, sum all missing or undercharged fees, subtract the undercharges from the overcharges, and verify the final account-level amount before invoking the credit tool. Do not apply a single credit amount based only on overcharges or only on undercharges.
Ours Trajectory: • Step 1: Verify the user and unlock the account and transaction-history tools. • Step 2: Retrieve transaction histories for all three checking accounts: chk_kj93a7b2e1_1, chk_kj93a7b2e1_2, and chk_kj93a7b2e1_3. • Step 3: Compute account-level net credits rather than treating each discrepancy as an isolated dispute. • Step 4: Invoke apply_checking_account_credit_5829 with three account-specific amounts: $9.50, $9.00, and $1.50. • Step 5: Confirm that all discrepancies were corrected and that no further user action is needed. Prediction:
Ours summarizes the final credits for all three accounts and the official Tau2 database check is correct. The final state matches the required account-level net-credit updates.
Baseline Contrast:
NoPlan collapses the task to two disputes and reports only $4.50 in total credits, missing most required account-level corrections. AMapReduce-Mem also fails the official verifier. The useful distinction is execution-level: the plan can say “review all accounts and compute credits,” but reward depends on binding the exact account IDs and net amounts in the final tool calls.
Figure 6: Executor memory helps when the high-level plan is conceptually correct but success depends on precise action arguments. Here the memory guides account-level arithmetic and tool-argument binding, preventing the agent from collapsing several discrepancies into an incomplete dispute summary.
Case Study: Gate Attribution for a Strict Deadline Failure Benchmark: VitaBench
Outcome: Failed episode, verifier score 0.00.0 with partial rubric score 0.750.75
Gate Label: learn_layer=planner  ∣\mid  confidence 0.900.90  ∣\mid  external_issue=false
 
Task Input:
The user wants to order specific Sichuan dishes from a previously visited restaurant to a temporary work location. The food should arrive before a friend’s lunch break starts at 13:00, and it must respect low-oil and low-salt dietary constraints.
Agent Plan:
The plan correctly selects Xiao Sichuan (Shifan Street Branch), verifies the temporary work address Jinzheng Haiyue International, filters the menu for Boiling Fish and garlic-flavored vegetables, and schedules delivery around the lunch deadline.
Final Prediction / Verifier Feedback:
The final answer confirms that the order is set, but the official evaluator finds one critical violation: the estimated delivery time is 2024-09-12 13:00:00. The rubric requires delivery strictly before 13:00, so exactly 13:00 does not satisfy the constraint.
Gate Diagnosis: • Plan sufficient: false. The plan handled store and item selection, but failed to encode a strict-inequality time buffer. • Execution followed plan: true. The executor carried out the planned schedule; the problem was not a wrong tool argument relative to the plan. • External issue: false. The failure is reusable and not caused by an unavailable API or environment error. • Routed layer: planner. The missing lesson concerns how to plan around strict “before” deadlines. Memory Written by the Gate:
For delivery orders that require arrival before a specific deadline, calculate the dispatch time by subtracting the estimated shipping duration plus a safety buffer from the deadline, and verify that the resulting delivery time is strictly less than the deadline. Avoid setting a delivery time that lands exactly on the boundary.
Figure 7: The gate converts a failed trajectory into a layer-specific memory. Because the executor followed the plan and the remaining error was a strict temporal-constraint mistake, the episode writes a planner memory rather than polluting the executor memory bank.

Planner-side decomposition.

Figure 5 shows why planner memory cannot be replaced by a generic summary of past successes. The task contains several retail operations that look similar in natural language but require different protocols: returned items belong to delivered orders, color and address edits belong to a pending order, and the requested tracking number belongs to a cancelled order. The retrieved planner memories directly encode this object-status decomposition. As a result, the agent plans the correct action family for each object before making tool calls, while the baselines produce plausible but verifier-failing summaries.

Executor-side precision.

Figure 6 highlights a different failure mode. The high-level plan is not the main challenge: all methods can state that the agent should review accounts and correct ATM-fee errors. The reward depends on execution details: retaining three account identities, computing net credits after subtracting missing fees, and invoking the credit tool with the exact amount for each account. The executor memory is useful because it is written at the same granularity as the eventual action arguments.

Layer-aware attribution.

Figure 7 illustrates why the gate uses trajectory-level diagnosis rather than only the binary verifier outcome. The delivery episode fails, but it should not penalize executor memory: the selected restaurant, address, and item choices are mostly correct, and the executor follows the planned schedule. The reusable error is the planner’s treatment of a strict “before” constraint as if equality were acceptable. The gate therefore writes a planner memory about strict deadline buffers and leaves the executor bank untouched. This keeps the memory bank more specific and reduces cross-layer contamination.