CHIME: Credit-Aware Hierarchical Memory Evolution for
Long-Horizon Agentic Planning
Abstract
Planning is a central capability that enables agents to decompose complex long-horizon tasks into manageable steps. Test-time search and training-based methods improve planning but incur high inference costs or require expensive training data. Self-evolving memory instead accumulates reusable experience from agent interaction outcomes into an external memory bank, so planning capability keeps improving at inference time without parameter updates. However, existing self-evolving memory methods share an inherent credit assignment problem: they rely on final task outcomes as feedback, but such outcomes conflate plan quality with execution errors and environmental factors, so the accumulated planning experience is often biased and noisy. To address this problem, we propose Credit-Aware HIerarchical Memory Evolution (CHIME), a self-evolving memory framework that maintains a separate planning bank and execution bank and follows an attribute-before-memorize principle: CHIME first attributes each task outcome to the plan, the execution, both, or neither, and then updates only the corresponding memory bank. Extensive experiments on four long-horizon agent benchmarks show that CHIME consistently outperforms state-of-the-art training-based and self-evolving memory baselines. Further analyses reveal several interesting findings. For example, CHIME accumulates effective memory with far fewer items. In addition, the learned memory values faithfully reflect downstream utility: high-quality planning memories are more valuable than execution memories. Finally, the accumulated memory effectively transfers across backbone models. Code will be released at https://github.com/ATH-MaaS/Marco-DeepResearch.
1Institute of Artificial Intelligence, Xiamen University
2Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage
of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism
3Zhejiang University
4Alibaba Group
1 Introduction
Long-horizon tasks pose a defining challenge for agentic systems: agents must coordinate decisions across interdependent steps while sustaining adherence to task constraints throughout execution Xie et al. (2024); Wang et al. (2024); Erdogan et al. (2025). Agentic planning addresses these requirements by decomposing the overall objective into modular, independently executable subgoals Wang et al. (2024); Shen et al. (2023); Zhang et al. (2026b); Liu et al. (2026) and establishing the dependencies and constraints among them Kim et al. (2023); Guo, Kingston, and Kavraki (2024); Lan et al. (2026), making it a central capability for solving long-horizon tasks.
This has motivated prior work to improve agents’ planning capability. Test-time search Yao et al. (2023a); Hao et al. (2023); Zhou et al. (2024); Yu et al. (2026); Team et al. (2026a) improves per-task performance by exploring multiple candidate plans or trajectories, but incurs high inference costs and does not retain reusable experience. To retain and reuse planning experience across tasks, subsequent work has pursued two directions: (1) Planner model training Erdogan et al. (2025); Si et al. (2026); Liu et al. (2026); Team et al. (2026b) internalizes planning capability into model parameters, but requires costly data collection and post-training, since identifying high-quality plans relies on extensive downstream executions; (2) Self-evolving memory Chen et al. (2026); Ouyang et al. (2026) instead accumulates experience in an external, training-free memory bank, offering a more efficient and scalable alternative. We therefore focus on improving planning within the self-evolving memory paradigm in this paper.
However, existing self-evolving memory methods share a fundamental limitation for agentic planning: they equate downstream task outcomes with plan quality. The same outcome can arise from different causes: success may result from a sound plan or from executor recovery from a flawed one, while failure may stem from a flawed plan, incorrect execution, or environmental issues. Using such outcomes directly as memory feedback therefore leads to a fundamental credit assignment failure. On the one hand, once biased experience is written into the memory bank, it is repeatedly retrieved and reused across tasks, allowing local misjudgments to accumulate into systematic planning errors. On the other hand, execution- or environment-level experience may be mistaken for planning memory; retrieving might low-level details then misguides high-level planning decisions instead of providing planning guidance.
To address this, we propose Credit-Aware HIerarchical Memory Evolution (CHIME), a self-evolving agentic planning framework that replaces outcome-based memory updates Ye et al. (2026) with an attribute-before-memorize process. Specifically, CHIME comprises three components: (1) Hierarchical Memory Bank explicitly separates planning and execution memory into a planning bank and an execution bank, retrieving stage-specific memories through similarity retrieval and value-based reranking to separately guide the agent’s planning and execution; (2) Credit Attribution Gate then reflects on the task, plan, execution, and retrieved memories to attribute reusable feedback to planning, execution, both stages, or neither, before any memory is written; it also generates stage-specific experience and identifies misleading memories; and (3) Credit-Aware Memory Evolution updates the values of retrieved memories using confidence-weighted feedback, while filtering, merging, or inserting new experience only into the attributed bank. In this way, CHIME evolves each bank only with memory attributed to that stage, preventing biased signals from persisting and execution-level memory from contaminating the planning bank.
Experiments on four long-horizon benchmarks with two backbone models show that CHIME consistently outperforms training-based and self-evolving memory baselines, improving the eval average over the strongest baseline by 2.96% and 3.68% on the two backbones, respectively. Further analyses reveal several interesting findings: (1) CHIME achieves the highest accuracy while retaining only 129 memories, versus 3,585 for the strong baseline; (2) the learned memory values faithfully reflect downstream utility, and planning memories bring more than twice the accuracy gain of execution ones (21.7%50.8% vs. 23.7%35.4%); (3) the Credit Attribution Gate is reliable, with repeated attributions agreeing in up to 97.9% of cases and over half of the failures rescued by the generated experience; (4) replacing model’s self-reflection credit with a stronger credit model yields a further gain of up to 3.12%; and (5) the accumulated memory effectively transfers across backbones, outperforming A-MapReduce’s transferred memory by up to 4.68%. These findings confirm that attribute-before-memorize is the key to CHIME’s effectiveness.
2 Related Work
Self-evolving memory.
Self-evolving memory methods treat an external memory bank as the agent’s evolvable parameters, distilling interaction trajectories into reusable experience or skills Yu et al. (2025); Wang et al. (2025); Zhang et al. (2026a). Unlike training-based methods, they keep model parameters frozen and update only the memory bank, avoiding costly trajectory collection and post-training Li et al. (2025); Wu et al. (2025). For example, ReasoningBank Ouyang et al. (2026) distills generalizable reasoning strategies from self-judged successful and failed trajectories, and UMEM Ye et al. (2026) jointly optimizes memory extraction and management with reinforcement learning.
Agentic planning.
Agentic planning decomposes complex objectives into structured steps and coordinates their execution. It underlies a broad range of agent systems Yao et al. (2023b); Erdogan et al. (2025); Liu et al. (2026); Lan et al. (2026). Long-horizon tasks often involve many interdependent subtasks. Without an explicit plan to organize them, agents easily lose track of progress, a failure known as the lost-in-the-middle problem Lan et al. (2026); Liu et al. (2026). As recent benchmarks introduce increasingly challenging tasks, improving planning capability has become a central research problem for agent systems Wong et al. (2025); Lan et al. (2025); Yang et al. (2025).
Improving agentic planning.
Existing efforts fall into three paradigms: (1) Test-time search explores multiple candidate plans or trajectories at inference time and selects among them with value estimates or environmental feedback Yao et al. (2023a); Hao et al. (2023); Zhou et al. (2024). For example, MiroThinker-H1 Team et al. (2026b) and WebAnchor Yu et al. (2026) select high-quality plans via rejection sampling; (2) Planner model training internalizes planning capability into model parameters Erdogan et al. (2025); Si et al. (2026); Yu et al. (2026). For example, TodoEvolve Liu et al. (2026) constructs a modular design space of planning architectures and trains a meta-planner with reinforcement learning to synthesize task-specific planning systems. However, such methods require costly trajectory collection and post-training, and the resulting planner is expensive to update continually Liu et al. (2026); and (3) Self-evolving memory offers a training-free and continual alternative, but few works evolve memory specifically for agentic planning Kagaya et al. (2024). A representative is A-MapReduce Chen et al. (2026), which evolves structured hints from past executions to improve task decomposition and result aggregation in long-horizon agentic search. However, all three paradigms evaluate agent’s plans by final task outcomes, implicitly equating outcome with plan quality. This equivalence does not hold, as outcomes are also shaped by execution details and environmental factors.
3 Methodology of CHIME
3.1 Task Formulation of Self-Evolving Agents
We consider a self-evolving agent that solves a stream of tasks. The agent’s policy parameters remain frozen, and adaptation is carried by an external memory bank that evolves across episodes. At episode , the agent receives a task and retrieves task-relevant experience from . Conditioned on and , the frozen policy produces a trajectory and receives an outcome from the environment and the evaluator. A memory update operator then distills the episode into reusable experience:
| (1) |
During evaluation, the memory bank is frozen, so that performance reflects experience accumulated before evaluation. Existing self-evolving memory methods mainly use the final outcome as the supervision signal for memory updates. For agentic planning, this signal is biased: the outcome depends on both the plan and the execution process, so it does not faithfully reflect plan quality.
3.2 Overview of CHIME
We propose CHIME (Figure 2), an attribute-before-memorize self-evolving framework for agentic planning. It attributes each task outcome to planning, execution, or external factors, and updates only the corresponding memory bank, which reduces biased feedback and cross-stage memory contamination. CHIME consists of three components: a Hierarchical Memory Bank that separates planning and execution experience, a Credit Attribution Gate that attributes each outcome to its responsible stage, and Credit-Aware Memory Evolution that updates the banks from the attributed credit.
Hierarchical Memory Bank.
Unlike existing self-evolving methods Ye et al. (2026), CHIME maintains hierarchical memory banks
| (2) |
which comprises a planning bank and an execution bank, indexed by . The planning bank stores strategic experience (e.g., subtask decomposition, dependency modeling, and constraint incorporation), whereas the execution bank stores operational experience (e.g., tool selection and invocation). Each bank consists of memory items
| (3) |
where describes the tasks to which the memory applies, serving as the key for memory retrieval and merging; is the reusable experience memory distilled from past episodes; is a value score estimated from reuse feedback, which guides reranking (Eq. 5) and pruning; and is the reuse count, which calibrates the value score (Eq. 3.2). At episode , CHIME receives a long-horizon task and constructs a planning query to retrieve from . The planner then generates a plan from and . Next, CHIME constructs an execution query from and retrieves from . The executor uses and to produce a trajectory , and the environment returns a binary outcome indicating task success. Finally, the credit attribution gate produces a diagnosis from the plan, trajectory, outcome, and retrieved memories, which drives the transition from to . One CHIME episode is summarized as
| (4) |
These steps correspond to the left, right, and feedback-loop parts of Figure 2.
Hierarchical Memory Retrieval.
Retrieval operates independently at each stage . Given the stage-specific query , CHIME retrieves the Top- memories from the corresponding bank according to , yielding the candidate set . For each candidate memory , CHIME adjusts its value score according to its reuse count: , where is the reuse count of , and is a smoothing constant. The factor shrinks the value estimates of rarely reused memories toward zero, so that they fall back to similarity-only ranking. Semantic similarity does not indicate whether a memory actually helps, so CHIME reranks the candidates by their credited value and retrieves the Top-:
| (5) |
where is the value weight, and bounds the value effect so that similarity remains the dominant ranking factor. As a result, memories validated as helpful in past episodes are preferred, whereas misleading ones are suppressed. The retrieved memories guide the corresponding planning or execution stage.
Credit Attribution Gate.
After planning and execution with the retrieved memories, the agent obtains the task outcome . The credit attribution gate then attributes this outcome to planning, execution, or external factors. CHIME implements the gate as a structured self-reflection of the frozen policy: the model is prompted to review the task, plan, trajectory, outcome, and retrieved memories, assess plan sufficiency, execution correctness, and the effects of external factors, and produce a diagnosis
| (6) |
where the attribution label takes one of four values: planning (e.g., an inadequate plan), execution (e.g., incorrect execution of a sufficient plan), both (e.g., distinct issues at both stages), and none (e.g., external causes, insufficient evidence, or no reusable experience); in successful episodes, it instead identifies the stage credited with the success; is the task scenario and experience generated for each stage ; is the attribution confidence; and contains retrieved memories identified as misleading. These outputs determine the subsequent memory value and content updates.
Credit-Aware Memory Evolution.
CHIME then uses the diagnosis to evolve the memory bank in two steps, both directed at the attributed stages: (1) Memory value evolution: CHIME evaluates each retrieved memory with a value feedback : the memory is rewarded when its stage is credited for a successful outcome, penalized when it is identified as misleading, and unchanged otherwise. The attribution label determines the memory stages to update:
| (7) |
The feedback for memory in stage is then
| (8) |
The feedback magnitude is the attribution confidence . CHIME updates the value score and reuse count using an online average:
| (9) |
The updated affects subsequent retrieval through Eq. 5. To maintain the quality of the memory bank, CHIME removes a memory when it has been reused sufficiently often yet remains low-valued, i.e., and . (2) Memory content evolution: For each attributed stage , CHIME updates the corresponding bank with the new experience in one of three ways: merging it into a sufficiently similar memory, inserting it as a new memory item, or directly discarding it.
| (10) |
where is the most similar memory in the corresponding memory bank, with similarity defined as ; acceptance requires and a concrete task scenario with actionable content. A newly inserted memory starts with a neutral value and zero reuse count, so its value is estimated from future reuse rather than inherited from the gate. Applying this update only to the attributed stages yields and limits cross-stage memory contamination.
4 Experimental Setup
| -Bench | VitaBench | BrowseComp-ZH | BFCL-v4 | Average | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Method | Train | Eval | Train | Eval | Train | Eval | Train | Eval | Train | Eval |
| Qwen3.5-Flash | ||||||||||
| Qwen3.5-Flash No-Plan | 21.68 | 22.32 | 9.40 | 8.89 | 20.20 | 21.25 | 71.33 | 70.86 | 30.65 | 30.83 |
| WebAnchor (Yu et al., 2026) | 24.77 | 23.45 | 14.76 | 11.94 | 17.68 | 21.98 | 69.36 | 70.26 | 31.64 | 31.91 |
| TodoEvolve (Liu et al., 2026) | 20.13 | 21.28 | 10.60 | 6.11 | 20.88 | 20.51 | 69.75 | 70.20 | 30.34 | 29.53 |
| A-MapReduce (Chen et al., 2026) | 19.81 | 20.55 | 12.02 | 6.67 | 18.35 | 18.31 | 73.84 | 72.45 | 31.01 | 29.50 |
| CHIME (Our Method) | 26.25 | 25.92 | 16.19 | 15.83 | 23.74 | 23.81 | 73.90 | 73.91 | 35.02 | 34.87 |
| DeepSeek-V4-Flash | ||||||||||
| DeepSeek-V4-Flash No-Plan | 31.42 | 32.86 | 16.31 | 11.11 | 26.94 | 23.81 | 63.75 | 64.15 | 34.61 | 32.98 |
| WebAnchor Yu et al. (2026) | 31.30 | 29.61 | 22.62 | 17.50 | 29.46 | 24.54 | 58.27 | 60.88 | 35.41 | 33.13 |
| TodoEvolve Liu et al. (2026) | 28.45 | 26.18 | 18.57 | 15.56 | 26.26 | 22.34 | 56.29 | 58.32 | 32.39 | 30.60 |
| A-MapReduce Chen et al. (2026) | 30.55 | 30.95 | 21.07 | 15.83 | 27.61 | 24.54 | 69.18 | 69.98 | 37.10 | 35.33 |
| CHIME (Our Method) | 33.15 | 36.28 | 22.98 | 21.11 | 29.97 | 31.50 | 65.89 | 67.15 | 38.00 | 39.01 |
| -Bench | VitaBench | BrowseComp-ZH | BFCL-v4 | Average | ||||||
| Method | Train | Eval | Train | Eval | Train | Eval | Train | Eval | Train | Eval |
| CHIME | 26.25 | 25.92 | 16.19 | 15.83 | 23.74 | 23.81 | 73.90 | 73.91 | 35.02 | 34.87 |
| Hierarchical Memory Design | ||||||||||
| w/o Plan Memory | 22.96 | 23.11 | 15.95 | 12.22 | 20.71 | 20.51 | 69.52 | 70.60 | 32.29↓2.7 | 31.61↓3.3 |
| w/o Execution Memory | 24.38 | 25.14 | 15.47 | 12.50 | 20.88 | 9.16 | 72.93 | 73.42 | 33.42↓1.6 | 30.06↓4.8 |
| w/o All Memory | 24.05 | 23.67 | 13.22 | 12.50 | 19.70 | 19.78 | 69.52 | 70.99 | 31.62↓3.4 | 31.74↓3.1 |
| w/o Hierarchical Bank | 24.87 | 22.76 | 15.60 | 11.11 | 21.04 | 19.41 | 72.63 | 71.79 | 33.54↓1.5 | 31.27↓3.6 |
| Credit Attribution and Memory Evolution | ||||||||||
| w/o Attribution Gate | 21.21 | 22.28 | 17.02 | 12.50 | 20.20 | 22.71 | 73.40 | 72.23 | 32.96↓2.1 | 32.43↓2.4 |
| w/o Credit Rerank | 24.32 | 24.75 | 16.31 | 10.83 | 21.72 | 23.44 | 73.75 | 72.28 | 34.03↓1.0 | 32.83↓2.0 |
| w/o Memory Evolution | 25.61 | 24.32 | 16.79 | 8.61 | 20.71 | 22.34 | 73.79 | 72.05 | 34.23↓0.8 | 31.83↓3.0 |
Baselines.
We compare CHIME with four baselines, ordered by increasing proximity to our setting: (1) No-Plan, a standard LLM agent without explicit planning or persistent memory; (2) WebAnchor, which improves a frozen policy through test-time search over sampled candidate plans (Yu et al., 2026; Team et al., 2026b); (3) TodoEvolve, which trains a meta-planner to generate task-specific planning architectures but accumulates no experience (Liu et al., 2026); and (4) A-MapReduce, the closest baseline to ours, which distills reusable insights from trajectories but updates its memory directly from task outcomes (Chen et al., 2026). All methods use the same benchmark runtimes, tools, and verifiers.
Benchmarks.
We evaluate CHIME on four long-horizon agent benchmarks: (1) -bench, multi-turn customer-service interactions (Barres et al., 2025); (2) VitaBench, multi-turn life-service tasks with changing user intent and large tool spaces (He et al., 2025); (3) BrowseComp-ZH, long-horizon InfoSeeking (Zhou et al., 2025); and (4) BFCL-v4, function calling, using Live and long-horizon Agentic groups (Patil et al., 2025).
Backbones.
We evaluate CHIME with two foundation model backbones: Qwen3.5-Flash model (35B total, 3B active) (Qwen Team, 2026a) and DeepSeek-V4-Flash model (284B total, 13B active) (DeepSeek-AI et al., 2026). The same backbone serves all roles in CHIME, including the planner, the executor, and the Credit Attribution Gate.
Test-Time Evaluation Protocol.
Following previous works (Ye et al., 2026; Ouyang et al., 2026), we process each benchmark as a sequential task stream and split it 7:3 into a train split and an eval split along the stream order. For benchmarks with multiple domains, we truncate each domain to the smallest domain size and split each domain independently, so that all domains contribute equally. The train split measures accumulation: memory-based methods update their memory from outcomes on this stream, and train performance reflects the utility of the accumulated experience for subsequent tasks. The eval split measures generalization: we freeze the memory banks accumulated on the train split and solve each eval task by directly retrieving the accumulated memory, without any further memory update, so eval performance reflects transfer of the accumulated memory to unseen tasks rather than further adaptation. To reduce variance from sample ordering, we run the protocol under three random orderings and report Avg@3 performance.
Implementation Details.
All evaluation-side LLM roles use Qwen3.6-Flash (Qwen Team, 2026b) under the official benchmark protocols (Barres et al., 2025; He et al., 2025; Zhou et al., 2025). Memory and -bench knowledge retrieval use Qwen3-Embedding-0.6B (Zhang et al., 2025). For retrieval and reranking, we set , , , and . For memory evolution, we use , , , and . All LLM calls use temperature , except WebAnchor plan generation (). Further settings are provided in Appendix.
5 Main Experimental Results
We analyze the main results (Table 1) in three parts: memory accumulation on the train split, memory transfer to unseen tasks on the eval split, and ablations of CHIME’s key designs.
Memory Accumulation (Train Split).
CHIME achieves the highest average with both backbones (35.02% on Qwen3.5-Flash and 38.00% on DeepSeek-V4-Flash), outperforming the strongest baseline by 3.38% and 0.90%. These gains show that CHIME’s accumulated memory guides subsequent tasks more effectively than baseline memory.
Memory Transfer (Eval Split).
With memory banks frozen, CHIME achieves the highest eval average with both backbones: it outperforms the strongest baseline by 2.96% on Qwen3.5-Flash and 3.68% on DeepSeek-V4-Flash. Compared with A-MapReduce, the closest baseline to ours, CHIME leads by 5.37% on Qwen3.5-Flash and 3.68% on DeepSeek-V4-Flash. These gains show that attributed memory transfers better than outcome-based memory.
Ablation Studies.
We group the ablations into two parts: hierarchical memory design, and credit attribution and memory evolution. Table 2 reports the results with Qwen3.5-Flash: (1) Hierarchical memory design: On the eval split, removing planning memory, execution memory, or both reduces the average performance by 3.26%, 4.81%, and 3.13%, respectively. Moreover, collapsing the two banks into a shared one (w/o Hierarchical Bank) reduces it by 3.60%. These results show that CHIME’s gains stem from stage-specific experience and its explicit separation, rather than from memory accumulation alone. (2) Credit attribution and memory evolution: Removing the Credit Attribution Gate (w/o Attribution Gate) reverts memory updates to direct outcome feedback and reduces the eval average by 2.44%. Besides, disabling value-based reranking (w/o Credit Rerank) and credit-aware memory evolution (w/o Memory Evolution) reduces it by 2.04% and 3.04%, respectively. These results indicate that attribution decides which stage receives feedback, while value-based reranking and credit-aware evolution decide which memory is reused and retained.
6 Analysis
This section investigates the sources of CHIME’s effectiveness through six research questions: memory quality (RQ1), evolution efficiency (RQ2), memory value utility (RQ3), credit attribution gate reliability (RQ4), gains from stronger credit models (RQ5), and cross-backbone transfer (RQ6).
RQ1: Does CHIME Accumulate High-Quality Memory?
We use GLM-5.2 as an external judge to evaluate every memory in the Qwen3.5-Flash train-split banks. Each memory receives three binary judgments: Correctness (valid stage-specific guidance), Non-Redundancy (no duplicated experience), and Generality (applicability beyond the source episode). Table 3 shows that CHIME achieves the highest pass rate on all three dimensions, with 81.04% of its memories passing all of them. Removing the Gate floods the bank with noisy outcome-level traces, dropping non-redundancy and generality below 31%; A-MapReduce’s insights reach only 36.08% correctness, as they are induced without credited trajectories. These results show that credit-aware construction yields valid, distinct, and reusable memories, supporting the attribute-before-memorize principle.
| Method | Corr. | Non-R. | Gen. | All |
|---|---|---|---|---|
| CHIME | 81.20 | 93.60 | 90.13 | 81.04 |
| A-MapReduce | 36.08 | 81.65 | 75.32 | 36.08 |
| w/o Attribution Gate | 37.09 | 25.77 | 30.53 | 21.38 |
RQ2: Does Stage-Attributed Evolution Improve Memory Efficiency?
We track accuracy against memory-bank size on -bench along the train stream. We freeze the memory bank at four accumulation checkpoints (25%, 50%, 75%, and 100% progress of the stream) for measurement. We compare CHIME with two outcome-based baselines: Outcome-based Evolution (w/o Attribution Gate) Ye et al. (2026) and A-MapReduce, which update memory directly from final outcomes. As shown in Figure 3, CHIME improves accuracy as accumulation proceeds while keeping the smallest bank. At the end of the stream, CHIME retains only 129 memories, compared with 3,585 for A-MapReduce, yet achieves the highest accuracy. In contrast, both counterparts keep growing their banks, but their accuracy peaks at the 75% checkpoint and declines at full accumulation: without attribution, misleading memories are written into the bank and accumulate, which increasingly disturbs retrieval. These results show that effective self-evolving memory depends on the quality of the accumulated memories, not their quantity.
RQ3: Do Learned Memory Values Reflect Downstream Utility?
We partition the retrieved planning and execution memories into low-, mid-, and high-value groups by evenly splitting their value range, and report the average task accuracy of each group. Figure 4 shows a monotonic increase at both stages: from low to high value, accuracy rises from 21.7% to 50.8% for planning memories and from 23.7% to 35.4% for execution memories. This correlation suggests that the learned values faithfully reflect memory quality, indicating that the Credit Attribution Gate attributes credit effectively. Notably, higher-value planning memories bring larger accuracy gains than their execution counterparts (29.1% vs. 11.7% between the low- and high-value groups), whereas low-value planning memories are even less useful than low-value execution ones (21.7% vs. 23.7%), underscoring the importance of planning in long-horizon tasks. Together with the value-reranking ablation in Table 2, these results support selecting reusable experience by learned values.
RQ4: Is the Credit Attribution Gate Reliable?
We assess the reliability of the Credit Attribution Gate on Qwen3.5-Flash along two dimensions: Stability, the agreement between the original attribution and the majority label of three samples; and Rescued, the fraction of failed episodes whose attributed failure no longer persists when the task is rerun with only the generated memory. As shown in Table 4, planning and execution labels reach an agreement between 87.8% and 97.9% on both benchmarks, and between 56.3% and 67.6% of the failures are rescued, indicating that the Gate assigns repeatable credit and that the generated experience effectively corrects the attributed failure.
| Benchmark | Stage | Stability | Rescued |
|---|---|---|---|
| VitaBench | planning | 91.7 | 67.6 |
| execution | 97.9 | 60.2 | |
| BFCL-v4 | planning | 87.8 | 64.7 |
| execution | 92.9 | 56.3 |
RQ5: Does Stronger Attribution Further Improve Performance?
The Credit Attribution Gate relies on the backbone’s self-reflection capability. We replace it with a stronger model, Qwen3.7-Flash, for both credit attribution and experience generation, and evaluate the resulting memory on the eval split. Table 5 shows consistent gains on all three benchmarks, from 2.25% to 3.12%. These gains indicate that more accurate attribution yields more effective experience, so CHIME benefits directly from stronger credit models.
| Benchmark | Self-Reflect | Qwen3.7-Flash | |
|---|---|---|---|
| -Bench | 23.80 | 26.92 | +3.12 |
| VitaBench | 12.50 | 15.00 | +2.50 |
| BFCL-v4 | 72.45 | 74.70 | +2.25 |
RQ6: Does CHIME’s Memory Transfer Across Backbones?
We copy the memory accumulated by a source backbone on a benchmark’s train split into a target backbone’s frozen eval run. As shown in Table 6, w/o Mem. denotes the target running without memory, and Self denotes the target with its self-accumulated memory. CHIME’s transferred memory outperforms A-MapReduce’s in all four settings by between 2.25% and 4.68%, and closely approaches the upper bound: matching or exceeding it on the Qwen3.5-Flash target model, and staying within 1.72% on BFCL-v4 for the DeepSeek-V4-Flash target model. These experimental results reveal that attributed memory thus encodes reusable guidance rather than backbone-specific patterns.
| Benchmark | w/o Mem. | A-Map. | CHIME | Self |
|---|---|---|---|---|
| DeepSeek-V4-Flash Qwen3.5-Flash | ||||
| -Bench | 24.45 | 21.72 | 26.40 | 25.92 |
| BFCL-v4 | 70.86 | 71.66 | 73.91 | 73.91 |
| Qwen3.5-Flash DeepSeek-V4-Flash | ||||
| -Bench | 30.04 | 28.22 | 30.82 | 36.28 |
| BFCL-v4 | 60.13 | 61.85 | 65.43 | 67.15 |
7 Conclusion
In this work, we presented CHIME, an self-evolving agentic planning framework based on our proposed attribute-before-memorize principle, which separates and updates planning and execution memory bank through stage-level credit attribution. Experiments on four long-horizon agent benchmarks show that CHIME consistently outperforms strong training-based and self-evolving memory baselines. Further analyses validate CHIME along six dimensions: memory quality, evolution efficiency, memory value utility, credit attribution gate reliability, credit-model scaling, and cross-backbone transfer. These results suggest that reliable memory evolution depends not only on what agents remember, but also on whether each experience is attributed to the decision stage where it can provide effective guidance.
References
- Barres et al. (2025) Barres, V.; Dong, H.; Ray, S.; Si, X.; and Narasimhan, K. 2025. -Bench: Evaluating Conversational Agents in a Dual-Control Environment. arXiv:2506.07982.
- Chen et al. (2026) Chen, M.; Zhang, G.; Chang, H.; Guo, Y.; and Zhou, S. 2026. A-MapReduce: Executing Wide Search via Agentic MapReduce. arXiv:2602.01331.
- DeepSeek-AI et al. (2026) DeepSeek-AI; et al. 2026. DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence. arXiv:2606.19348.
- Erdogan et al. (2025) Erdogan, L. E.; Lee, N.; Kim, S.; Moon, S.; Furuta, H.; Anumanchipalli, G.; Keutzer, K.; and Gholami, A. 2025. Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks. arXiv:2503.09572.
- Guo, Kingston, and Kavraki (2024) Guo, W.; Kingston, Z.; and Kavraki, L. E. 2024. CaStL: Constraints as Specifications through LLM Translation for Long-Horizon Task and Motion Planning. arXiv:2410.22225.
- Hao et al. (2023) Hao, S.; Gu, Y.; Ma, H.; Hong, J. J.; Wang, Z.; Wang, D. Z.; and Hu, Z. 2023. Reasoning with Language Model is Planning with World Model. arXiv:2305.14992.
- He et al. (2025) He, W.; Sun, Y.; Hao, H.; Hao, X.; Xia, Z.; Gu, Q.; Han, C.; Zhao, D.; Su, H.; Zhang, K.; Gao, M.; Su, X.; Cai, X.; Cai, X.; Yang, Y.; and Zhao, Y. 2025. VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications. arXiv:2509.26490.
- Kagaya et al. (2024) Kagaya, T.; Yuan, T. J.; Lou, Y.; Karlekar, J.; Pranata, S.; Kinose, A.; Oguri, K.; Wick, F.; and You, Y. 2024. RAP: Retrieval-Augmented Planning with Contextual Memory for Multimodal LLM Agents. arXiv:2402.03610.
- Kim et al. (2023) Kim, S.; Moon, S.; Tabrizi, R.; Lee, N.; Mahoney, M. W.; Keutzer, K.; and Gholami, A. 2023. An LLM Compiler for Parallel Function Calling. arXiv:2312.04511.
- Lan et al. (2026) Lan, T.; Henry, F.; Zhu, B.; Jia, Q.; Ren, J.; PU, Q.; Li, H.; Wang, L.; Xu, Z.; and Luo, W. 2026. Table-as-Search: Agentic Information Seeking is Table Completion. In Liakata, M.; Moreira, V. P.; Zhang, J.; and Jurgens, D., eds., Findings of the Association for Computational Linguistics: ACL 2026, 15088–15107. San Diego, California, United States: Association for Computational Linguistics. ISBN 979-8-89176-395-1.
- Lan et al. (2025) Lan, T.; Zhu, B.; Jia, Q.; Ren, J.; Li, H.; Wang, L.; Xu, Z.; Luo, W.; and Zhang, K. 2025. DeepWideSearch: Benchmarking Depth and Width in Agentic Information Seeking. arXiv:2510.20168.
- Li et al. (2025) Li, X.; Jiao, W.; Jin, J.; Dong, G.; Jin, J.; Wang, Y.; Wang, H.; Zhu, Y.; Wen, J.-R.; Lu, Y.; and Dou, Z. 2025. DeepAgent: A General Reasoning Agent with Scalable Toolsets. arXiv:2510.21618.
- Liu et al. (2026) Liu, J.; Jiang, Y.; Zhang, G.; Zhang, Z.; Chang, H.; Yin, Z.; Ren, Q.; and Yan, J. 2026. TodoEvolve: Learning to Architect Agent Planning Systems. arXiv:2602.07839.
- Ouyang et al. (2026) Ouyang, S.; Yan, J.; Hsu, I.-H.; Chen, Y.; Jiang, K.; Wang, Z.; Han, R.; Le, L. T.; Daruki, S.; Tang, X.; Tirumalashetty, V.; Lee, G.; Rofouei, M.; Lin, H.; Han, J.; Lee, C.-Y.; and Pfister, T. 2026. ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory. arXiv:2509.25140.
- Patil et al. (2025) Patil, S. G.; Mao, H.; Cheng-Jie Ji, C.; Yan, F.; Suresh, V.; Stoica, I.; and E. Gonzalez, J. 2025. The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models. In Forty-second International Conference on Machine Learning.
- Qwen Team (2026a) Qwen Team. 2026a. Qwen3.5: Towards Native Multimodal Agents.
- Qwen Team (2026b) Qwen Team. 2026b. Qwen3.6-35B-A3B: Agentic Coding Power, Now Open to All.
- Shen et al. (2023) Shen, Y.; Song, K.; Tan, X.; Li, D.; Lu, W.; and Zhuang, Y. 2023. HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in HuggingFace. In Advances in Neural Information Processing Systems.
- Si et al. (2026) Si, S.; Zhao, H.; Luo, K.; Chen, G.; Qi, F.; Zhang, M.; Chang, B.; and Sun, M. 2026. A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Tasks. arXiv:2510.05608.
- Team et al. (2026a) Team, C.; et al. 2026a. MiMo-V2-Flash Technical Report. arXiv:2601.02780.
- Team et al. (2026b) Team, M.; Bai, S.; Bing, L.; Lei, L.; Li, R.; Li, X.; Lin, X.; Min, E.; Su, L.; Wang, B.; Wang, L.; Wang, L.; Wang, S.; Wang, X.; Zhang, Y.; Zhang, Z.; Chen, G.; Chen, L.; Cheng, Z.; Deng, Y.; Huang, Z.; Ng, D.; Ni, J.; Ren, Q.; Tang, X.; Wang, B. L.; Wang, H.; Wang, N.; Wei, C.; Wu, Q.; Xia, J.; Xiao, Y.; Xu, H.; Xu, X.; Xue, C.; Yang, Z.; Yang, Z.; Ye, F.; Ye, H.; Yu, J.; Zhang, C.; Zhang, W.; Zhao, H.; and Zhu, P. 2026b. MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification. arXiv:2603.15726.
- Wang et al. (2025) Wang, Y.; Takanobu, R.; Liang, Z.; Mao, Y.; Hu, Y.; McAuley, J.; and Wu, X. 2025. Mem-: Learning Memory Construction via Reinforcement Learning. arXiv:2509.25911.
- Wang et al. (2024) Wang, Z.; Cai, S.; Chen, G.; Liu, A.; Ma, X.; and Liang, Y. 2024. Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents. arXiv:2302.01560.
- Wong et al. (2025) Wong, R.; Wang, J.; Zhao, J.; Chen, L.; Gao, Y.; Zhang, L.; Zhou, X.; Wang, Z.; Xiang, K.; Zhang, G.; Huang, W.; Wang, Y.; and Wang, K. 2025. WideSearch: Benchmarking Agentic Broad Info-Seeking. arXiv:2508.07999.
- Wu et al. (2025) Wu, R.; Wang, X.; Mei, J.; Cai, P.; Fu, D.; Yang, C.; Wen, L.; Yang, X.; Shen, Y.; Wang, Y.; and Shi, B. 2025. EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle. arXiv:2510.16079.
- Xie et al. (2024) Xie, J.; Zhang, K.; Chen, J.; Zhu, T.; Lou, R.; Tian, Y.; Xiao, Y.; and Su, Y. 2024. TravelPlanner: A Benchmark for Real-World Planning with Language Agents. arXiv:2402.01622.
- Yang et al. (2025) Yang, Y.; Lan, T.; Jia, Q.; Zhu, L.; Jiang, H.; Zhu, H.; Wang, L.; Luo, W.; and Zhang, K. 2025. HSCodeComp: A Realistic and Expert-level Benchmark for Deep Search Agents in Hierarchical Rule Application. arXiv:2510.19631.
- Yao et al. (2023a) Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T. L.; Cao, Y.; and Narasimhan, K. 2023a. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv:2305.10601.
- Yao et al. (2023b) Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; and Cao, Y. 2023b. ReAct: Synergizing Reasoning and Acting in Language Models. arXiv:2210.03629.
- Ye et al. (2026) Ye, Y.; Jiang, H.; Jiang, F.; Lan, T.; Du, Y.; Fu, B.; Shi, X.; Jia, Q.; Wang, L.; and Luo, W. 2026. UMEM: Unified Memory Extraction and Management Framework for Generalizable Memory. arXiv:2602.10652.
- Yu et al. (2025) Yu, H.; Chen, T.; Feng, J.; Chen, J.; Dai, W.; Yu, Q.; Zhang, Y.-Q.; Ma, W.-Y.; Liu, J.; Wang, M.; and Zhou, H. 2025. MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent. arXiv:2507.02259.
- Yu et al. (2026) Yu, X.; Zhang, L.; Feng, X.; Jiang, Y.; Qin, B.; Xie, P.; and Zhou, J. 2026. WebAnchor: Anchoring Agent Planning to Stabilize Long-Horizon Web Reasoning. arXiv:2601.03164.
- Zhang et al. (2026a) Zhang, S.; Wang, J.; Zhou, R.; Liao, J.; Feng, Y.; Li, Z.; Zheng, Y.; Zhang, W.; Wen, Y.; Li, Z.; Xiong, F.; Qi, Y.; Tang, B.; and Wen, M. 2026a. MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory. arXiv:2601.03192.
- Zhang et al. (2026b) Zhang, Y.; Jiang, S.; Li, R.; Tu, J.; Su, Y.; Deng, L.; Guo, X.; Lv, C.; and Lin, J. 2026b. DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints. arXiv:2601.18137.
- Zhang et al. (2025) Zhang, Y.; Li, M.; Long, D.; Zhang, X.; Lin, H.; Yang, B.; Xie, P.; Yang, A.; Liu, D.; Lin, J.; Huang, F.; and Zhou, J. 2025. Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models. arXiv:2506.05176.
- Zhou et al. (2024) Zhou, A.; Yan, K.; Shlapentokh-Rothman, M.; Wang, H.; and Wang, Y.-X. 2024. Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models. arXiv:2310.04406.
- Zhou et al. (2025) Zhou, P.; Leon, B.; Ying, X.; Zhang, C.; Shao, Y.; Ye, Q.; Chong, D.; Jin, Z.; Xie, C.; Cao, M.; Gu, Y.; Hong, S.; Ren, J.; Chen, J.; Liu, C.; and Hua, Y. 2025. BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese. arXiv:2504.19314.
Appendix A Implementation Details
| Component | Setting | Value |
| Models and decoding | Agent backbone | qwen3.5-flash or deepseek-v4-flash |
| Evaluation-side LLM | qwen3.6-flash | |
| Default decoding | Temperature ; thinking disabled | |
| WebAnchor plan generation | Four candidates; temperature | |
| Maximum input/output length | / tokens | |
| Request policy | -s timeout; at most 15 retries; 50-s base retry interval | |
| Retrieval and reranking | Embedding configuration | Qwen3-Embedding-0.6B; batch size 16; maximum length |
| Candidate/retrieved set sizes | per memory layer | |
| Candidate similarity threshold | ||
| Reranking parameters | , , and | |
| Near-tie relevance margin | ||
| Memory evolution | Update confidence threshold | |
| Experience-addition threshold | ||
| Initial memory state | ||
| Merge threshold | ||
| Pruning thresholds | and | |
| Memory capacity | Unlimited | |
| Evaluation protocol | Train/eval split | within each benchmark domain |
| Selected memory checkpoint | Epoch 2 for -Bench, VitaBench, and BFCL-v4; epoch 5 for BrowseComp-ZH | |
| -Bench | User-simulator LLM | qwen3.6-flash |
| Interaction controls | 100 steps; at most 10 errors; seed 42 | |
| Subprocess timeout | s | |
| Knowledge retrieval | Local embedding retrieval; Top- | |
| VitaBench | User-simulator LLM | qwen3.6-flash |
| Agent/user configuration | llm_agent/user_simulator | |
| Benchmark language | English | |
| Interaction controls | 100 steps; at most 10 errors; seed 42 | |
| BrowseComp-ZH | Search framework | smolagents |
| Search-agent backbone | qwen3.5-flash or deepseek-v4-flash, matching the evaluated backbone | |
| Search configuration | Top-; at most 10 search steps | |
| Search/page-fetch timeout | 120 s each | |
| Page-length limit | characters | |
| Answer judge | Deterministic qwen3.6-flash | |
| BFCL-v4 | Evaluator | qwen3.6-flash checker |
| Search configuration | Top-; at most 10 tool/search steps | |
| Shared execution | Episode concurrency | 16 for -Bench, BrowseComp-ZH, and BFCL-v4; 8 for VitaBench |
CHIME keeps the agent backbone frozen throughout, updating only the external planning and execution memory banks during accumulation and freezing them during transfer. These updates use only interaction trajectories and environment feedback, not reference answers. For consistent comparison, all methods share the same benchmark setup and evaluation-side models, and each baseline uses the best-performing configuration reported in its paper. All results are averaged over three random instance orderings (Avg@3).
We evaluate CHIME with Qwen3.5-Flash (Qwen Team, 2026a) and Deepseek-V4-Flash (DeepSeek-AI et al., 2026); in each setting, the same backbone handles planning, execution, and self-reflection in the Credit Attribution Gate. All evaluation-side LLM roles use Qwen3.6-Flash (Qwen Team, 2026b), while Qwen3-Embedding-0.6B (Zhang et al., 2025) serves as the retriever. Table 7 lists the complete implementation settings.
Appendix B Failure Analysis
Our analysis reveals two practical boundaries of CHIME.
Agent Capability.
CHIME augments the agent with memory guidance but leaves the backbone unchanged. As a result, even useful memory may not improve the outcome when the agent cannot follow the guidance or execute the required actions correctly. In the gate reliability analysis, rerunning failed episodes with the generated memory resolves 56.3–67.6% of attributed failures, but not all of them. The benefit of CHIME therefore remains bounded by the underlying capabilities of the agent.
Memory Transferability.
Our cross-backbone transfer results show that memory can transfer across backbones when the benchmark remains unchanged. Tasks within the same benchmark generally share the conditions under which a memory applies, including the available tools, valid actions, and success criteria. In contrast, we observe weaker transfer across benchmarks, where semantic similarity does not ensure that these conditions remain unchanged. CHIME’s memory is therefore most transferable across tasks with similar scenarios and operating conditions, while transfer across different benchmarks remains limited.
Appendix C Case Studies
We include three representative cases to illustrate how the memory bank is used and updated. The first case shows a planner-side benefit, where retrieved memories help decompose a complex request before execution. The second case shows an executor-side benefit, where retrieved memories help bind exact tool arguments. The third case shows the attribution gate itself: a failed trajectory is routed to the planner layer because the executor followed the plan and the reusable error lies in the planning abstraction.
Planner-side decomposition.
Figure 5 shows why planner memory cannot be replaced by a generic summary of past successes. The task contains several retail operations that look similar in natural language but require different protocols: returned items belong to delivered orders, color and address edits belong to a pending order, and the requested tracking number belongs to a cancelled order. The retrieved planner memories directly encode this object-status decomposition. As a result, the agent plans the correct action family for each object before making tool calls, while the baselines produce plausible but verifier-failing summaries.
Executor-side precision.
Figure 6 highlights a different failure mode. The high-level plan is not the main challenge: all methods can state that the agent should review accounts and correct ATM-fee errors. The reward depends on execution details: retaining three account identities, computing net credits after subtracting missing fees, and invoking the credit tool with the exact amount for each account. The executor memory is useful because it is written at the same granularity as the eventual action arguments.
Layer-aware attribution.
Figure 7 illustrates why the gate uses trajectory-level diagnosis rather than only the binary verifier outcome. The delivery episode fails, but it should not penalize executor memory: the selected restaurant, address, and item choices are mostly correct, and the executor follows the planned schedule. The reusable error is the planner’s treatment of a strict “before” constraint as if equality were acceptable. The gate therefore writes a planner memory about strict deadline buffers and leaves the executor bank untouched. This keeps the memory bank more specific and reduces cross-layer contamination.