Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems
Abstract
Large language model (LLM)–based multi-agent systems tackle complex reasoning by orchestrating how multiple agents are configured and how they collaborate. A central challenge is to adapt orchestration to the evolving collaboration state. Routing from the query alone cannot adapt to intermediate progress or errors, which hurts accuracy. Routing from the complete execution history supplies this missing context, but forces later decisions to process every prior step, including redundant or low-utility ones. This creates an execution-history overload that inflates cost. Effective orchestration instead requires a compact state that captures useful progress without accumulating redundant context. We propose Gated-Memory Routing, which conditions each decision on the query and a learned execution memory. A learned Memory Write Gate commits only non-redundant reasoning steps, and a learned Retrieval Gate supplies each agent a compact, relevant subset, so every decision conditions on a clean, informative state. At each step, the system selects the next role and backbone from this memory, while an Adaptive Halting Controller stops execution once the memory contains sufficient evidence for answering. Across five reasoning and code-generation benchmarks, our framework is both effective and efficient: it attains the best average accuracy, exceeding the strongest baseline by points, while reducing HumanEval inference cost by relative to that baseline. Code is available at https://github.com/rajibrhasan/gated-memory-routing.
1 Introduction
Large language model (LLM)–based multi-agent systems have become a powerful paradigm for complex reasoning: by decomposing a problem across agents that specialize, critique, or extend one another’s reasoning, they achieve strong performance on reasoning and code-generation benchmarks Chen et al. (2024c); Liang et al. (2024); Zhang et al. (2024); Qian et al. (2025). Their effectiveness depends not only on the underlying LLMs but on how execution is orchestrated: how many agents to invoke, which roles to assign, which backbone powers each agent, and what each agent may see. These decisions constitute the routing problem, and how well they are made determines whether collaboration yields genuine reasoning gains or redundant computation.
Early multi-agent systems address routing through fixed, hand-engineered designs, instantiating predetermined agents with manually specified roles and static collaboration patterns that remain constant across queries Hong et al. (2024); Qian et al. (2024); Li et al. (2023); Du et al. (2024); Wu et al. (2024). To relax this rigidity, later work makes the interaction structure adaptive by optimizing or pruning the communication graph Zhuge et al. (2024); Zhang et al. (2025a); Wang et al. (2025); Liu et al. (2024), and end-to-end routers configure agent count, roles, and backbones directly from the query Yue et al. (2025). These query-only routers commit routing decisions before any agent produces reasoning, so they cannot respond to partial progress, errors, or emerging gaps in the solution. A recent alternative conditions routing on the unfolding execution, sequencing agents over the evolving task state Dang et al. (2025). When that evolving state is carried forward with little filtering, however, each decision must reason over an undifferentiated, ever-growing context that mixes useful evidence with redundant or erroneous reasoning; this full-history routing exposes downstream agents to low-value context and inflates inference cost.
These two approaches have opposite limitations: conditioning on the query alone observes too little, committing before execution begins, which hurts accuracy, while conditioning on the full execution history (Figure 11(a)) observes too much, forcing each decision to process an unfiltered and growing trajectory, which inflates cost. A router should instead learn what to retain, what to surface, and how to act on the resulting state. This motivates conditioning routing on a structured, selective memory (Figure 11(b)): the system dynamically routes based on a gated memory, exposes only step-relevant context to downstream agents, and commits only the reasoning steps worth preserving. These decisions are tightly coupled: low-quality writes pollute later retrieval, overly broad retrieval distracts downstream agents, and fixed-depth execution either stops prematurely or continues spending computation after additional reasoning is unlikely to help.
Based on this view, we introduce Gated-Memory Routing, a framework in which a learned, gated execution memory serves as the shared state that every routing decision conditions on. A single jointly trained router acts on it through complementary components: a History-Aware Role Allocator and an LLM Router choose the agent and its backbone, a Retrieval Gate surfaces a compact, step-relevant subset of memory, a Memory Write Gate commits only high-utility, non-redundant reasoning steps, and an Adaptive Halting Controller decides when the gated state is sufficient to stop. Because routing is driven by this gated memory rather than the raw history, the interaction structure emerges dynamically during execution.
Contributions.
- •
Routing over a learned gated memory. We recast multi-agent orchestration as routing over a learned, selective execution memory: each step conditions on a filtered task-relevant state rather than the query alone (query-only) or the raw execution history (full-history).
- •
Memory curation and budget-aware halting. Learned write and retrieval gates control what memory stores and surfaces, keeping the state compact and high-signal, while a budget-aware halting policy consumes that clean state to control reasoning depth and thus cost, all trained end-to-end with role and backbone routing under a group-relative, cost-aware objective.
- •
Empirical effectiveness and efficiency. Across MATH, GSM-Hard, MBPP, HumanEval, and MMLU-Pro, Gated-Memory Routing attains the best average accuracy, exceeding the strongest baseline by points while cutting HumanEval inference cost by relative to that baseline.
2 Related Work
Fixed and Role-Based Multi-Agent Systems.
Representative LLM-based multi-agent systems design collaboration through predefined roles and static interaction graphs. MetaGPT Hong et al. (2024) enforces role-based workflows through Standard Operating Procedures, while ChatDev Qian et al. (2024) models software development as a fixed linear chain. These systems demonstrate the value of specialization, but their communication patterns are fixed before execution begins and cannot adapt the participating agents, context, or depth as execution unfolds.
Adaptive Topology and Graph Optimization.
Several methods make workflows more flexible by adapting interaction structures or optimizing communication graphs. GPTSwarm Zhuge et al. (2024) optimizes orchestration through learnable prompting, and AFlow Zhang et al. (2025b) searches over executable workflows. AgentPrune Zhang et al. (2025a) and AgentDropout Wang et al. (2025) induce sparsity by pruning connections or agents from predefined templates. DyLAN Liu et al. (2024) goes further by pruning low-contribution agents during inference using peer evaluation and consensus, and OMAC Li et al. (2026) jointly optimizes agent functionality and collaboration structure. These methods improve flexibility by optimizing which agents participate and how they are connected, but they do not learn which intermediate information subsequent orchestration decisions should condition on; DyLAN, for example, adapts by removing agents while still forwarding the unfiltered outputs of the remaining ones.
Query-Level LLM and MAS Routing.
LLM routing optimizes the trade-off between model cost and capability. RouteLLM Ong et al. (2025), RouterDC Chen et al. (2024b), R2-Router Xue et al. (2026), HW-Router Kabir et al. (2026), Securerouter Zhang et al. (2026) and FrugalGPT Chen et al. (2024a) route each query to a cost-effective model. In multi-agent systems, MASRouter Yue et al. (2025) trains an RL policy to configure execution components, including agent count, role assignment, and backbone selection, from the input query. These methods make routing more cost-aware and task-adaptive, but their decisions are primarily conditioned on the query rather than on the unfolding execution trajectory. They therefore cannot revise orchestration based on the quality, novelty, or redundancy of intermediate trajectory elements.
Routing over the Execution History.
Recent work addresses the query-only routing limitation by conditioning orchestration on the unfolding execution state. Evolving Orchestration Dang et al. (2025) uses RL to train a centralized orchestrator that selects agents from the evolving execution state and terminates through a designated stopping action. This makes routing responsive to execution, but the state retains the accumulated trajectory with limited filtering, so intermediate elements are carried forward to later decisions. Such full-history routing can lead to execution-history overload as execution deepens, forcing later decisions to process redundant, noisy, or low-utility intermediate content. Our framework instead routes through a gated execution memory, conditioning each decision on a filtered state rather than the unfiltered history, and outperforms Evolving Orchestration by average points across the benchmarks in Table 1. Our adaptive halting also relates to adaptive computation Graves (2016) and early-exit inference, which learn when to stop computing within a single model; in contrast, it halts multi-agent collaboration based on the gated memory state.
Memory-Optimized Agent Architectures.
Memory management is a core component of agent systems. MemGPT Packer et al. (2023) pages information in and out of context through a virtual memory hierarchy, AIOS Mei et al. (2025) manages memory as a shared OS-level resource for agents, and systems such as A-mem Xu et al. (2025) and MM Hatalis et al. (2023) organize memory through categorization and retrieval. Memory-R1 Yan et al. (2026) uses reinforcement learning to train a single agent’s policies for writing to and retrieving from a long-term memory over an external store. These systems primarily treat memory as a storage and retrieval substrate, optimizing what information should be preserved or recalled under a context budget. In contrast, Gated-Memory Routing treats the within-query execution memory of a multi-agent trajectory as an active control signal: write and retrieval decisions are trained jointly with role allocation, backbone routing, and halting under the task-level reward, so memory not only supplies context but also determines who acts next, what they read, and when collaboration stops.
3 Methodology
3.1 Problem Definition
We formalize gated-memory routing as a sequential decision process over a set of agent specializations , a pool of heterogeneous LLM backbones , and an execution memory maintained within a single query execution. At step , the observable state is , where holds the retained execution records from prior steps. The system selects a specialization and a backbone , retrieves a context subset , generates a reasoning step , and decides whether to commit the record , so memory evolves as a gated subset of the trajectory, . The process terminates at a realized depth when the halting policy fires or the maximum depth is reached; because halting is evaluated after each step, at least one agent always acts. An aggregator then produces the final answer from the terminal state . Throughout, boldface denotes the frozen sentence embedding of a quantity (e.g., ).
Writing for the full raw trajectory before step , routing strategies differ in the state each decision sees: query-only routing conditions on alone, full-history routing on and the entire , and gated-memory routing on , pairing the query with a learned selective memory derived from that keeps only selected records.
3.2 Framework Overview
Figure 2 illustrates Gated-Memory Routing. Its central object is the gated execution memory that every decision conditions on. A single jointly trained router acts over it through complementary components: a Memory Write Gate (Section 3.6) and Retrieval Gate (Section 3.5) keep the memory high-signal, so the History-Aware Role Allocator (Section 3.3) and LLM Router (Section 3.4) choose from a gated execution state, while the Adaptive Halting Controller (Section 3.7) reads that state to decide when to stop, making depth, and thus cost, adaptive. Because the memory stays compact and trustworthy, it can halt early without the overload full-history routing incurs. The trajectory is trained end-to-end with a single reward trading answer quality against cost (Section 3.8); encoders and representations appear in Appendix 8.
3.3 History-Aware Role Allocator
Role allocation chooses which agent specialization acts at each step. Rather than committing from the query alone or reacting to the full raw trajectory, our allocator conditions on the current state , so the next specialization complements the retained solution state rather than every intermediate step produced.
Because specializations such as a critic, verifier, or debugger carry semantic structure rather than being interchangeable IDs, we represent each role through its textual description rather than a categorical label. The description is encoded and projected through a Variational Autoencoder (VAE) Kingma and Welling (2014) into a continuous latent role embedding . A state encoder maps to a context vector , and the allocator defines a stochastic policy over the available roles :
This policy selects the specialization whose latent description best matches the query and current memory state, supporting transfer across related roles and extension to newly described specializations without changing the routing interface.
3.4 LLM Router
Backbone selection matches the current reasoning need to a model’s capability and cost. Because this need depends on the selected role and the progress already in memory (a difficult step may warrant a stronger model, while memory may already supply enough guidance for a cheaper one), we model LLM routing as a role-conditioned policy over the current state rather than a query-level choice.
The router combines the state context with the selected role representation via a learned projection to form a role-conditioned context . Each candidate LLM is represented by a continuous latent capability vector , derived from its natural-language description using the same description-encoding scheme as the role allocator, rather than by a fixed index. The router samples the backbone according to
This lets the system route each step to a model whose described capabilities match the current role and memory state.
3.5 Retrieval Gate
Even a selective memory should not be exposed wholesale to every agent: many retained records belong to other subproblems or roles irrelevant to the agent about to act, so passing all of them inflates token cost and buries the entries that matter. A fixed top- retriever is equally rigid, since some steps need broad context while others need one record or none. The Retrieval Gate therefore treats context construction as a step-adaptive selection, deciding independently for each record whether to surface it.
Each stored record is summarized by an embedding that fuses its role latent, backbone latent, and a projection of its response, and the current step forms a retrieval query by the analogous fusion of the query with the role and backbone just selected. The gate therefore asks which prior records a step of this configuration should read: for record , the retrieval logit and binary decision are
and the retrieved context is
Because the draws are independent, varies in size, contracting toward empty when memory is largely irrelevant and expanding when several records bear on the step. Independence keeps each decision local; cross-record redundancy is instead suppressed at write time (Section 3.6). The gate has no cost term of its own: retrieving more records lengthens the agent prompt and is paid for only through the trajectory reward (Section 3.8).
3.6 Memory Write Gate
What the system writes to memory determines the state on which all later decisions are based. If every intermediate element is appended, gated memory collapses back into full-history routing, and redundant records crowd memory and dilute retrieval. The Memory Write Gate therefore treats memory construction as a selective decision rather than a default append: a new record should enter memory only when it is both relevant to the task and novel with respect to the records already stored.
We score this relevance–novelty trade-off with a Maximal Marginal Relevance criterion Carbonell and Goldstein (1998), but replace its fixed components with learned, stochastic ones. Let denote the current context state (Section 3.3), the embedding of the new reasoning step, and the embeddings of stored steps, and let be a cosine similarity taken in a learned projection of these frozen embeddings, so the relevance and novelty geometry is trained rather than prescribed. The write score is
with a learned coefficient balancing relevance against redundancy with stored records; the novelty term is dropped when memory is empty, so the first record is admitted on relevance alone. Rather than greedily keeping the top-scoring record as in deterministic MMR, we sample the write decision , keeping memory construction trainable under the same policy objective as the routing decisions. The novelty criterion controls cross-record redundancy, allowing the retrieval gate (Section 3.5) to score records independently. As a result, memory remains a compact representation of execution progress rather than accumulating into a transcript.
3.7 Adaptive Halting Controller
Reasoning depth is a major driver of cost, yet a depth fixed from the query commits this budget before the system observes how execution unfolds. We therefore make depth a consequence of execution: after each step the Adaptive Halting Controller decides whether the memory holds sufficient evidence for aggregation or another agent should act.
Whether recent steps still add evidence is a trend no single-step snapshot captures, so the controller carries a recurrent state that integrates a pooled summary of the post-write memory across steps, kept separate from the query-attended routing context . It then samples a halt action
Because the decision is evaluated only after the step’s agent has executed, at least one agent always runs. If the stored records pass to the Aggregator LLM; otherwise execution continues to the maximum depth . The halt action is optimized under the same trajectory-level reward as the other decisions (Section 3.8), so reasoning depth co-adapts with routing.
3.8 Optimization
The routing stack couples several discrete stochastic decisions (role, backbone, retrieval, write, and halt), and their number varies with the halt-determined depth, so we optimize the trajectory with policy gradients rather than backpropagating through sampled actions. Because queries differ widely in difficulty, a single trajectory baseline is noisy; we instead train all policies jointly with a group-relative advantage in the spirit of GRPO Shao et al. (2024), using rollouts of each query as a per-query baseline without a learned critic.
For each query we sample trajectories that differ only in their sampled decisions. Trajectory has utility , where is task success, sums the per-step backbone costs, and sets the accuracy–cost trade-off (Appendix 9). Centering each utility within its group yields the advantage
a per-query baseline that cancels difficulty without a critic. We deliberately drop the usual within-group standard-deviation normalization: when rollouts for a query share near-identical rewards, the standard deviation approaches zero and inflates the advantage from negligible cost differences, whereas the centered advantage remains well-scaled. The same advantage weights every sampled action. With the summed log-probability of all role, backbone, retrieval, write, and halt actions in , the objective is
where is stop-gradient, regularizes the role and backbone encoders (Appendix 8), and is the mean per-step entropy of the routing and halt policies. The entropy term discourages premature deterministic behavior; Appendix 12 verifies empirically that the trained gates do not collapse to trivial always-write, never-write, retrieve-all, or retrieve-none solutions. The LLM backbones and sentence encoder are frozen, so the router is the only trained component, updated jointly by a single optimizer.
4 Experimental Setup
Datasets and Benchmarks.
We evaluate on five benchmarks spanning mathematical reasoning, program synthesis, and knowledge-intensive question answering: GSM-Hard Gao et al. (2023), MATH Hendrycks et al. (2021), HumanEval Chen et al. (2021), MBPP Austin et al. (2021), and MMLU-Pro Wang et al. (2024).
Baselines.
Our baselines span single-model inference; single-agent reasoning (CoT Wei et al. (2022), Complex-CoT Fu et al. (2023)); fixed-topology multi-agent collaboration (MacNet-Chain, -Tree, and -Complete Qian et al. (2025)); adaptive workflow optimization (AFlow Zhang et al. (2025b)); evolving orchestration (Puppeteer Dang et al. (2025)); and query-only routing (MASRouter Yue et al. (2025)).
LLM Backbones.
We route over five open-weight LLMs: llama-3.2-3B, llama-3.1-8B Grattafiori et al. (2024), mistral-nemo-12B Mistral AI Team (2024), qwen-2.5-14B, and qwen-2.5-32B Yang et al. (2024). This 3B–32B heterogeneous pool allows the router to trade off capability and inference cost. To quantify the accuracy–cost trade-off, we assign each backbone a price proportional to its parameter count, a hardware-independent proxy for per-token inference compute (Appendix 9). Backbone capability descriptions are listed in Appendix 16.
4.1 Implementation Details
We train with Adam optimizer using learning rate , GRPO group size , and queries per batch. We set maximum depth , cost penalty , entropy coefficient , and VAE regularizer weight . For the fixed-topology multi-agent baselines, we set the number of agents equal to our maximum depth , so every multi-agent method operates under the same agent budget. We use the same heterogeneous role profiles as MASRouter Yue et al. (2025), spanning programming agents with compiler access to research-oriented agents with external knowledge sources; full descriptions are in Appendix 14. For final aggregation, we use the LLM backbone selected most frequently across the query’s reasoning steps. The aggregation prompt is provided in Appendix 15.
| Method | LLM | Mul. | Rout. | MATH | GSM-Hard | MBPP | HumanEval | MMLU-Pro | Avg. |
| Single LLM | llama-3.2-3B | ✗ | ✗ | 43.27 | 25.85 | 52.20 | 62.79 | 34.00 | 43.62 |
| llama-3.1-8B | ✗ | ✗ | 43.75 | 39.87 | 57.60 | 69.78 | 48.25 | 51.85 | |
| mistral-nemo-12B | ✗ | ✗ | 42.07 | 30.11 | 56.00 | 68.22 | 27.75 | 44.83 | |
| qwen-2.5-14B | ✗ | ✗ | 75.72 | 64.58 | 71.80 | 82.95 | 65.25 | 72.06 | |
| qwen-2.5-32B | ✗ | ✗ | 78.37 | 61.52 | 79.00 | 84.37 | 66.02 | 73.86 | |
| CoT Wei et al. (2022) | qwen-2.5-14B | ✗ | ✗ | 77.88 | 67.52 | 70.80 | 82.17 | 62.16 | 72.11 |
| qwen-2.5-32B | ✗ | ✗ | 78.85 | 64.56 | 77.20 | 85.15 | 68.41 | 74.83 | |
| Complex-CoT Fu et al. (2023) | qwen-2.5-14B | ✗ | ✗ | 75.96 | 66.76 | 70.40 | 80.62 | 63.07 | 71.36 |
| qwen-2.5-32B | ✗ | ✗ | 77.16 | 65.32 | 78.40 | 85.47 | 68.41 | 74.95 | |
| MacNet-Chain Qian et al. (2025) | qwen-2.5-14B | ✓ | ✗ | 73.80 | 59.75 | 82.66 | 86.82 | 61.48 | 72.90 |
| qwen-2.5-32B | ✓ | ✗ | 77.64 | 61.93 | 80.73 | 83.80 | 67.95 | 74.41 | |
| MacNet-Tree Qian et al. (2025) | qwen-2.5-14B | ✓ | ✗ | 76.68 | 60.32 | 79.84 | 82.95 | 60.68 | 72.09 |
| qwen-2.5-32B | ✓ | ✗ | 77.88 | 62.31 | 80.71 | 86.02 | 67.73 | 74.93 | |
| MacNet-Complete Qian et al. (2025) | qwen-2.5-14B | ✓ | ✗ | 77.16 | 59.00 | 80.65 | 84.50 | 61.48 | 72.56 |
| qwen-2.5-32B | ✓ | ✗ | 79.09 | 62.22 | 81.20 | 86.05 | 67.16 | 75.14 | |
| AFlow Zhang et al. (2025b) | qwen-2.5-14B | ✓ | ✗ | 63.46 | 58.24 | 70.00 | 85.27 | 64.66 | 68.33 |
| qwen-2.5-32B | ✓ | ✗ | 77.64 | 67.05 | 76.40 | 84.50 | 68.41 | 74.80 | |
| Puppeteer Dang et al. (2025) | qwen-2.5-14B | ✓ | ✗ | 75.00 | 65.91 | 72.58 | 83.59 | 65.23 | 72.46 |
| qwen-2.5-32B | ✓ | ✗ | 79.09 | 68.75 | 74.80 | 85.16 | 68.64 | 75.29 | |
| MASRouter Yue et al. (2025) | LLM Pool | ✓ | ✓ | 74.31 | 66.00 | 79.20 | 85.16 | 66.70 | 74.27 |
| Ours | LLM Pool | ✓ | ✓ | 79.33 | 70.55 | 79.60 | 89.84 | 69.32 | 77.73 |
5 Results
5.1 Performance Analysis
Table 1 shows that Gated-Memory Routing achieves the strongest overall performance: it attains the best average accuracy and the top result on four of the five benchmarks, improving on the strongest baseline by points on average. The gains are most pronounced on GSM-Hard and HumanEval, where intermediate reasoning quality and selective context matter most; the one exception is MBPP, where a fixed multi-agent chain topology powered by qwen-2.5-14B edges ahead at pass@1 but trails on every other benchmark and on average. On MBPP our method is comparable to MASRouter ( vs. ) and ahead of Puppeteer-32B (), while several fixed MacNet configurations score highest. MBPP’s short, relatively uniform problems appear to benefit from a fixed pipeline built on a strong code model, leaving less room for adaptive orchestration to help; the benefit of gated memory is most visible where reasoning is longer and more heterogeneous.
The comparison is sharpest against the two paradigms our method is built to improve on. MASRouter routes from the query alone, committing every decision before any reasoning exists, and trails by points. Puppeteer orchestrates agents over the evolving execution state but does not route over the pool, running every step on the largest backbone in the pool (qwen-2.5-32B), and still trails by points. Gated-memory routing conditions each decision on a filtered state instead, and is the most accurate of the three on every benchmark, even though it routes over a heterogeneous pool rather than relying on the largest model alone. The gain is therefore not a matter of raw scale but of how well roles, backbones, and memory are coordinated as the trajectory unfolds. Appendix 11 reports the mean and standard deviation over three training seeds for our method and both routing baselines; the ranking of the three methods is identical under every seed.
5.2 Cost Analysis
Figure 3 places our method on the Pareto frontier of accuracy against inference cost on HumanEval, where it costs less than MASRouter and less than Puppeteer (qwen-2.5-32B) while remaining more accurate than both, gaining on cost and accuracy at once rather than trading one for the other. The GSM-Hard frontier in Appendix 10 shows the same pattern. Appendix 9 additionally reports approximate FLOPs and batch-amortized wall-clock time per query on HumanEval and MBPP, which give the same ordering.
These savings come from where computation is spent. Because each decision conditions on a compact, gated memory rather than the full execution history, the halting controller can stop once that memory carries enough evidence, so easy queries terminate early and only the hard ones run deep. The gates contribute indirectly: a clean memory makes halting trustworthy and spares downstream agents the redundant context that inflates prompts under full-history routing. Section 5.5 compares the three paradigms directly: gated-memory routing is at least as accurate as full-history routing while costing appreciably less, and is both more accurate and cheaper than query-only routing.
5.3 Ablation Study
| Configuration | GSM-Hard | HumanEval | ||
| Acc. | Cost | Acc. | Cost | |
| Full system | 70.55 | 0.587 | 89.84 | 0.032 |
| w/o role allocator | 68.37 | 0.400 | 85.94 | 0.043 |
| w/o LLM router | 56.53 | 0.555 | 71.88 | 0.026 |
| w/o retrieval gate | 69.22 | 0.488 | 85.94 | 0.034 |
| w/o write gate | 66.67 | 0.621 | 84.38 | 0.037 |
| w/o halting | 70.08 | 0.840 | 86.72 | 0.086 |
| w/o both memory gates | 67.99 | 0.447 | 83.59 | 0.033 |
Table 2 reports a leave-one-out ablation, each variant replacing one module with a default: w/o role allocator uses random role selection, w/o LLM router random backbone selection, w/o retrieval gate exposes the full memory to each agent, w/o write gate writes every response to memory, and w/o halting runs every query to the maximum depth . The effects split cleanly along the division of labor the method is built around. The two routing decisions carry accuracy: dropping the LLM router is by far the most damaging change (a -point drop on GSM-Hard), and dropping the role allocator costs a smaller but consistent amount, so matching model capacity and specialization to the current state is what drives answer quality.
The memory gates also serve accuracy through the quality of the memory state: removing either lowers accuracy, because unfiltered context dilutes what each agent and the router condition on. Disabling both gates at once (w/o both memory gates) lowers accuracy by points on GSM-Hard and on HumanEval, so selective memory matters even with the router and role allocat Halting is the lever that controls cost: without it, inference cost rises by over on GSM-Hard and more than doubles on HumanEval, while accuracy also slips, since most queries are solved before the depth limit and the extra steps add computation without improving answers. The w/o halting variant also provides an equal-depth comparison with the fixed-topology baselines: executing all six steps, it reaches on GSM-Hard, above every MacNet-32B variant (–, Table 1) at lower cost (Appendix 10), because each step is still routed to an appropriate backbone rather than the largest one. Halting also reduces training-time computation: disabling it raises the number of training rollout steps by about on both benchmarks (Appendix 13).
5.4 Effect of the Depth Budget
Figure 4 varies the maximum depth on MATH with adaptive halting left in place. Accuracy increases with the budget but saturates: most of the gain is realized by , with under a point more from to as cost keeps climbing. Small budgets truncate useful reasoning on the harder problems, while larger ones mostly add cost, since halting already stops easy queries early regardless of the ceiling. The default thus sits near the knee of the accuracy–cost curve, capturing nearly all the attainable accuracy before the budget begins buying cost more than accuracy.
5.5 Routing-Paradigm Comparison
We contrast Gated-Memory Routing with the two routing paradigms it is built to improve on. Query-only routing fixes every decision from the input query before execution begins; we instantiate it with MASRouter, the strongest query-only router among our baselines. Under full-history routing, the per-step role and backbone allocation is conditioned on the accumulated history, and all intermediate reasoning is forwarded as context to the next agent. Gated-memory routing is our full system. Table 3 compares the three.
| GSM-Hard | HumanEval | |||
| Routing paradigm | Acc. (%) | Cost | Acc. (%) | Cost |
| Query-only (MASRouter) | 66.00 | 1.012 | 85.16 | 0.057 |
| Full-history | 70.27 | 0.985 | 89.06 | 0.068 |
| Gated-memory (ours) | 70.55 | 0.587 | 89.84 | 0.032 |
Full-history routing conditions each decision on the entire accumulated trajectory, increasing context length and reasoning overhead. The additional context does not appear to improve accuracy here: full-history routing is roughly on par with the curated memory ( vs on GSM-Hard and vs on HumanEval), while the gated memory attains comparable accuracy at appreciably lower cost (about on GSM-Hard and over on HumanEval). Query-only routing tends to be the least accurate of the three. These results suggest that, in our setting, curating the memory rather than forwarding the full history largely preserves accuracy while reducing cost.
6 Conclusion
We presented Gated-Memory Routing, a framework for dynamic multi-agent coordination that conditions routing decisions on a learned, gated execution memory rather than on the query alone or the full raw history. A Memory Write Gate and a Retrieval Gate keep this memory compact and task-relevant, enabling role and backbone routing to operate over a filtered execution state. An Adaptive Halting Controller uses the same memory to decide when to stop, adapting reasoning depth and cost to each query. Across five benchmarks, our method attains the best average accuracy, exceeding the strongest baseline by points, while reducing HumanEval inference cost by relative to that baseline. Future work will extend gated-memory routing to open-ended generation, larger-scale agent ecosystems, and pools of more recent backbone models, and explore richer reward signals beyond binary task success.
Limitations
Our evaluation uses closed-domain tasks with verifiable answers, enabling automatic reward computation; open-ended generation would require reward models, human evaluation, or other supervision. Our efficiency analysis: parameter count provides a hardware-independent compute proxy but does not capture latency, memory pressure, or batching, while FLOPs and wall-clock measurements in Appendix 9 cover only two benchmarks. Finally, we evaluate open-weight Qwen2.5, Llama-3, and Mistral models; backbones can be substituted through their capability profiles and per-token costs without changing the framework.
References
- Program synthesis with large language models. arXiv preprint arXiv:2108.07732. External Links: Link Cited by: §4.
- The use of mmr, diversity-based reranking for reordering documents and producing summaries. In Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’98, New York, NY, USA, pp. 335–336. External Links: ISBN 1581130155, Link, Document Cited by: §3.6.
- FrugalGPT: how to use large language models while reducing cost and improving performance. Transactions on Machine Learning Research. Note: Featured Certification External Links: ISSN 2835-8856, Link Cited by: §2.
- Evaluating large language models trained on code. CoRR abs/2107.03374. External Links: Link, 2107.03374 Cited by: §4.
- RouterDC: query-based router by dual contrastive learning for assembling large language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §2.
- AgentVerse: facilitating multi-agent collaboration and exploring emergent behaviors. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024, External Links: Link Cited by: §1.
- Multi-agent collaboration via evolving orchestration. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §1, §2, §4, Table 1.
- Improving factuality and reasoning in language models through multiagent debate. In Forty-first International Conference on Machine Learning, External Links: Link Cited by: §1.
- Complexity-based prompting for multi-step reasoning. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023, External Links: Link Cited by: §4, Table 1.
- PAL: program-aided language models. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. External Links: Link Cited by: §4.
- The llama 3 herd of models. External Links: 2407.21783, Link Cited by: §4.
- Adaptive computation time for recurrent neural networks. ArXiv abs/1603.08983. External Links: Link Cited by: §2.
- Memory matters: the need to improve long-term memory in llm-agents. In Proceedings of the 2023 AAAI Fall Symposia, Arlington, Virginia, USA, October 25-27, 2023, C. W. Geib and R. P. A. Petrick (Eds.), pp. 277–280. External Links: Link, Document Cited by: §2.
- Measuring mathematical problem solving with the MATH dataset. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), External Links: Link Cited by: §4.
- MetaGPT: meta programming for a multi-agent collaborative framework. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §1, §2.
- HW-Router: hardware-aware routing for scalable multi-LLM serving. In Proceedings of the 63rd Design Automation Conference (DAC), Cited by: §2.
- Scaling laws for neural language models. CoRR abs/2001.08361. External Links: Link, 2001.08361 Cited by: §9.
- Auto-encoding variational bayes. In 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Conference Track Proceedings, Y. Bengio and Y. LeCun (Eds.), External Links: Link Cited by: §3.3.
- CAMEL: communicative agents for ”mind” exploration of large language model society. In Thirty-seventh Conference on Neural Information Processing Systems, External Links: Link Cited by: §1.
- OMAC: a holistic optimization framework for LLM-based multi-agent collaboration. In Forty-third International Conference on Machine Learning, External Links: Link Cited by: §2.
- Encouraging divergent thinking in large language models through multi-agent debate. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 17889–17904. External Links: Link, Document Cited by: §1.
- A dynamic LLM-powered agent network for task-oriented agent collaboration. In First Conference on Language Modeling, External Links: Link Cited by: §1, §2.
- AIOS: LLM agent operating system. In Second Conference on Language Modeling, External Links: Link Cited by: §2.
- Mistral nemo. Note: https://mistral.ai/news/mistral-nemo/Accessed: 2026-08-28 Cited by: §4.
- RouteLLM: learning to route LLMs from preference data. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §2.
- MemGPT: towards llms as operating systems. arXiv preprint arXiv:2310.08560. External Links: Link Cited by: §2.
- ChatDev: communicative agents for software development. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 15174–15186. External Links: Link, Document Cited by: §1, §2.
- Scaling large language model-based multi-agent collaboration. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, External Links: Link Cited by: §1, §4, Table 1, Table 1, Table 1.
- DeepSeekMath: pushing the limits of mathematical reasoning in open language models. CoRR abs/2402.03300. External Links: Link, Document, 2402.03300 Cited by: §3.8.
- MMLU-pro: a more robust and challenging multi-task language understanding benchmark. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, External Links: Link Cited by: §4.
- AgentDropout: dynamic agent elimination for token-efficient and high-performance LLM-based multi-agent collaboration. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 24013–24035. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §1, §2.
- Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho (Eds.), External Links: Link Cited by: §4, Table 1.
- AutoGen: enabling next-gen LLM applications via multi-agent conversations. In First Conference on Language Modeling, External Links: Link Cited by: §1.
- A-mem: agentic memory for LLM agents. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §2.
- R2-router: a new paradigm for LLM routing with reasoning. In Forty-third International Conference on Machine Learning, External Links: Link Cited by: §2.
- Memory-r1: enhancing large language model agents to manage and utilize memories via reinforcement learning. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 12805–12825. External Links: Link, Document, ISBN 979-8-89176-390-6 Cited by: §2.
- Qwen2.5 technical report. CoRR abs/2412.15115. External Links: Link, Document, 2412.15115 Cited by: §4.
- MasRouter: learning to route LLMs for multi-agent systems. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 15549–15572. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §1, §14, §2, §4, §4.1, Table 1.
- Cut the crap: an economical communication pipeline for LLM-based multi-agent systems. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §1, §2.
- AFlow: automating agentic workflow generation. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §2, §4, Table 1.
- Exploring collaboration mechanisms for LLM agents: a social psychology view. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 14544–14607. External Links: Link, Document Cited by: §1.
- SecureRouter: encrypted routing for efficient secure inference. CoRR abs/2604.15499. External Links: Link, Document, 2604.15499 Cited by: §2.
- GPTSwarm: language agents as optimizable graphs. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024, R. Salakhutdinov, Z. Kolter, K. A. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (Eds.), Proceedings of Machine Learning Research, Vol. 235, pp. 62743–62767. External Links: Link Cited by: §1, §2.
7 Algorithm
We provide detailed pseudocode for Gated-Memory Routing in Algorithm 1.
8 Representations, Encoders, and Training Interface
This appendix specifies the encoders and representations underlying the routing modules of Section 3, and how the policy is trained against the frozen backbones.
Sentence encoder.
All text is embedded by a single frozen sentence encoder (all-MiniLM-L6-v2), producing -dimensional vectors; following the convention of Section 3, for any text , so is the query embedding and the embedding of reasoning step . The same encoder embeds the query, the role and backbone descriptions, and every agent response. Its parameters are not updated during training, and response embeddings are detached before they enter memory, so no gradient flows back through the encoder or through the agent outputs.
Role and backbone latents.
Each role profile and each backbone description is embedded once by and mapped to a continuous latent by a variational encoder (a small VAE with a -dimensional latent), yielding the role embeddings and backbone embeddings used by the role allocator (Section 3.3) and LLM router (Section 3.4). The reconstruction and KL terms of these encoders form the regularizer in the training objective.
State encoder.
The state is encoded into the context vector used by the routing modules. The query is linearly projected to a -dimensional vector , and each retained step becomes a token formed from its role latent and backbone latent , additively modulated by a sigmoid gate over the response embedding , so that response content, and not only routing identity, informs the state. A two-layer Transformer encoder contextualizes the tokens , the query attends over them to produce a pooled history vector , and . The execution memory is therefore encoded as an ordered set of per-step records rather than a single averaged vector.
Memory-record embedding.
The retrieval gate (Section 3.5) forms the query vector and, for each stored record , an embedding , where is a learned projection of . The retrieval logit is the scaled cosine similarity with learnable scale and bias , and record is admitted by . Memory holds at most records, so ranges from empty to entries, with setting its typical length.
Write-gate similarities.
Training interface.
The backbones are frozen and queried as an environment: at each step the selected model is called through an OpenAI-compatible vLLM server, and the returned text is embedded and detached as described above. The router is optimized purely by the score-function (policy-gradient) estimator: the group-relative advantage (Section 3.8) weights the trajectory log-probability , where collects the step- retrieval decisions. Gradients reach only the routing, gating, halting, and embedding parameters; a single optimizer updates all of them jointly, and the backbones are never differentiated through.
9 Cost Calculation for Open-Weight Backbones
Because the five backbones are open-weight and served locally, they do not carry a public per-token price. To give the cost-aware objective a hardware-independent and monotonic cost signal, we set each model’s per-token reference price proportional to its (active) parameter count in billions: an input rate of and an output rate of per million tokens, with output weighted more heavily to reflect the higher cost of autoregressive decoding. For dense transformers, per-token inference compute (FLOPs) grows approximately linearly with , so pricing each token by model size makes the reported cost proportional to a size-weighted token count: a hardware-independent proxy for relative inference compute across the pool, rather than the dollar price of any specific API or the wall-clock cost on particular hardware. Both the training cost term of Section 3.8 and the reported test cost of a trajectory are the additive sum of its per-step backbone costs, where is the cost of step ’s call (its input and output tokens times these rates). As a relative-compute proxy it abstracts away attention cost that scales with context length, batching, and memory bandwidth. Table 4 lists the backbone pool and the resulting rates.
| Backbone | Size | Cost in/out |
| llama-3.2-3B | 3B | 0.009 / 0.03 |
| llama-3.1-8B | 8B | 0.024 / 0.08 |
| mistral-nemo-12B | 12B | 0.036 / 0.12 |
| qwen-2.5-14B | 14B | 0.042 / 0.14 |
| qwen-2.5-32B | 32B | 0.096 / 0.32 |
To corroborate the size-weighted proxy with hardware-relevant measurements, Table 5 reports, for HumanEval and MBPP, an approximate FLOPs count and the batch-amortized wall-clock time per query for our method and the two strongest baselines. FLOPs are estimated from the per-call token records as FLOPs per processed token Kaplan et al. (2020), summed over all calls with the parameter count of the model serving each call; this omits the sequence-length-dependent attention term. Wall-clock time is the end-to-end inference time for the test set divided by the number of queries, measured under the same serving configuration for all methods. Our method has the lowest FLOPs and wall-clock time on both benchmarks, and the ordering matches the proxy.
| HumanEval | MBPP | |||
| Method | PFLOPs/query | Time/query (s) | PFLOPs/query | Time/query (s) |
| Ours | 0.107 | 1.5 | 0.145 | 2.6 |
| MASRouter | 0.200 | 4.4 | 0.289 | 6.4 |
| Puppeteer-32B | 0.144 | 3.0 | 0.163 | 3.2 |
10 Per-Dataset Accuracy and Cost
Tables 6 and 7 report per-method accuracy and total test-set cost on GSM-Hard and HumanEval, the data underlying the Pareto comparison of Figure 3. Cost is the size-based reference cost of Appendix 9. Figure 5 shows the GSM-Hard accuracy–cost frontier, the companion to the HumanEval frontier in the main text.
| Method | LLM | Acc. (%) | Cost |
| Single LLM | llama-3.2-3B | 25.85 | 0.025 |
| llama-3.1-8B | 39.87 | 0.046 | |
| mistral-nemo-12B | 30.11 | 0.050 | |
| qwen-2.5-14B | 64.58 | 0.047 | |
| qwen-2.5-32B | 61.52 | 0.071 | |
| CoT | qwen-2.5-14B | 67.52 | 0.076 |
| qwen-2.5-32B | 64.56 | 0.167 | |
| Complex-CoT | qwen-2.5-14B | 66.76 | 0.108 |
| qwen-2.5-32B | 65.32 | 0.230 | |
| MacNet-Chain | qwen-2.5-14B | 59.75 | 0.629 |
| qwen-2.5-32B | 61.93 | 3.421 | |
| MacNet-Tree | qwen-2.5-14B | 60.32 | 0.696 |
| qwen-2.5-32B | 62.31 | 2.433 | |
| MacNet-Complete | qwen-2.5-14B | 59.00 | 0.933 |
| qwen-2.5-32B | 62.22 | 4.099 | |
| AFlow | qwen-2.5-14B | 58.24 | 0.134 |
| qwen-2.5-32B | 67.05 | 2.165 | |
| Puppeteer | qwen-2.5-14B | 65.91 | 0.232 |
| qwen-2.5-32B | 68.75 | 0.445 | |
| MASRouter | LLM Pool | 66.00 | 1.012 |
| Ours | LLM Pool | 70.55 | 0.587 |
| Method | LLM | Acc. (%) | Cost |
| Single LLM | llama-3.2-3B | 62.79 | 0.001 |
| llama-3.1-8B | 69.78 | 0.0025 | |
| mistral-nemo-12B | 68.22 | 0.002 | |
| qwen-2.5-14B | 82.95 | 0.003 | |
| qwen-2.5-32B | 84.37 | 0.009 | |
| CoT | qwen-2.5-14B | 82.17 | 0.004 |
| qwen-2.5-32B | 85.15 | 0.012 | |
| Complex-CoT | qwen-2.5-14B | 80.62 | 0.005 |
| qwen-2.5-32B | 85.47 | 0.014 | |
| MacNet-Chain | qwen-2.5-14B | 86.82 | 0.019 |
| qwen-2.5-32B | 83.80 | 0.059 | |
| MacNet-Tree | qwen-2.5-14B | 82.95 | 0.024 |
| qwen-2.5-32B | 86.02 | 0.084 | |
| MacNet-Complete | qwen-2.5-14B | 84.50 | 0.033 |
| qwen-2.5-32B | 86.05 | 0.119 | |
| AFlow | qwen-2.5-14B | 85.27 | 0.014 |
| qwen-2.5-32B | 84.50 | 0.224 | |
| Puppeteer | qwen-2.5-14B | 83.59 | 0.029 |
| qwen-2.5-32B | 85.16 | 0.047 | |
| MASRouter | LLM Pool | 85.16 | 0.057 |
| Ours | LLM Pool | 89.84 | 0.032 |
11 Variance across Training Seeds
Table 1 reports one training run per method. To quantify run-to-run variability, we retrained our method and the two strongest baselines with three independent seeds each and report the per-benchmark mean and sample standard deviation in Table 8, together with the overall mean (the macro-average over benchmarks for each seed, then mean and standard deviation across seeds). Our method has the highest mean on all five benchmarks and improves the overall mean by points over MASRouter and over Puppeteer-32B. The ranking is identical under every seed: the lowest overall score among our seeds () exceeds the highest among Puppeteer’s () and MASRouter’s ().
| Method | GSM-Hard | MATH | MBPP | HumanEval | MMLU-Pro | Overall |
| MASRouter | 65.37 0.55 | 74.37 2.28 | 76.87 2.21 | 81.51 3.94 | 64.09 4.83 | 72.44 1.88 |
| Puppeteer-32B | 66.98 1.87 | 77.17 4.41 | 76.41 2.29 | 86.72 2.06 | 68.98 1.01 | 75.25 0.52 |
| Ours | 69.76 1.45 | 78.37 1.88 | 80.40 1.21 | 88.54 1.19 | 69.62 1.17 | 77.35 0.63 |
12 Gate Behavior at Test Time
Table 9 checks that the trained gates do not collapse to trivial policies. The write rate is the fraction of reasoning steps committed to memory; the retrieved fraction is the mean number of retrieved items divided by the number available at that step (the memory is empty at step 0). Write rates lie strictly inside on all three benchmarks, and the retrieval gate becomes more selective as memory grows rather than exposing all or none of it. The per-step entropy of the gate distributions also stays well above zero throughout training on all three benchmarks.
| Benchmark | Mean depth | Write rate | Retrieved fraction, steps 1–5 |
| GSM-Hard | 2.92 | 39.6% | 16 / 24 / 14 / 12 / 11% |
| MBPP | 3.79 | 32.8% | 9 / 30 / 20 / 15 / 12% |
| HumanEval | 3.30 | 36.5% | 91 / 48 / 33 / 24 / 21% |
13 Halting and Training-Time Computation
Adaptive halting also reduces training-time computation, since shorter trajectories mean fewer rollout steps per update. Table 10 compares the mean realized depth and the total number of rollout steps over training with and without halting.
| Benchmark | Depth, halting | Depth, no halting | Rollout steps, halting vs. none |
| GSM-Hard | 3.4 | 6.0 | 26.5k vs. 47.9k () |
| HumanEval | 3.3 | 6.0 | 6.3k vs. 11.5k () |
14 Agent Role Profiles
The role allocator selects among the same heterogeneous role profiles used by MASRouter Yue et al. (2025), grouped into three domains: code-oriented roles, commonsense/knowledge roles, and mathematical-reasoning roles. Each profile is a short natural-language description that we encode and project into the latent role space (Section 3.3). At instantiation, this description, together with the role’s reasoning strategy and the benchmark’s output-format requirements, forms the agent’s system prompt, while the user prompt supplies the query and any retrieved context. The descriptions are listed below. The imbalance across domains reflects the original MASRouter catalog, which we adopt unchanged so that both methods select from the same roles. To test dependence on the full catalog, we randomly subsampled it to roles, keeping roughly half of each domain, and retrained the full system: accuracy is on GSM-Hard and on HumanEval, against and with all roles, so performance does not hinge on the full role set.
Code.
- •
AlgorithmDesigner. Specifies the design of the algorithm, including explanations, usage instructions, and API references, optionally giving pseudocode for the main logic; replies concisely.
- •
ProgrammingExpert. A programming expert who, given a function signature and docstring, writes the full implementation (restating the signature) in a single Python code block.
- •
BugFixer. A programming expert who restates the signature and returns a corrected full implementation in a Python code block.
- •
ReflectProgrammer. A programming expert who reflects on prior attempts and returns a full implementation in a Python code block.
- •
PlanSolver. Produces the pseudocode of the target function.
- •
ProjectManager. Oversees the overall code structure, suggests optimal design patterns for maintainability and flexibility, and avoids over-engineering; replies concisely.
- •
TestAnalyst. Identifies problems in the current code from test data and feedback, supplies special cases and boundary conditions to watch, and points out potential errors; replies concisely.
Commonsense and Knowledge.
- •
KnowledgeExpert. A knowledgeable question-answering expert who analyzes step by step and selects the correct answer.
- •
Reflector. Re-examines the question-answering process step by step and selects the correct answer.
- •
Critic. Points out potential issues in other agents’ analyses point by point and gives a critical opinion before the final result.
- •
Scientist. A scientist with natural-science knowledge who provides a thorough solving process, including necessary proofs and explanations.
- •
Economist. An experienced economist (macroeconomics, microeconomics, financial markets) who gives well-reasoned, evidence-based answers.
- •
Historian. Analyzes cultural, economic, political, and social events from primary sources to reason about the past.
- •
WikiSearcher. Lists the key Wikipedia entities that should be looked up to solve the problem.
Mathematical Reasoning.
- •
MathSolver. A math expert who produces a solving process from the hints supplied by other agents.
- •
Mathematician. A mathematician skilled at math games, arithmetic, and long-horizon planning.
- •
MathTeacher. Teaches the solution step by step as if to a student.
- •
MathAnalyst. First derives the solution symbolically (variables as letters), then substitutes values to compute the result.
- •
Inspector. Checks whether the problem-solving logic, calculations, and any accompanying code are correct and consistent, then gives its own step-by-step solution.
- •
AlgorithmEngineer. Integrates step-by-step reasoning with Python code to solve the problem.
- •
ProgrammingExpert. Analyzes the problem and writes functions, combining reasoning with Python code.
- •
SoftwareDeveloper. Designs efficient solutions and provides clear, concise functions.
- •
Engineer. An experienced engineer who solves the problem from engineering knowledge.
- •
Scientist. Provides a detailed solving process with necessary proofs and explanations.
- •
Economist. Applies economic reasoning to give evidence-based answers.
- •
CertifiedAccountant. Analyzes financial problems and returns correct calculations and solutions.
15 Aggregator Prompt
Once the halting controller stops the trajectory, the Aggregator LLM receives the query and the retained execution records and produces the final answer. The aggregator is instantiated with the backbone the router selected most frequently during the trajectory (Section 4.1). Across benchmarks the aggregator shares a common framing in its system prompt, namely to weigh the analyses and results of the other agents, identify errors, and commit to a single most-reliable answer, while the user prompt enforces the benchmark-specific output format. The boxes below paraphrase the system and user prompts for each benchmark.
16 Backbone LLM Profiles
The LLM router selects among the five open-weight backbones through their natural-language capability descriptions, which a variational encoder maps into the latent backbone space (Section 3.4). The descriptions below are the text encoded for each model; each description also states the model’s reference input/output price (Table 4), so the encoded text covers both capability and cost.
- •
llama-3.2-3B. Meta’s compact 3-billion-parameter, text-only instruction-tuned model with a 128k-token context window. The cheapest and fastest choice in the pool, with solid general reasoning and instruction following; suited to easy queries where quality is less critical.
- •
llama-3.1-8B. Meta’s widely used 8-billion-parameter instruction-tuned model, offering solid general-purpose reasoning and instruction following at very low cost; a reliable baseline across a broad range of tasks.
- •
mistral-nemo-12B. Mistral AI’s 12-billion-parameter dense model built in partnership with NVIDIA, with strong multilingual performance, a 128k-token context window, and solid general reasoning and coding at modest cost; a reliable mid-size workhorse.
- •
qwen-2.5-14B. Alibaba Cloud’s 14-billion-parameter dense instruction-tuned model with a 128k-token context window; strong on mathematics, coding, and general reasoning, punching above its size class as a capable mid-tier option for moderately hard queries.
- •
qwen-2.5-32B. Alibaba Cloud’s flagship dense model in this pool, with frontier-level performance among open-weight models in its size class on mathematics, coding, and complex reasoning; reserved for the hardest queries where smaller models fall short.