Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems
Abstract
Adapting the communication topology of an LLM multi-agent system to each query improves both accuracy and efficiency, yet current designers treat this as conditional graph generation: a variational, autoregressive, or diffusion decoder searches the adjacency space, and a graph-network proxy trained on utility and a structural cost such as edge count ranks the sampled candidates. We argue that this formulation is misaligned with the problem. Empirically, topologies that survive a reward filter collapse to about six distinct graphs even when the codebook capacity grows from 8 to 64; edge count is negatively correlated with measured token consumption (Pearson ), so sparsifying the graph makes inference more expensive; and a message-passing scorer over agent-profile nodes is adjacency-invariant whenever agents share a profile—the default configuration of published benchmarks—so it cannot rank candidates at all in that regime. These three facts motivate Codebook Agent: a vector-quantized autoencoder compresses successful topologies into a query-independent 16-entry codebook; a reward-weighted MLP maps the query embedding to a distribution over codes; and an MLP proxy that reads the flattened adjacency, regressed on measured utility and per-task normalized token cost, reranks the top decoded candidates in a single batched forward pass. With no iterative search and no message passing at test time, Codebook Agent is the most accurate method on all six benchmarks we compare (84.6 average against 83.0 for the strongest prior designer), emits a topology in 2.4 ms, and uses 21.9–33.2% fewer LLM tokens. Our code is available here: https://github.com/jinxiy1104/CodebookAgent.
University of California, Los Angeles
1 Introduction
Multi-agent systems built from LLM agents solve reasoning, question answering, and coding tasks by exchanging intermediate outputs over a communication topology (Wu et al. 2024; Hong et al. 2024; Qian et al. 2024; Du et al. 2024). That topology is not bookkeeping: it decides which agents see which messages, and both accuracy and token consumption move with it (Qian et al. 2025; Zhuge et al. 2024; Zhang et al. 2025b). The field has therefore progressed from hand-designed structures to learned ones, and most recently to query-conditioned generators that emit a fresh topology for every input (Zhang et al. 2025c; Jiang et al. 2026).
The latest designers share a natural but costly recipe. A variational (Zhang et al. 2025c), autoregressive (Li et al. 2026), or diffusion decoder (Jiang et al. 2026) searches the adjacency space; a graph network over agent-profile nodes then ranks candidates with a utility head and a structural cost head, typically the edge count (Zhuge et al. 2024; Zhang et al. 2025b; Zhang et al. 2025c; Jiang et al. 2026). The recipe looks principled—generate, then score—yet it quietly assumes that the useful design space is large enough to need a generative decoder, that tracks inference cost, and that message passing can tell two topologies apart. Our measurements reject all three assumptions.
Figure 1 sketches the consequence. First, the useful design space is a short list, not a manifold: topologies that solve their task fall into about six distinct codes even as codebook capacity grows from 8 to 64, and the best fixed topology stays within 1.4 accuracy points of every generated one we measure—so capacity spent modeling adjacency is spent where the problem does not live. Second, the structural cost surrogate is inverted: edge count correlates with measured tokens at , because sparse communication yields longer completions, so minimizing maximizes the bill it was introduced to reduce. Third, on the homogeneous teams that dominate published benchmarks, a message-passing scorer over profile nodes is adjacency-invariant: every candidate receives the same score, the ranking does not exist, and hundreds of guided diffusion steps reproduce a constant.
If the problem is to choose among a few graphs under a measured cost, the matching designer is not another decoder. Codebook Agent therefore amortizes topology design into three feed-forward pieces: a vector-quantized autoencoder (van den Oord et al. 2017) that indexes successful topologies into 16 codes, a reward-weighted MLP that maps the query embedding to a distribution over those codes, and an MLP proxy that reads the flattened adjacency and is regressed on measured utility and per-task normalized token cost. At test time we decode the top codes, score them in one batched proxy call, and execute the winner—2.4 ms, with no sampling loop and no message passing.
Our contributions:
- ❶
We characterize topology design rather than assume its form: the reward-surviving space collapses to about six graphs independently of model capacity, the common structural cost surrogate is inverted w.r.t. measured tokens, and message-passing proxies are topology-blind on homogeneous teams.
- ❷
We propose Codebook Agent, the minimal designer these facts admit, a discrete codebook, a reward-weighted code predictor, and one batched pass of an execution-grounded proxy, with no iterative search and no message passing.
- ❸
2 Related Work
Communication topologies for LLM agents.
Early frameworks fix the topology by hand (Hong et al. 2024; Qian et al. 2024; Wu et al. 2024; Chen et al. 2024c; Du et al. 2024). A second line optimizes it: GPTSwarm learns edge distributions with policy gradients (Zhuge et al. 2024), DyLAN selects agents dynamically (Liu et al. 2024), search-based systems explore agentic designs offline (Hu et al. 2025; Zhang et al. 2025d), pruning methods sparsify a fixed structure to save tokens (Zhang et al. 2025b; Wang et al. 2025), and Optima trains the agents for efficiency (Chen et al. 2025). Closest to us are query-conditioned generators (Zhang et al. 2025c; Jiang et al. 2026). We share their problem statement, one topology per query, and differ in what we assume about it: they model the adjacency space as something to be generated and scored by a graph network, and we show it is a short list to be selected from and scored by measured cost.
Graph generation and discrete latents.
Deep graph generators are predominantly iterative, whether autoregressive (You et al. 2018) or score-based (Jo et al. 2022; Vignac et al. 2023; Zhou et al. 2024), inheriting a sampling cost (Ho et al. 2020) that faster samplers reduce but do not remove (Song et al. 2021); one-shot VAE decoders exist for small graphs (Simonovsky and Komodakis 2018). We need one small graph per query under a latency constraint, so we discretize the output space with vector quantization (van den Oord et al. 2017; Razavi et al. 2019) instead of sampling from a continuous process. VQGraph tokenizes local structure to improve GNN-to-MLP distillation (Yang et al. 2024; Zhang et al. 2022; Tian et al. 2023; Wu et al. 2023; Hinton et al. 2015); we borrow its structure-token regularizer, but our proxy has no teacher to distill from: on homogeneous teams a message-passing teacher is constant per query, so we train the MLP directly on measured rewards.
Cost of multi-agent inference.
Multi-agent pipelines multiply LLM calls, and returns diminish or invert as calls grow (Li et al. 2024; Chen et al. 2024a; Smit et al. 2024), with many observed failures attributable to coordination rather than single-agent ability (Cemri et al. 2025). Cost-aware model selection addresses the single-call case (Chen et al. 2024b). These results motivate treating measured tokens, not a structural surrogate, as the cost objective of topology design.
3 Problem Setup and Background
Topology design.
A team of agents has profile texts embedded as by a frozen sentence encoder (). A topology is a directed adjacency matrix without self-loops, where shows the output of agent to agent , and a decision node aggregates the final answers following GDesigner (Zhang et al. 2025c). Executing the team on query under returns a utility and a token cost . Writing for the query embedding, the goal is a generator mapping to a topology maximizing , with a normalized cost and . All methods in this paper are trained on the same records: per benchmark, 50 training tasks executed under 6 fixed topologies (complete, chain, star, and three Erdos-Renyi samples) with real LLM agents, giving 300 tuples .
Two design axes.
A query-conditioned designer is determined by two independent choices, which we vary separately in the experiments. Axis A, the candidate generator, maps to one or more topologies; the incumbent choice is an iterative or continuous decoder over the adjacency space, variational (Zhang et al. 2025c), autoregressive (Li et al. 2026), or a diffusion reverse process with steps and candidates per step (Jiang et al. 2026), which we measure at 301 to 396 ms per query. Axis B, the scorer, ranks candidates; the incumbent choice, shared across the learned-design literature (Zhuge et al. 2024; Zhang et al. 2025b; Zhang et al. 2025c; Jiang et al. 2026), is a message-passing network over a graph whose nodes carry the profile embeddings and whose edges are the candidate :
| (1) |
with a two-dimensional head regressed onto utility and a structural cost label, most commonly the edge count . We implement Eq. (1) as a GAT (Veličković et al. 2018) and treat it as one level of Axis B, so it can be swapped into our own pipeline with everything else held fixed. The introduction’s three observations—a collapsed design space, an inverted cost surrogate, and adjacency-blind message passing on homogeneous teams—determine how we instantiate both axes below.
4 Method
Codebook Agent instantiates Axis A as a discrete codebook with a query-conditioned prior, and Axis B as an MLP over flattened adjacencies trained on measured utility and token cost. Let be the execution records of Section 3. Define the composite reward
| (2) |
with : is the token count normalized by the mean token count of the same task, so difficulty cancels and within-task rankings remain. Using measured tokens rather than is essential: on our records edge count correlates with at , so an edge-count head pushes the system toward more expensive graphs. The codebook, code predictor, and reranking proxy are trained on with no further LLM calls. Figure 2 shows offline collection, the three training stages, and one-pass test-time generation.
4.1 Topology Codebook
If reward-surviving topologies occupy only a handful of distinct graphs, the generator should index that short list rather than search the adjacency manifold. We therefore compress observed topologies with a vector-quantized autoencoder. Flatten the off-diagonal entries and encode
| (3) | ||||
with and codebook size . The codebook is updated by exponential moving averages (decay ), with commitment coefficient , straight-through gradients, and reinitialization of codes unused for 50 batches (van den Oord et al. 2017; Razavi et al. 2019). A mirror-image decoder maps to edge logits ; the diagonal is fixed to zero and the graph remains directed. Writing , the training objective is
| (4) |
where weights positive entries by the zero-to-one ratio of the training set to counter edge sparsity and is the stop-gradient. We fit Eq. (4) only on records with , so the codebook stores topologies that solved their task. Neither encoder nor decoder is conditioned on the query: after training, code decodes to the binary topology
| (5) |
and all query dependence is carried by the predictor below. We train for 200 epochs with Adam at learning rate and batch size 16.
4.2 Reward-Weighted Code Prediction
At test time there is no target topology to encode, so we learn a conditional prior over codes. Passing every training record through the frozen encoder and quantizer yields indices . For each distinct condition we form the soft target
| (6) |
which concentrates mass on codes whose topologies earned high reward under Eq. (2), so the prior already prefers cheap, successful graphs. The predictor is an MLP (, dropout ) trained by soft cross-entropy
| (7) |
This is the only place the query enters generation, so the map from query to candidate topologies is a single matrix pipeline.
4.3 Execution-Grounded MLP Proxy
The scorer must (i) read the adjacency itself and (ii) regress measured tokens, not . The first requirement is forced by the homogeneous-team regime: when , any message-passing layer of Eq. (1) aggregates identical features over self-loops, so every node state—and thus the pooled score—is independent of . The same holds for mean and sum aggregators; there is nothing to distill from such a teacher on anonymous teams. We therefore score candidates with an MLP over the flattened adjacency,
| (8) |
and regress directly on the measured targets,
| (9) |
Because covers the topology space thinly, we regularize with a structure-token objective in the spirit of VQGraph (Yang et al. 2024). A separate reconstruction-only quantizer with codes is fit over an augmented pool of topologies (static families, Erdős–Rényi graphs at eight densities, single-edge perturbations, and Bernoulli samples of smoothed edge profiles, capped at 100 per condition). Writing for the soft code assignment of a pool graph, an auxiliary head predicts under
| (10) |
and the total proxy loss is with . The augmented pool contributes only to Eq. (10); reward supervision comes exclusively from real executions. The proxy is trained for 200 epochs with Adam at learning rate and batch size 256.
4.4 One-Pass Inference
Given a query embedding , let be the top codes under , and let after removing duplicates under Eq. (5). The selected topology is
| (11) |
where is evaluated for all survivors in one batched forward pass. Generation therefore costs one predictor pass, at most decoder passes, and one dense proxy pass—a fixed number of matrix products, with no sampling loop, no sparse graph construction, and no message passing on the test-time path.
| Method | General/ Math reasoning | Code generation | Avg. | ||||
|---|---|---|---|---|---|---|---|
| GSM8K | MATH | MultiArith | SVAMP | MBPP | HumanEval | ||
| single-agent prompting | |||||||
| Vanilla | 87.0 | 47.6 | 97.2 | 88.5 | 72.4 | 73.1 | 77.63 |
| CoT | 86.50.5 | 48.00.4 | 96.70.5 | 89.00.5 | 74.01.6 | 74.41.3 | 78.100.47 |
| SC-CoT | 87.0 0.0 | 49.62.0 | 97.2 0.0 | 89.51.0 | 75.02.6 | 75.01.9 | 78.881.25 |
| multi-agent collaboration | |||||||
| LLM-Debate | 89.02.0 | 50.02.4 | 97.80.6 | 90.52.0 | 76.44.0 | 75.01.9 | 79.782.15 |
| LLM-Blender | 87.50.5 | 48.40.8 | 97.80.6 | 89.00.5 | 75.63.2 | 75.01.9 | 78.881.25 |
| DyLAN | 89.52.5 | 50.02.4 | 97.80.6 | 90.52.0 | 77.04.6 | 76.33.2 | 80.182.55 |
| learned agent and topology design | |||||||
| GPTSwarm | 88.51.5 | 49.62.0 | 97.2 0.0 | 90.52.0 | 76.03.6 | 75.62.5 | 79.571.94 |
| ADAS | 85.51.5 | 44.43.2 | 96.70.5 | 88.00.5 | 71.01.4 | 70.62.5 | 76.031.60 |
| AFLOW | 90.53.5 | 52.44.8 | 96.70.5 | 90.01.5 | 78.46.0 | 76.93.8 | 80.823.19 |
| MaAS | 91.54.5 | 53.66.0 | 97.20.0 | 91.53.0 | 79.06.6 | 76.93.8 | 81.623.99 |
| AgentDropout | 90.03.0 | 51.64.0 | 98.31.1 | 91.02.5 | 78.05.6 | 75.62.5 | 80.753.12 |
| G-Designer | 91.54.5 | 52.44.8 | 98.31.1 | 91.53.0 | 79.67.2 | 76.93.8 | 81.704.07 |
| ARG-Designer | 92.55.5 | 54.06.4 | 98.91.7 | 92.54.0 | 80.07.6 | 77.54.4 | 82.574.94 |
| TopoDIM | 93.06.0 | 55.07.4 | 98.91.7 | 92.54.0 | 81.08.6 | 77.54.4 | 82.985.35 |
| GTD | 93.56.5 | 55.57.9 | 98.21.0 | 93.04.5 | 80.48.0 | 77.54.4 | 83.025.39 |
| Codebook Agent (ours) | 94.87.8 | 56.58.9 | 99.42.2 | 95.46.9 | 83.511.1 | 78.15.0 | 84.626.99 |
5 Experiments
We design the evaluation to stress-test the three claims of the introduction with four questions: (Q1) Does amortized Codebook Agent match or beat iterative query-conditioned designers on accuracy, or does indexing a short codebook trade quality for speed? (Q2) Is the useful design space truly a short list—do fixed topologies already sit within a point of generated ones, and does extra codebook capacity go unused? (Q3) Does an execution-grounded MLP over cut tokens where the incumbent edge-count GNN cannot rank homogeneous teams, and is random selection enough? (Q4) Does one-pass codebook generation deliver the latency and token savings of Figure 1 without sacrificing transfer across LLM backends?
5.1 Setup
Benchmarks. Reasoning: GSM8K (Cobbe et al. 2021), MATH (Hendrycks et al. 2021b), MultiArith (Roy and Roth 2015), SVAMP (Patel et al. 2021). Code: MBPP (Austin et al. 2021), HumanEval (Chen et al. 2021) (Pass@1). Transfer: MMLU (Hendrycks et al. 2021a) under Qwen-3-8B (Table 2). Evaluation uses the fixed test splits of GSM8K (1314), MATH (500), MultiArith (180), SVAMP (1000), MBPP (500), HumanEval (160), and MMLU (1530); design-axis accuracy/token numbers use 200 tasks per benchmark (full MultiArith/HumanEval where smaller). Teams. Math-format: MathSolver with FinalRefer. Code: Programming Expert with FinalRefer. MMLU: Knowledgeable Expert. Homogeneous profiles are the default (six of seven design-axis settings); HumanEval also runs a heterogeneous four-role team. Backbone gpt-4o-mini (temp. , top- , max tokens ) unless stated; embeddings are frozen all-MiniLM-L6-v2 ().
Baselines. Single-agent: Vanilla, CoT (Wei et al. 2022), SC-CoT (Wang et al. 2023). Multi-agent: LLM-Debate (Du et al. 2024), LLM-Blender (Jiang et al. 2023), DyLAN (Liu et al. 2024). Learned design: GPTSwarm (Zhuge et al. 2024), ADAS (Hu et al. 2025), AFLOW (Zhang et al. 2025d), MaAS (Zhang et al. 2025a), AgentDropout (Wang et al. 2025), G-Designer (Zhang et al. 2025c), ARG-Designer (Li et al. 2026), TopoDIM (Sun et al. 2026), GTD (Jiang et al. 2026). Metrics. Accuracy (Pass@1 for code); median topology-generation latency (3 warmup 50/100 timed calls on one GPU); mean LLM tokens/query; end-to-end wall clock. Single evaluation run per configuration.
5.2 Main Results
Table 1 reports gpt-4o-mini accuracy; Table 2 the Qwen-3-8B transfer; Figures 3 and 4 the latency and token frontiers. We highlight four observations corresponding to Q1–Q4.
(Q1) Amortized selection clears the prior designer frontier; it does not trade accuracy for a codebook. Table 1 puts Codebook Agent best on every column at 84.62 average: over Vanilla (77.63), over DyLAN (80.18), and over the strongest prior topology designer GTD (83.02). The Vanilla gap is smallest on MultiArith (; near ceiling) and largest on MBPP (). Against GTD the gains are spread, not concentrated: GSM8K, MATH, MultiArith, SVAMP, MBPP, HumanEval. Reading by family: single-agent prompting floors the high 70s; multi-agent collaboration lifts into the low 80s under hand-designed structure; query-conditioned generators (G-Designer through GTD) form the prior frontier at . Codebook Agent sits above that frontier on all six benchmarks—so indexing a short list of graphs is not a speed concession; it matches or exceeds the iterative adjacency search it replaces.
(Q2) The useful design space is a short list, and which fixed graph wins is not knowable a priori. Fixed topologies need no designer, so they measure how much of the payoff lives in the choice itself. Across seven settings the best fixed family (fully connected / chain / star) stays within 1.4 accuracy points of every generated topology we measure, at comparable tokens; on MATH, fully connected is above every generated configuration. Capacity spent modeling the adjacency manifold is therefore spent where the problem does not live—the premise of Axis A as a codebook. The choice is nonetheless load-bearing: which fixed graph wins changes with the setting (fully connected on MATH, star on homogeneous HumanEval, chain on MMLU), with gaps up to 4.3 accuracy points and a token factor (chain vs. fully connected on MATH). Identifying the winner for a new setting requires real LLM executions—the same cost class as our one-time 300-record collection—so a designer earns its keep by amortizing that selection per query. Figure 5(b) closes the loop on capacity: as grows from 4 to 64, accuracy moves by at most 1.5 points (GSM8K) / 2.5 (HumanEval), and for every the encoder uses at most six codes. Collapse is a property of reward-surviving topologies, not of an undersized quantizer.
(Q3) Measured-token MLP scoring is where cost is won; the incumbent GNN cannot rank anonymous teams. We hold Axis A (codebook) and the candidate set fixed, and swap only Axis B. On six homogeneous settings Eq. (1) is constant in , so every candidate in the short list receives the same score; the edge-count head is then the only separator, and because correlates with tokens at , it points toward expensive graphs. Empirically this incumbent scorer, reranking our codebook candidates, costs 1711 tokens/query on GSM8K and 918 on homogeneous HumanEval—worse than uniform random over the same candidates (1249 / 750), and not comparable to the full incumbent pipeline of Q4. Heterogeneous HumanEval is the exception on accuracy (incumbent 78.8 vs. ours 78.1); our largest token saving () is also there. Replacing our proxy with random, or dropping rerank and taking the predictor’s top-1 (Figure 5c), shows the ranking is load-bearing for cost: random costs – more tokens at equal or lower accuracy; top-1 is already cheap (913 tokens on GSM8K) and rerank adds – accuracy points. The chain fixed topology is the extreme positive control for the inverted surrogate: sparsest connected, most token-expensive on every math-format benchmark.
(Q4) One-pass generation is faster and – cheaper in tokens, and the ordering transfers. Figure 3 places MATH accuracy against median generation latency: iterative generators sit at 301–396 ms; Codebook Agent at 2.4–2.5 ms (–). The gap is structural— sparse GNN evaluations vs. one predictor pass, decodes, and one batched MLP call (Eq. (11))—and design falls below of end-to-end latency. Figure 4 shows the accuracy–token frontier on MATH and MBPP; Codebook Agent occupies the lower-right corner. Against the full incumbent pipeline (iterative decoder edge-count GNN), mean tokens drop – (e.g. GSM8K, MATH, HumanEval-homo, HumanEval-het), and wall clock follows (MATH min; MMLU min). Table 2 repeats the accuracy comparison with Qwen-3-8B: absolute numbers fall, but Codebook Agent still leads at 74.0 vs. GTD 72.7 / G-Designer 72.1, with MATH and MMLU over GTD—so the codebook is not an artifact of one API.
Overall. Q1–Q4 close the loop opened in the introduction: when reward-surviving topologies are few, accuracy is preserved by selecting among them (Q1–Q2), cost is won by scoring measured tokens rather than on anonymous teams (Q3), and amortization makes that selection essentially free at test time while transferring across backends (Q4). The ablations below isolate each mechanism under controlled Axis A/B swaps.
5.3 Ablations
We ablate the mechanisms behind Eqs. (4)–(11) and test whether the main-table ordering survives a backbone change. Figure 5 varies four controls on GSM8K and homogeneous HumanEval under a fixed training budget; Table 2 repeats the accuracy comparison with Qwen-3-8B agents.
(a) Team size . Figure 5(a): retrain per on the same 50-task budget. GSM8K stays flat (–); HumanEval rises . Coding benefits from a wider expert pool, but the same index-and-select pipeline tracks that gain without architectural change—so the method is not a four-node trick, and topology selection remains useful as the team grows.
(b) Codebook size . Figure 5(b): moves accuracy by (GSM8K) / (HumanEval). For every the encoder uses at most six codes—extra capacity stays idle under EMA with dead-code reset—so the design space, not the quantizer, is what collapsed. Searching the full adjacency at test time cannot be justified by needing more capacity. The used-code count plateaus near the same six survivors across both benchmarks, which is why enlarging past 16 yields almost no accuracy return under a matched training budget.
(c) Rerank signal. Figure 5(c): same Top- candidates; only the selector changes. Tokens on GSM8K / HumanEval are MLP /, random /, GNN /, with accuracy within points—the ranking buys cost, not large Pass@1. The GNN is worse than random because on homogeneous teams Eq. (1) is constant in and its head favors expensive graphs ( with tokens). Only the MLP, reading and regressing measured , separates candidates on the true objective. Keeping the short list fixed isolates the critic: any token gap here is attributable to Eq. (11) alone.
(d) Cost weight . Figure 5(d): is costliest (/ tokens) with no accuracy gain. Tokens fall to (GSM8K) and (HumanEval at ); accuracy peaks at /. Default sits on the frontier: the proxy must both see structure and be asked to care about cost. The prior already prefers cheap codes (top-1: on GSM8K), and rerank adds – accuracy points. Pushing more negative continues to cut tokens but begins to trade accuracy, so is a deliberate operating point rather than an unconstrained minimum-cost choice.
Qwen-3-8B transfer (Table 2). Same teams and 300-record protocol, backbone swapped to Qwen-3-8B. Absolute numbers drop (CoT avg. ), but Codebook Agent still leads every column (//, avg. ) over GTD () and G-Designer (). MATH ( vs. GTD) and MMLU (; three-expert team) show the gain is not confined to saturated arithmetic or four MathSolvers. Design never calls the agent LLM (Eq. (11)), so swapping the backbone changes only the rewards in —the short-list recipe is not an API artifact. Relative ordering among adaptive designers is preserved under the weaker model, which suggests that the amortized index, not gpt-4o-mini idiosyncrasies, drives the main-table gaps.
| Method | GSM8K | MATH | MMLU | Avg. |
|---|---|---|---|---|
| CoT | 87.8 | 55.0 | 59.6 | 67.5 |
| GPTSwarm | 90.93.1 | 60.05.0 | 63.53.9 | 71.54.0 |
| MaAS | 91.23.4 | 60.45.4 | 63.94.3 | 71.84.3 |
| G-Designer | 91.53.7 | 60.25.2 | 64.65.0 | 72.14.6 |
| GTD | 92.34.5 | 61.86.8 | 63.94.3 | 72.75.2 |
| Ours | 92.74.9 | 63.58.5 | 65.86.2 | 74.06.5 |
6 Conclusion
We find that effective multi-agent topologies form a short list, that denser graphs need not cost fewer tokens, and that message-passing critics can miss structure on homogeneous teams. Accordingly, Codebook Agent indexes a small discrete set of topologies and selects among them with a lightweight proxy (Eqs. (4)–(11)), leading all six benchmarks and the Qwen-3-8B transfer while cutting design latency to 2.4 ms and LLM tokens by 21.9–33.2%.
References
- Program synthesis with large language models. arXiv preprint arXiv:2108.07732. Cited by: §5.1.
- Why do multi-agent LLM systems fail?. In Advances in Neural Information Processing Systems 38, Datasets and Benchmarks Track, Cited by: §2.
- Are more LLM calls all you need? towards the scaling properties of compound AI systems. In Advances in Neural Information Processing Systems 37, Cited by: §2.
- FrugalGPT: how to use large language models while reducing cost and improving performance. Transactions on Machine Learning Research. Cited by: §2.
- Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374. Cited by: §5.1.
- AgentVerse: facilitating multi-agent collaboration and exploring emergent behaviors. In Proceedings of the 12th International Conference on Learning Representations, Cited by: §2.
- Optima: optimizing effectiveness and efficiency for LLM-based multi-agent system. In Findings of the Association for Computational Linguistics: ACL 2025, Cited by: §2.
- Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168. Cited by: §5.1.
- Improving factuality and reasoning in language models through multiagent debate. In Proceedings of the 41st International Conference on Machine Learning, Cited by: §1, §2, §5.1.
- Measuring massive multitask language understanding. In Proceedings of the 9th International Conference on Learning Representations, Cited by: §5.1.
- Measuring mathematical problem solving with the MATH dataset. In Advances in Neural Information Processing Systems 34, Datasets and Benchmarks Track, Cited by: §5.1.
- Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531. Cited by: §2.
- Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems 33, Cited by: §2.
- MetaGPT: meta programming for a multi-agent collaborative framework. In Proceedings of the 12th International Conference on Learning Representations, Cited by: §1, §2.
- Automated design of agentic systems. In Proceedings of the 13th International Conference on Learning Representations, Cited by: §2, §5.1.
- LLM-blender: ensembling large language models with pairwise ranking and generative fusion. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Cited by: §5.1.
- Dynamic generation of multi-LLM agents communication topologies with graph diffusion models. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Cited by: §1, §1, §2, §3, §5.1.
- Score-based generative modeling of graphs via the system of stochastic differential equations. In Proceedings of the 39th International Conference on Machine Learning, Cited by: §2.
- More agents is all you need. Transactions on Machine Learning Research. Cited by: §2.
- Assemble your crew: automatic multi-agent communication topology design via autoregressive graph generation. In Proceedings of the 40th AAAI Conference on Artificial Intelligence, Cited by: §1, §3, §5.1.
- Dynamic LLM-agent network: an LLM-agent collaboration framework with agent team optimization. In Proceedings of the 1st Conference on Language Modeling, Cited by: §2, §5.1.
- Are NLP models really able to solve simple math word problems?. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Cited by: §5.1.
- ChatDev: communicative agents for software development. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Cited by: §1, §2.
- Scaling large language model-based multi-agent collaboration. In Proceedings of the 13th International Conference on Learning Representations, Cited by: §1.
- Generating diverse high-fidelity images with VQ-VAE-2. In Advances in Neural Information Processing Systems 32, Cited by: §2, §4.1.
- Solving general arithmetic word problems. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Cited by: §5.1.
- GraphVAE: towards generation of small graphs using variational autoencoders. In Proceedings of the 27th International Conference on Artificial Neural Networks, Cited by: §2.
- Should we be going MAD? a look at multi-agent debate strategies for LLMs. In Proceedings of the 41st International Conference on Machine Learning, Cited by: §2.
- Denoising diffusion implicit models. In Proceedings of the 9th International Conference on Learning Representations, Cited by: §2.
- TopoDIM: one-shot topology generation of diverse interaction modes for multi-agent systems. In Findings of the Association for Computational Linguistics: ACL 2026, Cited by: §5.1.
- Learning MLPs on graphs: a unified view of effectiveness, robustness, and efficiency. In Proceedings of the 11th International Conference on Learning Representations, Cited by: §2.
- Neural discrete representation learning. In Advances in Neural Information Processing Systems 30, Cited by: §1, §2, §4.1.
- Graph attention networks. In Proceedings of the 6th International Conference on Learning Representations, Cited by: §3.
- DiGress: discrete denoising diffusion for graph generation. In Proceedings of the 11th International Conference on Learning Representations, Cited by: §2.
- Self-consistency improves chain of thought reasoning in language models. In Proceedings of the 11th International Conference on Learning Representations, Cited by: §5.1.
- AgentDropout: dynamic agent elimination for token-efficient and high-performance LLM-based multi-agent collaboration. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Cited by: §2, §5.1.
- Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems 35, Cited by: §5.1.
- Extracting low-/high-frequency knowledge from graph neural networks and injecting it into MLPs: an effective GNN-to-MLP distillation framework. In Proceedings of the 37th AAAI Conference on Artificial Intelligence, Cited by: §2.
- AutoGen: enabling next-gen LLM applications via multi-agent conversation. In Proceedings of the 1st Conference on Language Modeling, Cited by: §1, §2.
- VQGraph: rethinking graph representation space for bridging GNNs and MLPs. In Proceedings of the 12th International Conference on Learning Representations, Cited by: §2, §4.3.
- GraphRNN: generating realistic graphs with deep auto-regressive models. In Proceedings of the 35th International Conference on Machine Learning, Cited by: §2.
- Multi-agent architecture search via agentic supernet. In Proceedings of the 42nd International Conference on Machine Learning, Cited by: §5.1.
- Cut the crap: an economical communication pipeline for LLM-based multi-agent systems. In Proceedings of the 13th International Conference on Learning Representations, Cited by: §1, §1, §2, §3.
- G-designer: architecting multi-agent communication topologies via graph neural networks. In Proceedings of the 42nd International Conference on Machine Learning, Cited by: §1, §1, §2, §3, §3, §5.1.
- AFlow: automating agentic workflow generation. In Proceedings of the 13th International Conference on Learning Representations, Cited by: §2, §5.1.
- Graph-less neural networks: teaching old MLPs new tricks via distillation. In Proceedings of the 10th International Conference on Learning Representations, Cited by: §2.
- Unifying generation and prediction on graphs with latent graph diffusion. In Advances in Neural Information Processing Systems 37, Cited by: §2.
- GPTSwarm: language agents as optimizable graphs. In Proceedings of the 41st International Conference on Machine Learning, Cited by: §1, §1, §2, §3, §5.1.