Skip to main content
archive
Search Submit Donate Log in
Press Enter to search · Advanced search

Multiagent Systems

  • New submissions
  • Cross-lists
  • Replacements

See recent articles

Showing new listings for Thursday, 1 October 2026

Total of 34 entries
Showing up to 2000 entries per page: fewer | more | all

New submissions (showing 9 of 9 entries)

[1] arXiv:2609.38327 [pdf, html, other]
Title: Absorbing State Phase Transitions in Multi-Agent Search
Wenwen Zheng, Yuzhe Yang, Helen Qu, Xin Eric Wang, Haewon Jeong
Subjects: Multiagent Systems (cs.MA)

Nontrivial dynamics can emerge in large language model (LLM)-based multi-agent systems, and preliminary evidence exists that formalisms from statistical mechanics can be effective at modeling and predicting such behaviors. In parallel, designing multi-agent communication topology for optimal task-solving is an active research question. In this paper, we focus on predicting the success of multi-agent search tasks using the formalism of absorbing state phase transitions. We first taxonomize search tasks into four types, informed by classical results in combinatorial search. We then theoretically derive a critical communication degree $d_c$, the minimum number of agents each agent can communicate with, above which incorrect hypotheses do not proliferate uncontrollably and the search enters the solved state. Finally, we evaluate frontier LLM-based multi-agent systems on real-world search and discovery tasks, software configuration debugging and physical mechanism discovery, and find that agreement with theory is mixed. LLM agents may not communicate with their neighbors and can develop strategies that are individually beneficial but limits the benefits of collaboration.

[2] arXiv:2609.38482 [pdf, html, other]
Title: PANDA: A Decentralized Architecture with Flexible Orchestration for Scalable, Fault-Tolerant Multi-Agent Systems
Matthew D. Laws, Cristina Nita-Rotaru
Comments: 19 pages, 7 figures, 5 tables
Subjects: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)

Existing architectures for LLM-based multi-agent systems (MAS) cannot reliably and efficiently solve multi-step tasks at scale: they struggle to support large numbers of agents and concurrent tasks, tolerate failures, govern agent interactions, and accommodate the diverse planning and execution patterns different tasks require. We present PANDA, a decentralized architecture that connects a large collective of heterogeneous, independently administered agents, letting them discover each other's capabilities and self-organize into small specialized teams per task. PANDA scales by decoupling collective communication from team communication, allowing agents to participate in multiple teams simultaneously, load-balancing tasks across the collective, and scheduling concurrent work within each agent. PANDA further separates the underlying architecture from the orchestration strategy, supporting three planning and execution patterns (star, chain, and mesh) that can be selected according to the structure and requirements of each task. PANDA detects infrastructure and orchestration failures and recovers affected tasks by dynamically replanning around failed components. Finally, to provide governance without a centralized service that would limit scalability, PANDA uses a web-of-trust model to constrain agent interactions to established trust relationships. We evaluate PANDA on the HotPotQA benchmark, demonstrating that it scales to thousands of agents, assembles teams in milliseconds, matches state-of-the-art accuracy at up to 8x the efficiency, and sustains 100% task completion under faults where existing systems fail.

[3] arXiv:2609.38516 [pdf, html, other]
Title: From Solo to Social Learning: Characterizing Recursive Social Improvement in LLMs
Kunal Jha, Max Kleiman-Weiner, Natasha Jaques
Subjects: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

Large language models (LLMs) can now improve themselves by revising the instructions they follow, and LLM agents are increasingly orchestrated to work together on complex problems. However, self-improvement methods typically optimize one system at a time, and multi-agent frameworks often have every model work toward a shared goal. We ask a different question. When each agent pursues its own reward, can self-improving LLMs learn from one another well enough to improve the whole population? We call this capability recursive social improvement. We study populations that revise skill files and choose whether, when, and whom to copy from. Independent search, learning from peers, and acting all share one token budget. In controlled environments, established social-learning algorithms benefit from peers, but three LLMs do not. They earn less reward per token than solo learners, and explore too narrowly or run out of tokens before acting. We then let the models write and revise their own skills. Observing peers changes how they improve, helping one model find useful skills sooner and another spend less on private search. Neither, however, outperforms independent learners at the same cost. Skills are copied, revised, and passed on, so one discovery can seed further search. Yet these exchanges concentrate the population around fewer independent discoveries. Together, these results show that LLMs can make learning more efficient by copying from peers, but not yet more effective.

[4] arXiv:2609.38643 [pdf, html, other]
Title: Recursive Organization Improvement: A Modeling Specification for Human--Agent Organizations
Zilong Wang
Comments: 18 pages, 3 figures, 9 tables;
Subjects: Multiagent Systems (cs.MA); Human-Computer Interaction (cs.HC)

Stronger AI agents do not automatically produce better organizations: teams must also learn which work arrangements to retain and when to reconsider them. We propose a modeling specification for recursive organization improvement and evaluate it through an executable checker, a public-record mapping, and controlled simulation. The specification connects actor-visible histories, organizational memory, decision rights, and evidence-carrying change contracts. The mechanism study crosses six decision rules, three memory conditions, and three task environments under fixed resource ceilings. In a stationary environment, cumulative evidence raises balanced evaluation's normalized net value per task from 0.45224 to 0.48007. Repeated reassessment's disadvantage relative to this comparator falls from 0.01702 with reset evidence to 0.00007 with cumulative evidence. A reversal of the best workflow reveals the opposite cost: indefinite retention delays adaptation, while a finite window restores eventual performance at a transition cost. In exploratory controls, matching trial acquisition and label reuse reduces the apparent reassessment gain from 0.00607 to 0.00191. Program replacement adds no stable benefit across the tested reversal times. The study identifies evidence acquisition, reuse, and timely updating as mechanisms that must be separated from evaluator replacement when assessing organizational improvement.

[5] arXiv:2609.38662 [pdf, html, other]
Title: CollabFlow: Recursive Self-Improvement of Agent Collaboration
Xiao Huang, Mingda Zhang, Junming Zhang, Qiang Huang, Hanwen Zhang, Yue Dai, Zijia Wang, Xiaoying Tang
Subjects: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)

Recursive self-improvement (RSI) lets a system improve from its own outcomes; in LLM-based multi-agent systems, Agents refine one another within a task, and outcomes improve how they collaborate across tasks. However, existing multi-agent collaboration leaves this loop open: collaboration is pre-defined at the operator level, topology-only learning keeps verbatim exchange that propagates errors, and reward maximization on a system's own outcomes concentrates on a few teams. To address these challenges, we propose CollabFlow, an RSI system of Learned Agent Collaboration: a trainable Collab-Director constructs teams of complete Agents, a frozen executor runs them, and each round's outcomes retrain the director. Within each round, the edges of a collaboration graph carry protocols of Evidence-Conditioned Communication: a receiver adopts a differing answer only when the sender's evidence is stronger by a margin, so the director learns who communicates and how. Across rounds, we further propose Collaborative Trajectory Balance (CTB), a flow-based objective that credits each team once across its construction orders and targets a reward-proportional distribution over teams, so several good teams stay in play. We also bound how far this self-generated target moves between rounds, which shrinks as records accumulate. On twelve datasets, CollabFlow outperforms all baselines and keeps improving across rounds. Code is available at this https URL.

[6] arXiv:2609.38761 [pdf, html, other]
Title: Where Do Multi-Agent Systems Fail? Evidence-Grounded Diagnosis of Collective Mechanisms
Zhengye Han
Subjects: Multiagent Systems (cs.MA)

When a multi-agent system answers correctly, it is tempting to conclude that its agents shared, checked, and used information as intended. Yet a system can break one of its collective mechanisms, the rules that govern how agents route, admit, store, and act on shared information, and still return the right answer, while a wrong answer rarely reveals which mechanism failed. We ask what evidence from an execution is sufficient to conclude that a particular mechanism was violated. Our answer is a diagnostic contract, which separates what counts as a violation from which execution records can establish one, and concludes that a violation is supported, ruled out, or unknown; removing records can make this conclusion unknown but never reverse it. We test contracts for four mechanisms by replaying executions from the step where a mechanism acts, once unchanged, once with the mechanism broken, and once with it restored. Broken mechanisms often left the answer correct. An LLM diagnoser detected many more violations from internal records than from public outputs, yet with identical records a generic prompt often claimed certainty the records did not support, which prompts stating the contracts largely avoided. The contracts also applied, in narrow form, to mechanisms in independently developed systems, but a diagnostic behavior that was nearly perfect on our benchmark degraded on an independently developed workflow. A correct outcome is therefore no substitute for records of how collective mechanisms operated, and agreement on one benchmark does not show that a diagnoser transfers to another system.

[7] arXiv:2609.38960 [pdf, html, other]
Title: Fast and Scalable Multi-Agent Distribution Matching via Partitioned Optimal Transport
Kooktae Lee, Ruchika Singh
Subjects: Multiagent Systems (cs.MA); Robotics (cs.RO); Systems and Control (eess.SY); Optimization and Control (math.OC)

This paper presents a scalable optimal-transport-based framework for terminal distribution matching in multi-agent systems. While optimal transport provides a natural way to measure distributional mismatch and assign agents to a desired spatial distribution, global discrete transport can become computationally expensive for large-scale systems. We address this bottleneck by partitioning agents and target samples into spatially corresponding blocks and solving smaller local transport problems. Under a mass-balance condition, the resulting restricted coupling remains feasible for the global problem and provides an upper bound on the Wasserstein cost. The local assignments generate target locations for finite-horizon agent control, applicable to both linear and nonlinear dynamics. By alternating local assignment and control, we establish a cycle-to-cycle descent guarantee for the resulting transport surrogate. The proposed framework therefore enables scalable terminal distribution matching while retaining a rigorous connection to the Wasserstein objective. The technical soundness of the proposed results is validated through simulations.

[8] arXiv:2609.39211 [pdf, html, other]
Title: Consensus and Factual Dynamics in Large Populations of Interacting Language Models
Emanuele Ricco, Elia Onofri, Vincenzo Sammartino, Roberto Di Pietro
Subjects: Multiagent Systems (cs.MA)

Large Language Model (LLM) agents are increasingly deployed as populations of interacting entities, in which consensus --agreement on a shared answer-- emerges as a collective, unengineered behaviour. Prior work on LLM consensus shows that agents can cross-verify their answers and converge towards more factual responses, treating agreement as a proxy for correctness. However, these studies usually fix a single interaction structure, leaving open how consensus depends on how agents interact. We address this gap by introducing RHEON, a physics-inspired framework that recasts a population drawn from a single frozen model as an evolving $O(n)$ spin system on a ladder of interaction geometries of increasing effective dimension --from a 1D ring to a full-coupling mean-field graph-- with the sampling temperature $T$ as the tunable source of thermal disorder, evolved through a Glauber-like asynchronous dynamics. Sweeping RHEON across $432$ configurations of prompt, population size, communication topology, and sampling temperature yields Eraclitus-4.7M, a tagged evolutionary corpus of $4.7$ million responses. We find that agents reach their strongest consensus gain within the first few update sweeps and that increasing the number of neighbours per agent accelerates convergence on average. We further show that whether a configuration settles on factually correct or hallucinated consensus is not predictable from its initial state alone, and that the hallucination-minimising temperature depends on how the agents are coupled, so the common near-greedy default is not automatically the safest. Finally, semantic agreement correlates positively with factual convergence, and interaction strengthens the association, yet never enough for unanimity to certify correctness.

[9] arXiv:2609.39889 [pdf, html, other]
Title: Solving Multi-Agent Sokoban via LaCAM
Keisuke Okumura
Subjects: Multiagent Systems (cs.MA)

Sokoban, a puzzle game in which an agent pushes boxes onto unlabelled target locations in a grid world, is a long-standing benchmark planning problem. While it is easy to see the connection to practical applications such as warehouse logistics with autonomous forklifts, its multi-agent counterpart has remained underdeveloped. This is because Multi-Agent Sokoban is substantially more difficult due to factors specific to multi-agent planning, such as the rapidly growing branching factor as the number of agents grows and the need to handle integrated task assignment and collision-free pathfinding. In this paper, we show that a scalable planner for Multi-Agent Sokoban can be designed by leveraging recent advances in multi-agent pathfinding (MAPF). Specifically, our Sokoban-LaCAM efficiently solves instances involving tens of agents and boxes while preserving both completeness and eventual optimality guarantees. This provides evidence that MAPF can serve as a powerful primitive for solving broader collective automation problems.

Cross submissions (showing 18 of 18 entries)

[10] arXiv:2609.38270 (cross-list from cs.CR) [pdf, html, other]
Title: VirusCascade: Hijacking Collaborative Reflection in LLM-Powered Recommender Agents
Yurong Hao, Wen Zhou, Guowei Guan, Tiantong Wu, Fuyao Zhang, Wei Yang Bryan Lim
Comments: Accepted by NDSS 2027
Subjects: Cryptography and Security (cs.CR); Machine Learning (cs.LG); Multiagent Systems (cs.MA)

Advancing beyond traditional static scoring models, LLM-powered agentic recommender systems (LLM-ARS) instantiate users and items as autonomous agents, whose semantic states are dynamically refined through a recurrent process known as collaborative reflection. While this mechanism improves recommendation quality, it simultaneously introduces a systemic vulnerability: adversarial evidence injected into a single agent can be rationalised into a legitimate preference narrative, written back into memory, and propagated to other agents through interaction contexts. We term the local rationalisation process reflection laundering, and its system-wide escalation through collaborative reflection collaborative-reflection hijacking. Existing attacks on recommender systems, whether based on interaction-level data poisoning or text-level adversarial perturbations, assume static pipelines and thus cannot exploit this recurrent, multi-agent amplification pathway. To bridge this gap, we first conduct a controlled vulnerability analysis that establishes two exploitable properties underlying collaborative-reflection hijacking: reflective persistence and cross-agent propagation. Then building on these findings, we propose VirusCascade, the first black-box targeted promotion attack that jointly shapes semantic and structural attack surfaces: the former ensures the target item is naturally rationalised as satisfying broad user preferences, the latter positions it for system-wide propagation. Extensive experiments on four real-world datasets across diverse LLM-ARS architectures demonstrate that VirusCascade consistently achieves state-of-the-art targeted exposure under evaluated stealth constraints, reaching a mean E@20 of 0.384 and surpassing the strongest baseline by an absolute margin of +0.185.

[11] arXiv:2609.38274 (cross-list from cs.CL) [pdf, html, other]
Title: Which Models Work Well Together? Measuring Heterogeneity for LLM Team Selection
Liangyu Teng, Hengsong Liu, Juncen Guo, Jingyu Zhang, Yang Liu, Jing Liu, Liang Song
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG); Multiagent Systems (cs.MA)

The performance ceiling of an LLM team is constrained not only by individual model capabilities, but also by inter-member error resonance and predictive differences. Although heterogeneous teaming is often observed to be effective in practice, existing approaches lack complementarity metrics that are computable, interpretable, and optimizable, leaving team composition to rely on heuristics. We propose a heterogeneity-driven team selection framework that performs offline profiling to characterize individual capability along with two complementary signals: one captures decorrelation in error patterns to reduce co-failures, while the other measures divergence in predictive behavior to capture strategy diversity. We formulate team selection as a standardized quality--complementarity combinatorial objective and apply an efficient greedy search to select a small team from a candidate pool. Experiments across multiple benchmarks demonstrate that our framework consistently outperforms quality-only baselines under controlled candidate pools and team sizes, establishing reusable selection principles for multi-LLM systems.

[12] arXiv:2609.38324 (cross-list from physics.soc-ph) [pdf, html, other]
Title: Multi-agent discussion gains less when dissent is withheld
Chand Sahil Mansuri, Xin Wang, Mengying Li, Bryan Acton, Rory Eckardt, Dhaval Patel, Sadamori Kojaku
Subjects: Physics and Society (physics.soc-ph); Computation and Language (cs.CL); Multiagent Systems (cs.MA)

Multi-agent systems of LLMs add discussion to majority voting and are therefore expected to be more capable. However, empirical reports conflict on whether discussion improves accuracy or leads to an incorrect consensus. Here, we introduce a parsimonious model that explains when discussion improves accuracy and when it ends in an incorrect consensus, built from four behaviors repeatedly observed in LLM agents: (1) withholding dissent, (2) internalizing a stated answer, (3) reconsidering after seeing dissent, and (4) correcting toward the correct answer. The model shows that discussion can overturn an incorrect initial majority only when the withholding rate $c$ is below a critical rate $c^* = \gamma/(\gamma + a)$, set by the net correction rate $\gamma$ and the internalization rate $a$. We estimate these rates from conversation logs with a Bayesian method and place LLM teams relative to $c^*$. As the model predicts, the gain from discussion shrinks as withholding rises, across LLMs and on a hidden profile benchmark, HiddenBench, and MedEInst. Instructing agents not to withhold dissent increases this gain. Turning reasoning off also increases the gain, because reasoning raises the internalization rate $a$ and keeps agents from reconsidering a minority answer. These findings reconcile the conflicting reports and identify when discussion outperforms majority voting.

[13] arXiv:2609.38660 (cross-list from cs.CL) [pdf, html, other]
Title: Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation
Haibo Jin, Xinjie Li, Najmeh Sadoughi, Yang Liu, Yibo Wang, Zhu Liu, Yuzong Liu
Comments: 49 pages
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multiagent Systems (cs.MA)

Long-form subtitle translation requires reasoning over discourse and cultural context spanning episodes or entire series, while maintaining consistent terminology and style. Existing single-LLM methods are largely sentence-level, and multi-agent systems often use static workflows that do not adapt to scene complexity or production context. We propose SMART, a Self-evolving Multi-Agent system for long-foRm subtitle Translation. During test-time training, SMART builds persistent series-level memory and translates a subset of sentences through a dynamic router and Mixture-of-Agents layer with tools for terminology verification, subtitle constraint validation, and contextual retrieval. A judge-refiner loop scores candidates and uses textual critiques to update agent prompts and routing policies without retraining the underlying LLMs. During test-time inference, the evolved configuration translates the remaining series. We also introduce Subtitle Arena, covering 14 genres, 2--198 episodes per series, production years 1959--2023, and 15 target locales, together with SubMQM, a subtitle-adapted MQM framework with seven dimensions and 19 error categories. SMART achieves the best overall MQM score in all 15 Subtitle Arena directions, reducing average penalty by 6.9% over the strongest competing agent system. On the public MuSC benchmark, SMART obtains the best model result across all four language pairs and also achieves the best human-evaluation result, with an overall score of 4.50/5.

[14] arXiv:2609.38731 (cross-list from math.OC) [pdf, html, other]
Title: Decentralized Decision-Making among Heterogeneous Autonomous Vehicles: An $α$-Potential Game Framework
Anran Hu, Zhexin Wang, Yufei Zhang, Xuan Di
Subjects: Optimization and Control (math.OC); Computer Science and Game Theory (cs.GT); Multiagent Systems (cs.MA)

We study noncooperative multi-vehicle games among heterogeneous autonomous vehicles, where each vehicle adopts a decentralized closed-loop policy based on its own state, and optimizes an objective that depends on other vehicles through potentially asymmetric interaction weights. We develop an $\alpha$-potential game framework that reduces the computation of an approximate Nash equilibrium (NE) to the minimization of a single auxiliary $\alpha$-potential function. We explicitly construct this $\alpha$-potential, establish the existence of its minimizers, and characterize the equilibrium approximation error $\alpha$ in terms of interaction asymmetry. We further introduce vehicle-specific scaling to reduce the effective interaction asymmetry, thereby tightening the equilibrium approximation and, in important cases, recovering an exact NE despite asymmetric interactions. We also derive social-efficiency guarantees for the potential-selected policies, revealing how the interaction structure shapes worst-case efficiency. Numerical experiments demonstrate the flexibility of the framework in capturing heterogeneous vehicle interactions, collision and obstacle avoidance, lane changing and overtaking under different traffic configurations, and priority-based intersection crossing.

[15] arXiv:2609.38760 (cross-list from cs.RO) [pdf, html, other]
Title: HALO: Heterogeneous Allocation Via Localized Observations for the Vehicle Routing Problem
Andrew Meighan, Hyungsub Kim, Or Dantsker
Comments: 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
Subjects: Robotics (cs.RO); Multiagent Systems (cs.MA)

Scalable robotic fleets have become increasingly popular for various applications such as package delivery, warehouse management, and military operations. Prior fleet control algorithms solve centralized routing problems with up to $1{,}000$ tasks in controlled environments, yet they fail to consider realistic constraints such as limited observation and communication ranges typical of decentralized fleets. Thus, deploying existing fleet control algorithms into real-world settings is currently infeasible.
To tackle this, we propose Heterogeneous Allocation via Localized Observations (HALO) to solve the Vehicle Routing Problem (VRP). HALO is a hybrid method that splits the VRP into allocation and routing portions to provide onboard, real-time solutions to robots in dynamic environments. During the allocation phase, HALO utilizes a heterogeneous graph neural network framework with unique message passing layers to explicitly separate the learning of spatial distributions and task-to-robot compatibility.
Evaluation results on a partially observable, online variant of the VRP show HALO significantly outperforms the heuristic baseline while maintaining similar solution quality to an all-knowing offline variant of HALO. While HALO is explicitly designed for partially observable environments, it imposes no strict upper bound on the observation space allowing us to test HALO on the traditional static, single-depot VRP. Here, HALO outperforms state-of-the-art architectures strictly optimized for the static variant of the VRP by up to $14.06\%$. Throughout all testing, this framework maintains the quickest execution times which emphasizes its potential for large-scale, real-time deployment.

[16] arXiv:2609.38822 (cross-list from cs.IR) [pdf, html, other]
Title: SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale
Guanqun Yang, Wenlong Zhang, Tian Shi, Ping Wang
Comments: Accepted at AACL-IJCNLP 2026. Code at this https URL
Subjects: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)

Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into this http URL directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated retrieval loop that rewrites queries and refines candidates inside the agent's decision loop, paying LLM tokens on every task. We present SkillSeek, an open-source two-stage skill retriever built from the standard IR recipe (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP). Across a $4 \times 11$ grid of pool, backbone, and method on the 89-task SkillsBench benchmark, SkillSeek reaches observed parity with the LLM-mediated loop of Liu et al. at essentially no extra cost: plain bm25 alone records a pass rate at or above their refined loop on three of four settings, and a small cross-encoder covers the remaining difference on the fourth. A first-stage recall ceiling explains the pattern, and total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the no-skill baseline). Under the SkillsBench tasks and OpenHands harness we tested, this positions the standard IR recipe as a strong default for agent-skill retrieval, with LLM-mediated alternatives a natural fit for cases where deterministic methods fall short.

[17] arXiv:2609.38914 (cross-list from cs.AI) [pdf, html, other]
Title: Risk-Aware Adaptive Evaluation: Finding High-Impact Failures Under Limited Budgets
Priyanath Maji, Spandan Ghose Chowdhury
Comments: Accepted to 40th Conference on Neural Information Processing Systems (NeurIPS 2026). Workshop: Evaluation of Interactive Agents
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)

Evaluating interactive agents is expensive. Agent behavior is stochastic, so reliability must be measured over repeated trials, but failures are rare and differ widely in how much they matter. Standard benchmarks spend this budget uniformly: a read-only lookup is sampled as often as an irreversible payment action. We instead formulate evaluation as a sequential allocation problem. Given a fixed trial budget and a set of scenarios whose failure behavior is unknown, which scenarios should be run, and run again? We propose a risk-aware contextual Thompson Sampling policy that combines a pre-execution scenario context vector and a fixed impact score with the failure outcomes observed during evaluation, and we test it by offline replay over 70 $\tau$-bench airline scenarios and 824 recorded trials. Our main result is at the smallest budget: with only 50 trials ($6\%$ of the corpus), the policy recovers $86\%$ of the impact-weighted failures an oracle could find, compared to $25\%$ for uniform allocation. It discovers $3.5\times$ more impact-weighted failures (215.4 vs. 62.2) with the same number of trials, delivers $5\times$ the discovery per dollar, and cuts the budget wasted on scenarios that never fail from $34\%$ to $2.8\%$. The rest of our analysis demonstrates and qualifies this result: a budget sweep shows the advantage shrinks as the budget approaches the corpus size, and paired significance tests show that scenario context helps mainly at small budgets while posterior-based exploration helps at moderate ones. Risk-aware adaptive allocation therefore helps most exactly where evaluation budget is scarcest.

[18] arXiv:2609.39045 (cross-list from cs.CL) [pdf, html, other]
Title: RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement
Wenyi Wu, Minghao Fu, Jieyu You, Kun Zhou, Siqi Liu, Aayush Salvi, Yiheng Lin, Ce Zhang, Xiaohan Lan, Jiahui Zhu, Yujie Zhong, Qi She, Biwei Huang
Subjects: Computation and Language (cs.CL); Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG); Multiagent Systems (cs.MA)

Recent advances in large language models have made automatic game generation increasingly feasible, yet reliably improving generated games beyond a playable version remains challenging. Naive iterative refinement can easily overfit a small set of test cases, producing fragile games with unresolved bugs, missing behaviors, and poor generalization to broader player interactions. We introduce RSIGame, an autonomous agentic game development framework with recursive self-improvement. RSIGame organizes development into complementary local and global loops. Concretely, a local explore-diagnose-improve loop broadly explores the executable game, diagnoses and prioritizes discovered issues, and performs evidence-grounded revision, where an evolving checklist continually accumulates new testing and improvement guidance. A global loop tracks overall quality, preserves the best checkpoint, and detects saturation or regression over long-horizon development. Beyond test-time improvement, RSIGame further internalizes successful development experience into the generator through training. Across 140 GameCraft-Bench tasks, two game engines, and five generators, RSIGame consistently improves game quality under matched development budgets. Notably, experience internalization enables Qwen3.8-27B to reach 61.38 on Godot and 58.53 on Phaser, exceeding GPT-5.5 one-shot scores while reducing Qwen's generation tokens by 11 times.

[19] arXiv:2609.39143 (cross-list from cs.AI) [pdf, html, other]
Title: RefCon: Iterative Refinement and Contrastive Memory Extraction for Context-Evolving Agent
Ubaidillah Ariq Prathama, Bo Liu, Yeo Boon Hong, Yu-Xuan Huang, Yangkai Ding, Tao Yu
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)

Long-horizon agent interactions generate useful but noisy experience, and retraining models to absorb it is expensive. Context-evolving agents therefore need memory extraction methods that improve with more test-time compute without relying on gold labels. We propose RefCon, which combines sequential self-refinement with parallel self-contrast to extract higher-quality memories without gold labels. Evaluated on AppWorld and BFCL-V3 across multiple context-evolving agent frameworks, RefCon delivers strong and consistent gains, including relative improvements of 21.6% on ACE and 16.6% on ReMe over no-scaling baselines, while a diversity-focused variant (DivCon) achieves a 35.5% gain on ReasoningBank. RefCon consistently outperforms existing baselines without ground-truth labels, and generalizes across model scales and to software engineering tasks, where it surpasses even ground-truth baselines. We further analyze the accuracy-token trade-off and scaling behavior, showing RefCon maintains favorable efficiency and continues to improve as more trajectories are used, unlike diversity-only scaling which saturates earlier.

[20] arXiv:2609.39302 (cross-list from quant-ph) [pdf, html, other]
Title: SQD-Agent: LLM-driven agentic framework for Quantum Chemistry workflows
Kislaya Tiwari, Anupama Ray
Comments: Accepted in IEEE International Conference on High Performance Computing (HiPC 2026)
Subjects: Quantum Physics (quant-ph); Multiagent Systems (cs.MA)

Quantum algorithms and quantum hardware are advancing towards a promising paradigm for scientific applications. However, translating domain-specific problems into executable hybrid quantum-classical workflows remains a significant barrier for application researchers due to the required expertise in quantum algorithms, nuances in quantum programming, and hardware-aware system integration. At the same time, AI and primarily LLM based agents are increasingly capable of interpreting natural-language intent, reasoning over complex workflows, and translating high-level objectives into executable code and building computational pipelines. In this work, we introduce SQD Agent, an LLM-based agentic framework that translates natural-language user intent into executable workflows for Quantum Chemistry applications where algorithms from the Sample-Based Quantum Diagonalization (SQD) family are used. By automating this translation, SQD Agent reduces the level of human expertise and configuration overhead required, thereby simplifying experimentation in hybrid quantum-classical settings for application researchers new to quantum. SQD Agent adopts a modular and extensible architecture that supports seamless integration of heterogeneous quantum backends, classical solvers, and workflow components, ensuring adaptability to rapidly evolving quantum ecosystems. The framework further incorporates interactive capabilities for on-demand profiling, bottleneck analysis, resource optimization, intelligent result caching, and convergence visualization. Key features include quantum chemistry experiments, error mitigation on real quantum hardware, together with analysis of candidate mitigation schemes in terms of their potential error-recovery behavior and computational budget, helping users understand their practical trade-offs and decide which strategies to explore in subsequent experiments.

[21] arXiv:2609.39637 (cross-list from cs.RO) [pdf, html, other]
Title: Prediction is Better than Detection: Traffic Congestion Control using Drones
Samira Hayat, Christian Raffelsberger
Comments: 7 pages, 9 figures
Subjects: Robotics (cs.RO); Multiagent Systems (cs.MA)

A central question in deploying teams of mobile robots for persistent monitoring is how task performance scales with fleet size, and whether this scaling holds once sensing drives downstream action rather than mere observation. We study this question for a team of drones performing traffic-jam detection and prediction in a simulated road network, whose reports drive an adaptive traffic-signal controller in closed loop. We build a multi-agent simulation, with vehicles following Nagel-Schreckenberg cellular-automaton dynamics and drones patrolling junctions via a round-robin policy, and sweep fleet size, traffic level, and network size to evaluate detection rate, detection delay, and prediction rate. We show how performance plateaus for fleet size approximating the number of junctions being monitored, and offer a general fleet-provisioning rule for persistent-monitoring deployments. More significantly, adapting the signal on a predicted jam, rather than a detected one, roughly doubles the resulting reduction in jam duration, showing that the value of onboard prediction in a sensing-to-action pipeline can exceed the value of adding more robots. Prediction accuracy, not sensing coverage, is now the binding constraint on further improvement, pointing to onboard inference, not fleet size, as the more promising direction for future work.

[22] arXiv:2609.39727 (cross-list from cs.AI) [pdf, html, other]
Title: OverForge: Reasoning Through Strategies and Tactics Helps Cooperative Lifelong Adaptation
Oana Madalina Fron, Ojas Shirekar, Chirag Raman
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multiagent Systems (cs.MA)

Cooperative language-model agents must coordinate over long horizons and adapt to changing environments and to partners with unfamiliar conventions, yet existing agents map observations to actions without separating persistent coordination strategies from their tactical execution. We introduce OverForge, a training-free hierarchical architecture that separates strategic reasoning over roles and divisions of labour from tactical reasoning over actions within each agent's private, partner-conditioned world model. A metacognitive Prefrontal Cortex Module couples the two levels by forming strategy-action branches, imagining their consequences with a forward model, and committing when confident. In OvercookedV2, OverForge delivers 7 soups in a connected kitchen versus 3 for each flat LLM baseline, retains agreed roles, and adopts roles proposed by unfamiliar partners. Ablations and a fixed-strategy probe show that persistent strategies guide tactical adaptation while each reasoning level contributes to coordination. Memory restarts show that cross-episode partner knowledge supports task performance and partner prediction, linking the hierarchy to continual adaptation.

[23] arXiv:2609.39788 (cross-list from cs.AI) [pdf, html, other]
Title: Safety of Latent Communication in Multi-Agent Systems
Muhammad Huzaifa, Sina Mavali, Thorsten Eisenhofer
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)

Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task performance without requiring harmful target responses. Across three communication topologies and four safety benchmarks, this attack raises the mean harmful-compliance score from 27.9 with benignly trained links to 76.9. Compared with direct supervised optimization, it also achieves higher average accuracy on two benign utility benchmarks. Adapting the rewards toward safer behavior also enables repair of compromised links, substantially reducing harmful compliance across all evaluated attacks without updating the agents. Overall, our results show that safety alignment requires considering the multi-agent system as a whole.

[24] arXiv:2609.39964 (cross-list from cs.AI) [pdf, html, other]
Title: AIMS: An Agentic AI Framework for Sim-to-Real Multi-Modal ISAC
Yijie Bian, Kai Zhang, Wei Guo, Zixin Wang, Shenghui Song, Jun Zhang, Khaled B. Letaief
Subjects: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)

Multi-modal integrated sensing and communication (ISAC) enables environmental perception and reliable connectivity for intelligent wireless networks. Data-driven multi-modal ISAC models depend heavily on annotated real-world data to learn relationships across sensing and wireless observations, thereby constraining scalable deployment. Although synthetic data generation reduces the burden, adapting existing simulation pipelines to a target deployment requires consistent scene, sensing, wireless, and learning configurations, while mismatches among these coupled components impair sim-to-real transferability. To address the challenge, we propose an agentic artificial intelligence (AI) framework for sim-to-real multi-modal ISAC, named AIMS. Given a natural-language deployment request specifying the target task, deployment conditions, and real-data budget, AIMS derives a deployment-specific sim-to-real configuration and coordinates its execution to produce a deployment-specific task model. A two-agent architecture coordinates scene construction with task learning. A scene construction agent generates geographically grounded, synchronized sensing and wireless records from shared physical states, while a scene understanding agent configures task-relevant modalities and mixture-of-experts (MoE) learning for zero-shot inference or few-shot adaptation. Structured domain knowledge guides dependency-aware planning, while validation evidence supports feedback-driven revision of affected decisions. Experiments on the real-world DeepSense~6G dataset demonstrate improved vehicle detection and beam prediction over the considered simulation and fusion baselines. A separate orchestration benchmark evaluates task interpretation, dependency reasoning, and feedback-driven replanning across diverse deployment requests, showing improved plan correctness with structured domain knowledge and validation feedback.

[25] arXiv:2609.39982 (cross-list from cs.CL) [pdf, html, other]
Title: Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
Minki Kang, Ryo Hachiuma, Shaokun Zhang, Subhashree Radhakrishnan, Yonggan Fu, Jindong Jiang, Mingjie Liu, Ehsan Hosseini-Asl, Yi Dong, Yu-Chiang Frank Wang, Byung-Kwan Lee
Comments: Project page: this https URL
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)

Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. To study these questions, we introduce Mid-Harness, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged. With a TMAX-9B generator, more action sampling yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator. On TerminalBench-Lite, a GPT-5.6 Sol verifier raises Pass@1 from 50.00% for the base agent to 68.03% with 8 sampled actions. When the same TMAX-9B model serves as the verifier, pairwise verification performs best among the evaluated verification mechanisms. Distilling responses from the stronger verifier into TMAX-9B further improves Pass@1, while leaving the action generator unchanged. With TMAX-9B on TerminalBench-Lite, combining action and trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone. Mid-Harness also improves performance across additional models, benchmarks, and harnesses. These findings identify action scaling as a promising target for test-time compute scaling in terminal agents.

[26] arXiv:2609.40102 (cross-list from cs.RO) [pdf, html, other]
Title: Passive Stiffness Shaping in Cable-Suspended Aerial Manipulation via Movable Compliant Anchors
Antonio Franchi, Amr Afifi
Subjects: Robotics (cs.RO); Multiagent Systems (cs.MA); Systems and Control (eess.SY)

Cable-suspended aerial manipulation offers a lightweight architecture for cooperative transportation and physical interaction, yet the passive mechanical response perceived at the load remains insufficiently understood and systematically exploited. This work interprets aerial vehicles as movable compliant anchors and develops a gravity-aware quasi-static theory for predicting and shaping the passive Cartesian stiffness of a suspended load. The formulation applies to an arbitrary number of aerial vehicles connected to a point load by taut, straight, inextensible cables. At a selected gravity-loaded equilibrium, aerial-anchor compliance and transverse cable geometric compliance combine in series within each leg, while the leg stiffnesses act in parallel on the load. For isotropic aerial-anchor behavior, each leg is exactly equivalent to a virtual unilateral elastic cable, revealing an axial--transverse stiffness decomposition governed by the equilibrium tension. These results define a nonlinear map from commanded-anchor configuration to passive load stiffness, whose differential enables local constraint-preserving shaping through anchor repositioning. A dynamic rigid-body validation framework with nonlinear vehicle control, elastic-damped tendons, and environmental contact is defined to assess when and to what extent the derived stiffness remains predictive beyond the assumptions of the analytical model.

[27] arXiv:2609.40269 (cross-list from cs.AI) [pdf, html, other]
Title: Belief-Aware Multi-Agent Path Finding under Map Uncertainty
Viraj Parimi, Shao-Hung Chan, Han Zhang, Jingkai Chen, Brian Williams
Comments: Under review
Subjects: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA); Robotics (cs.RO)

Multi-Agent Path Finding (MAPF) aims to find collision-free paths for multiple agents in a shared environment. Classical MAPF assumes that all static obstacles are known in advance, but real-world environments can change unexpectedly due to fallen objects, spills, or other local disturbances. When such changes are spatially correlated, an observation can inform traversability estimates beyond the observed location. Prior approaches address uncertainty in traversability through contingent plans or replanning based on direct observations, but do not leverage this spatial dependence to infer the traversability of nearby unobserved locations. As a result, they cannot use one observation to anticipate nearby unobserved obstacles that may cause costly rerouting later. We focus on Belief-Aware MAPF, where map discrepancies are fixed during execution but initially unknown, and observations can be informative beyond the observed location. We propose Multi-Agent Gaussian belief Inference for Coordination (MAGIC), a framework that updates a shared belief about traversability online based on agents' observations. MAGIC uses a Gaussian Markov Random Field and Gaussian Belief Propagation to approximately infer traversability and construct detour-aware costs for standard MAPF planners. Our experiments on MAPF benchmarks show that MAGIC reduces the executed sum of costs compared to existing approaches on 96.3% of instances, across several planner families and teams of up to 800 agents, demonstrating its applicability to large-scale MAPF problems.

Replacement submissions (showing 7 of 7 entries)

[28] arXiv:2607.25566 (replaced) [pdf, html, other]
Title: ARCHER: Agentic Rule and Compliance Harness for Executable Regulations
Chiraag Singh Anand, Xue Wen Tan, Lionel Teo, Eric Tan
Comments: Accepted at NeurIPS 2026 Workshop on AI for Verifiable Coding
Subjects: Multiagent Systems (cs.MA); Computational Engineering, Finance, and Science (cs.CE); Software Engineering (cs.SE)

Verifying building compliance requires validating thousands of rules against large Building Information Modeling (BIM) designs, which is laborious, capital-intensive, and unscalable. Existing Automated Compliance Checkers (ACCs) are often difficult to generalize across different scenarios, as they are typically developed for highly specific rule sets and use cases. In addition, many ACCs are proprietary, meaning the underlying verification code is not released to end users, so users cannot verify whether their regulatory intent can be accurately captured. We introduce ARCHER (Agentic Rule and Compliance Harness for Executable Regulations), a test-driven, deterministically orchestrated multi-agent program-synthesis harness that generates auditable verification code from regulatory Codes of Practice, enabling transparent, adaptable, and scalable compliance checking. To characterize what makes agentic synthesis work, we evaluate a taxonomy of six harnesses of increasing agentic sophistication across four backbone models, spanning realistic data-governance tiers (from frontier third-party APIs to a fully on-premise open-weights model) on a novel dataset derived from real-world compliance scenarios. ARCHER's deterministic multi-agent orchestration achieves the highest accuracy for every backbone, improving mean union accuracy by 82% over a naive single-pass prompting baseline. Our cost-accuracy analysis further shows that using the ARCHER harness, a self-hosted open-weights model can reach 97.8% of frontier-API accuracy at a quarter of the cost, making data-sovereign compliance checking practical.

[29] arXiv:2507.16221 (replaced) [pdf, html, other]
Title: Unbeatable imitation of a friend
Masahiko Ueda
Comments: 27 pages, 6 figures
Subjects: Physics and Society (physics.soc-ph); Multiagent Systems (cs.MA); Systems and Control (eess.SY)

Imitation sometimes achieves success in multi-agent situations even though it is very simple. In repeated games, success of imitation has been characterized by unbeatability against other agents. Previous studies specified conditions under which imitation is unbeatable against co-players in repeated symmetric games, and clarified that the existence of unbeatable imitation is strongly related to the existence of payoff-controlling strategies, called zero-determinant strategies. However, these results were on the situations in which one repeated game is played by one group of players. It was pointed out that imitation of other players in the same group and imitation of other players in the same role in other groups generally result in different outcomes. Here, we investigate the existence condition of unbeatable imitation in the latter situations, described by two concurrent repeated games played by two different groups of players. We find that the existence condition of unbeatable imitation is stronger than the existence condition of unbeatable zero-determinant strategies, whereas both are very limited. Furthermore, we also find that adding exploration can make an imitation strategy unbeatable even in games where imitation without exploration can be beaten. Our findings suggest a strong relation between unbeatable imitation and payoff control even in the two concurrent repeated game situations.

[30] arXiv:2604.20348 (replaced) [pdf, html, other]
Title: Bimanual Robot Manipulation via Multi-Agent In-Context Learning
Alessio Palma, Indro Spinelli, Vignesh Prasad, Luca Scofano, Yufeng Jin, Georgia Chalvatzaki, Fabio Galasso
Comments: Accepted at CoRL 2026
Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)

Large Language Models (LLMs) have emerged as powerful reasoning engines for embodied control. In particular, In-Context Learning (ICL) enables off-the-shelf, text-only LLMs to predict robot actions without any task-specific training while preserving their generalization capabilities. Applying ICL to bimanual manipulation remains challenging as the high-dimensional joint action space and tight inter-arm coordination constraints rapidly overwhelm standard context windows. To address this, we introduce BiCICLe (Bimanual Coordinated In-Context Learning), the first framework that enables standard LLMs to perform few-shot bimanual manipulation without fine-tuning. BiCICLe frames bimanual control as a multi-agent leader-follower problem, decoupling the action space into sequential, conditioned single-arm predictions. Evaluated on 13 tasks from the TWIN benchmark, BiCICLe achieves 70.5% average success rate, outperforming the best training-free baseline by 6.1 percentage points and surpassing most supervised methods. We also demonstrate superior real-world performance on 3 tasks without hardware-specific retraining. The project page is available at this https URL

[31] arXiv:2607.09139 (replaced) [pdf, html, other]
Title: Control Laguerre Tessellation: Semi-discrete Optimal Transport Over Control Systems
Ripon C. Sarker, Abhishek Halder
Subjects: Optimization and Control (math.OC); Machine Learning (cs.LG); Multiagent Systems (cs.MA); Systems and Control (eess.SY)

We study the optimal transport of optimally controlled agents from a compactly supported absolutely continuous source to a discrete target measure. The ground cost for the transport is induced by the optimal cost of the agents' motion. When this ground cost satisfies the twist condition, the optimal transport map is given almost everywhere in terms of a Laguerre tessellation of the state space. We refer to this control-theoretic generalization of Laguerre tessellation as Control Laguerre Tessellation (CLT), and illustrate it for two ground costs induced by linear controlled agents with minimum energy and minimum time objectives.

[32] arXiv:2607.22465 (replaced) [pdf, html, other]
Title: ORACLE: Agentic AI Orchestrator Routing Via Adaptive Verifier Calibration Feedback
Ritik Raj, Souvik Kundu, Dheemanth Joshi, Tushar Krishna
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)

Modern enterprise agent deployments consist of a heterogeneous pool of large language models (LLMs) having diverse capabilities and cost. Existing model routing strategies optimize the quality-cost trade-off, while providing request-level static decisions. More recent solutions address agentic routing as a task-level selection with a serial verifier based router feedback loop. However, their fixed verifier suitable for homogeneous workloads may not generalize to heterogeneous batches of agentic tasks (example: coding, general conversational). Additionally, due to the verifier placement in the critical path of the loop, serving quality may be affected during multiple concurrent requests routing. To mitigate these issues, we present ORACLE. It is a concurrency-aware online routing mechanism that aligns adaptive routing with adaptive verification for feedback. ORACLE acts as a training-free drop-in 'feedback loop' on top of any model-selection policy to first classify the task type and then dynamically assigns a task-appropriate verifier. Further, we develop a delayed feedback strategy for concurrent requests that largely removes verifier latency from the dispatch critical path. We then present a post-routing dispatch scheduler, namely DISC. DISC reserves each task's peak KV footprint at admission and dispatches to an alternate backend when the reward gain from reduced wait exceeds the reward loss from lower accuracy. Extensive evaluation on SWE-bench, tau2-bench, and Terminal-Bench 2.0 shows that ORACLE improves the accuracy-cost frontier by up to 7 percentage points over state-of-the-art routing baselines, while ORACLE with DISC improves program throughput by up to 1.8x.

[33] arXiv:2609.26927 (replaced) [pdf, html, other]
Title: Building Socio-Affective Artificial Intelligence for Interactive Multi-Agent Simulations
David Berga
Comments: 37 pages, 27 figures, 42 tables
Subjects: Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Computer Science and Game Theory (cs.GT); Human-Computer Interaction (cs.HC); Multiagent Systems (cs.MA)

The objective of this article is to provide design principles and a software architecture for enabling interaction between humans and multiple agents in simulated dynamic worlds. This connects the current era of general artificial intelligence (AI/AGI) with the proliferation of transformer-based conversational agents and the increased computational capabilities. Given an overview of current and previous multi-agent theories of mind (socially and affectively-aware agents), the existence of an integrative design of agent interactions with themselves and with humans must be crucial for understanding how to create sustainable and governance in future human-agent reasoning systems. In this work is presented a software "AGIMUD" that integrates: A. socially-aware reasoning and emotion in agent behavior and interaction, B. a design of human multimodal scheme for human users, artificial agents and simulated worlds, and C. distributing the AI processing through the network to enable multiple autonomous agents. These integrations allow the dynamic world recreation as multi-user dungeons (MUDs) where both agents and humans can interact simultaneously in real time. Find the code online in this https URL.

[34] arXiv:2609.32259 (replaced) [pdf, html, other]
Title: Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
Vincent-Daniel Yun, Woosang Lim, Haneul Yoo, Sungjoo Yoo, Murali Annavaram, Sai Praneeth Karimireddy
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)

Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundancy, but prefill-free transfer across model families must handle differences in tokenization, model depth, and KV representations. To address these issues, we propose \textit{HeteroFold}, a prefill-free cross-family KV cache transfer method that keeps both the sender and receiver frozen. HeteroFold aligns model structures, maps the sender cache into the receiver space, and calibrates it to preserve receiver behavior. Across six transfer directions, HeteroFold achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings. It also matches text-based communication on the multi-agent benchmark. At 32K context length, Llama-3.1-8B$\rightarrow$Ministral-3-14B transfer is $10.7\times$ faster than Native Prefill and $1.18$--$1.47\times$ faster than the state-of-the-art prefill-free baselines, Dense Latent and KV Ridge. These results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill.

Total of 34 entries
Showing up to 2000 entries per page: fewer | more | all
We gratefully acknowledge support from our major funders, member institutions, , and all contributors.
About · Help · Contact · Subscribe · Copyright · Privacy · Accessibility · Operational Status (opens in new tab)
Major funding support from
Simons Foundation Simons Foundation International Schmidt Sciences