Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems
Abstract
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through textual reflection. Despite strong empirical results, these systems lack a unified account of coordination, memory improvement, and the role of external verification. We model orchestrator–worker interaction as a bilevel coordination game: under bounded coupling, the workers’ local-update game is an approximate potential game whose equilibrium slack is controlled by decomposition quality. We then analyse reflection as stochastic movement over semantic memory states. For free-form reflection, we derive a finite-time upper bound, prove worst-case tightness, and give a positive lower bound under a falsifiable persistent-harm condition. We further prove an information-theoretic impossibility result: no gate that observes only the generated transcript can improve uniformly over text-indistinguishable environments, whereas an environment-grounded gate can. Motivated by this separation, we introduce Stochastic Reflective Memory Ascent (SRMA), which accepts a candidate memory only after a grounded evaluation risk strictly decreases. Under calibration and non-degenerate corrective mass, SRMA converges exactly, geometrically or polynomially; matching constructions show that both rate regimes are order-tight. We also provide confidence gating for stochastic evaluation and re-anchoring guarantees for piecewise-stationary environments. Experiments instantiate these objects with environment-grounded metrics and test the predicted coordination and drift laws. On 500 SWE-bench instances, the complete Kimi-based system resolves versus a public mini-SWE-agent reference. Code available at https://github.com/YihangChen9/Bilevel-Coordinated-Reflection.
1UCL Centre for Artificial Intelligence 2University of Liverpool 3Huawei
1 Introduction
Multi-agent LLM systems have become a common recipe for tasks too large or structured for a single agent: an orchestrator decomposes the task, worker models solve the pieces, and the team improves by reflecting—writing critiques, hypotheses, and lessons into a shared textual memory that conditions subsequent generations (Wu et al. 2024; Hong et al. 2024; Shinn et al. 2023; Benkovich and Valkov 2026; Qian et al. 2025). Because model weights are frozen at test time, memory editing is the principal adaptation channel (Zhou et al. 2025; Xu et al. 2025; Zhang et al. 2025b), and such loops often work better when grounded by a test harness, simulator, execution engine, or formal checker.
The dominant account of these systems is nevertheless procedural. Existing frameworks (Zhang et al. 2025a; Hu et al. 2025; Dang et al. 2025; Wang et al. 2025) specify who communicates with whom and which buffer is updated, but not the strategic object that the agents stabilise to or the quantity that reflection improves. This leaves three unresolved questions. First, how does the orchestrator’s decomposition quality control worker coordination? Second, when does unconditional reflection plateau rather than converge? Third, why can an external verifier succeed where a stronger text-only critic may still fail?
We address these questions in a single framework. The orchestrator–worker pipeline is modelled as a bilevel coordination game whose follower subgame is an approximate potential game, and textual memory editing as a stochastic process over a discrete semantic state space. For free-form reflection, a one-sided drift condition yields a finite-time upper bound that is tight in the worst case; a universal positive floor requires an additional, explicitly testable persistent-harm condition—unconditional commitment alone is not enough.
We then isolate the informational role of verification: in two environments with identical text-generation laws but opposite meanings for the same reflections, any possibly randomised, history-dependent gate that observes only the transcript behaves identically and therefore cannot improve both—even an ideal text-only judge—whereas a grounded verifier distinguishes the pair and recovers geometric convergence.
Motivated by this separation, we introduce Stochastic Reflective Memory Ascent (SRMA), which commits a candidate memory only when a fixed grounded evaluation protocol certifies a strict decrease in verifier risk. Under calibration and non-degenerate corrective mass, SRMA converges exactly at order-tight geometric or polynomial rates; a confidence gate handles stochastic probes, and re-anchoring restores per-segment convergence under piecewise stationarity.
The theory is instantiated on a hidden-cap resource contest, Overcooked with an exact BFS value table, and SWE-bench (Jimenez et al. 2024). The controlled environments expose the strategic, memory, and drift quantities directly without an LLM-as-judge; on SWE-bench the complete Kimi-based system resolves instances () versus for the public mini-SWE-agent v2 reference.
In summary, we contribute: (1) a bilevel coordination game linking decomposition coupling to follower equilibrium slack (Sec. 3.1); (2) a two-sided drift analysis of free-form reflection, tight in the worst case, with a universal lower bound under persistent harmful commitment (Sec. 3.2); (3) an impossibility theorem for self-contained text-only gates, with a grounded comparator that converges geometrically (Sec. 3.3); (4) SRMA, with exact convergence, order-tight rates, and a finite-probe confidence extension (Sec. 3.4); and (5) mechanism-level validation on Resource Contest and Overcooked plus end-to-end results on SWE-bench (Sec. 4).
2 Related Work
Multi-agent LLM frameworks. Orchestrator–worker architectures such as AutoGen (Wu et al. 2024), MetaGPT (Hong et al. 2024) and Agyn (Benkovich and Valkov 2026) show strong empirical performance but offer no convergence analysis; failure modes such as hallucination cascades are documented empirically (Liu et al. 2026; Cemri et al. 2025). We provide the missing game-theoretic and stochastic-approximation foundations.
Self-reflection, self-evaluation, and grounding. Reflexion (Shinn et al. 2023) and Self-Refine (Madaan et al. 2023) improve outputs by appending self-generated critiques but may plateau, and correlated self-evaluation bias (Zheng et al. 2023; Panickssery et al. 2024; Wu et al. 2026) weakens model-based judges in practice. Our drift analysis separates a worst-case floor from the persistent-harm condition needed for a universal lower bound, and our indistinguishable-environment theorem shows that without an environment-dependent signal even an ideal text-only gate cannot be uniformly correct.
Potential games and drift analysis. Our followers’ subgame builds on exact and approximate potential games (Monderer and Shapley 1996; Candogan et al. 2011; Christodoulou and Gairing 2014) and weakly coupled team problems (Srikant and Başar 1992). The convergence analysis uses Foster–Lyapunov drift (Hajek 1982; Meyn and Tweedie 2009), classical stochastic approximation (Robbins and Monro 1951; Borkar 2008; Bertsekas and Tsitsiklis 2000) and, for the gated regime, multiplicative and variable drift theorems from randomised search heuristics (Doerr et al. 2012; Johannsen 2010; Lehre and Witt 2021); the recursion is the discrete stochastic analogue of Polyak–Łojasiewicz-type conditions (Karimi et al. 2016; Chung 1954). Two-timescale bilevel structure follows Borkar (1997); Hong et al. (2023).
3 Methodology
Longer derivations are deferred to the supplementary material.
3.1 Problem Formulation: Bilevel Coordination Game
We formalise the resolution of a complex user query . The objective is a joint structured output maximising a global utility (logical correctness, constraint satisfaction). In a naive single-agent paradigm the entire output is generated directly from the query via the frozen LLM kernel, , which for large tasks induces context dilution and reasoning degradation (Liu et al. 2024; Levy et al. 2024; Du et al. 2025). Contemporary systems instead let an orchestrator partition the task among workers (Wu et al. 2024; Hong et al. 2024; Liu et al. 2025).
We model this as a bilevel coordination game. The orchestrator (Leader) generates a strategy profile ,
| (1) |
assigning subtask to worker (Follower), who generates a local sub-solution ; the global output is .
Unlike the idealised independent decomposition of classical potential-game analyses (Monderer and Shapley 1996), real multi-agent LLM systems exhibit non-trivial cross-worker interactions: shared variables, common interfaces, joint constraints (Liu et al. 2026). We adopt a weakly coupled decomposition in the spirit of Srikant and Başar (1992); Candogan et al. (2011).
Assumption 1 (Weakly Coupled Decomposability).
Each worker action set is finite. The worker payoff is the local objective , while the system-level objective is . The global utility admits
| (2) |
where is an undirected interaction graph induced by , with each edge counted once, and . Let be worker ’s coupled neighbours and .
When the system reduces to the independent case; and jointly quantify decomposition quality. Since LLM generation is stochastic, the system objective is the expected global utility .
Lemma 1 (Approximate Potential Game).
Under Assumption 1 and fixed , the workers’ subgame is an -approximate potential game with potential and slack
| (3) |
Proof.
If worker unilaterally deviates from to ,
| (4) |
where the coupling residual sums at most terms each bounded by (since ), so . Every unilateral deviation thus changes the potential within of the local utility change (Candogan et al. 2011; Christodoulou and Gairing 2014). ∎
A rational worker performs -better-response updates: . If no worker has such a deviation, the current profile is by definition already an -approximate Nash equilibrium, so the dynamics below are well defined in all cases.
Theorem 1 (Convergence of the Followers’ Subgame).
Under Lemma 1, iterated -better-response updates converge in finitely many steps to a profile satisfying, for every worker ,
| (5) |
Thus is an -approximate pure-strategy Nash equilibrium of the explicitly defined local-payoff game.
Proof sketch.
Each update raises the potential by a strictly positive amount (Lemma 1); finite and bounded imply finite termination. Full proof in the supplementary material. ∎
Leader’s objective and decomposition quality.
The orchestrator anticipates the followers’ equilibrium and solves . Because depends on , the leader’s objective contains an explicit decomposition-quality term:
Corollary 1 (Leader’s Decomposition Trade-off).
Let and . For any -approximate equilibrium ,
| (6) |
Hence the leader maximises a lower bound that trades achievable local utility against coupling: a good decomposition simultaneously raises and shrinks . (Proof in the supplementary material.)
3.2 Dual-Memory Drift Dynamics and Hallucination Floors
LLM weights are frozen, so adaptation proceeds by editing external, non-parametric memories: an execution memory shared by workers and a strategy memory used by the orchestrator. For a fixed decomposition , let
and rescale utility so that the sub-optimality lies in . Let denote the history up to the -th memory update.
The key distinction is whether a proposed reflection is committed unconditionally or evaluated before it enters memory. Unconditional commitment alone does not imply a positive asymptotic error: a universal lower bound requires an explicit condition that harmful commitments keep injecting non-vanishing expected error. We therefore separate an upper guarantee, its worst-case tightness, and a genuine lower bound under persistent harmful drift.
Regime A: free-form reflection.
When every generated reflection is appended, corrective information and hallucinated information (Huang et al. 2025; Ji et al. 2024) are mixed in the same update. We summarise their net conditional effect by the following one-sided drift condition.
Assumption 2 (One-Sided Free-Form Drift).
There exist and such that
| (7) |
Here is the available corrective drift and is the mean residual error load from committed, ungrounded content; is a first-moment quantity, not a variance.
Theorem 2 (Finite-Time Upper Bound).
If and , then, for ,
| (8) |
Consequently, . (Proof in the supplementary material.)
Theorem 2 is an upper guarantee only; the next result is the strongest conclusion available from Assumption 2 alone.
Proposition 1 (Worst-Case Tightness).
For every and , there exists a free-form process satisfying Assumption 2 with and such that
| (9) |
Hence the upper bound cannot be uniformly improved over the one-sided drift class.
Proof sketch.
The deterministic recursion with maps into itself, attains (7) with equality, and converges to its unique fixed point . ∎
A lower bound that applies to every process requires a lower drift condition, directly testable by regressing the next-step error on the current error in free-form trajectories.
Assumption 3 (Persistent Harmful Commitment).
There exist and such that, on every reachable state,
| (10) |
The parameter upper-bounds how much of the current error can be removed in one expected update, whereas is a persistent net error load that remains because harmful reflections are committed without screening.
Theorem 3 (Universal Lower Bound).
Corollary 2 (Two-Sided Error Tube).
An operational corollary in the supplementary material re-expresses the tube via estimable per-step correction and harm rates.
Leader’s outer loop.
On the slower timescale, define for . Under the analogous upper drift condition with and , the same affine recursion yields the finite-episode bound . We use only this finite-episode statement and make no asymptotic leader-regret claim.
3.3 Why Grounding Is Necessary: Impossibility of Self-Contained Gates
The fundamental informational requirement is grounding: access to a signal whose law depends on the environment rather than only on the generated transcript. We formalise this through a pair of environments that are indistinguishable at the text level.
Text processes and gates.
Let a memory be a finite reflection sequence with append operation . Fix an initial memory and a proposal kernel over . An environment assigns a sub-optimality to every reachable memory. Across the class considered below, the environment changes this semantic value but not the proposal kernel or any other text-level law.
Definition 1 (Self-Contained Gate).
A self-contained gate is any possibly randomised, history-dependent acceptance rule measurable with respect to the generated text process and its internal randomness only. A grounded gate may additionally observe an environment-dependent signal, such as realised reward, simulator state, test execution, or a formal-checker result.
This class contains text-only LLM-as-judge systems. Correlated-evaluation bias can make such judges weaker in practice (Panickssery et al. 2024); the result below applies even to an ideal gate with unlimited text-processing capacity.
Ambiguous-pair construction.
Let be disjoint and satisfy for every reachable ; all remaining proposals are inert. Fix and define and . In environment , an accepted proposal from applies to the current error; in environment , the roles of and are swapped. Thus the same text is corrective in one environment and harmful in the other. Let and assume both environments start at the same .
Theorem 4 (Self-Gating Impossibility).
For every self-contained gate and every horizon ,
| (14) |
Moreover, if and the gate accepts at least one proposal from with positive probability by time , then the inequality is strict. In contrast, the free-form rule accepts everything and satisfies in both environments, whereas the grounded gate that observes accepts only the corrective class and satisfies in both environments.
Proof sketch.
Couple both environments with shared proposal and gate randomness; the accepted class-label sequence is then identical under and , and the reflection identity yields , strictly when and an ambiguous proposal is accepted with positive probability. The free-form and grounded rates follow from the induced affine recursions. Full proof in the supplementary material. ∎
Remark 1 (Scope of the impossibility result).
The theorem is minimax over text-indistinguishable environments; textual self-evaluation remains useful when the transcript itself certifies correctness (a fully checkable proof). When truth depends on external state—hidden caps, API responses, simulator state, an evolving repository—judge capacity cannot substitute for grounding.
3.4 SRMA: Verifier-Gated Reflection
Theorem 4 establishes why the gate must have access to an environment-separating signal. Exact convergence additionally requires the gate to compare a fixed error functional of the memory state, rather than two uncontrolled one-shot samples from a stochastic generator. We therefore separate the stochastic proposal mechanism from the grounded evaluation protocol.
Definition 2 (Verifier and Evaluation Risk).
A verifier is a deterministic map with deterministic score . Let be a fixed deterministic evaluation protocol, such as an exact planner or decoding with fixed randomness. The verifier risk of memory is
| (15) |
The pair is fixed independently of the reflection proposal distribution. It is grounded when its score depends on an environment signal that is not determined by the generated transcript alone. Grounding supplies information; calibration below connects the score to task utility.
Definition 3 (Verifier-Gated SRMA Update).
Given , compute the evaluation output and diagnostic . Sample a reflection and form . Accept iff
| (16) |
On acceptance set ; otherwise retain .
Gating on two stochastic one-shot outputs would not suffice: sample variation could accept a memory with worse expected performance. Exact guarantees therefore assume deterministic or exact expected-risk evaluation; a finite-sample extension follows below.
Assumption 4 (Verifier Calibration).
There exists such that, for every reachable memory,
| (17) |
Thus zero verifier risk certifies zero task sub-optimality. This assumption is appropriate for exact value tables and complete formal checkers; on incomplete test suites, our theorem concerns verifier risk only.
Assumption 5 (Non-Degenerate Corrective Mass).
There exist and such that, whenever ,
| (18) |
Assumption 6 (Proportional Accepted Decrement).
There exists such that
| (19) |
Proposition 2 (Monotone Multiplicative Drift).
Theorem 5 (Exact Verifier Convergence and Rates).
Proof sketch.
Taking expectations in (21) and applying Jensen’s inequality to gives ; the rates follow by the standard multiplicative/variable-drift comparison. is non-increasing and non-negative, hence converges almost surely, and forces the limit to be zero. Calibration transfers the bound to the utility gap. ∎
Proposition 3 (Rate Tightness for Verifier-Gated Reflection).
Proof sketch.
Accept with probability and set on acceptance, so both assumptions hold with equality. For , has constant expected increment, and convexity of gives (24); for , . Full proof in the supplementary material. ∎
Proposition 4 (Confidence-Gated Stochastic Evaluation).
Suppose deterministic is unavailable and instead for an i.i.d. score . At round , estimate the current and candidate risks with independent probes and let
| (25) |
Accept only when . Then, with probability at least , every accepted update strictly decreases the true expected verifier risk.
Proof sketch.
Hoeffding’s inequality bounds each of the two estimation errors by with joint failure probability at most ; a union bound over rounds completes the argument. Exact convergence requires deterministic or exact expected-risk evaluation, or with summable . ∎
Proposition 5 (Piecewise-Stationary Re-Anchoring).
Remark 2 (Falsifiability and rate prediction).
The exponent is observable: the geometric regime is linear in versus , the polynomial regime in versus with slope ; estimating from acceptance frequencies and from trajectory decay gives the closed-loop calibration of Sec. 4. The free-form drift parameters are likewise estimable from conditional drift regressions.
Practical realisation.
Algorithm 1 probes the candidate under the same fixed protocol and commits only a strict improvement; recomputing the current risk enables the re-anchoring of Proposition 5, and under stochastic evaluation line 8 is replaced by the test of Proposition 4.
4 Experiments
We evaluate the theory on Resource Contest (RC; Table 2), Overcooked (Table 1), and SWE-bench (Table 4). RC and Overcooked use frozen MiniMax-M2.7 agents; unless noted otherwise, results are meanstandard deviation over five seeds. SWE-bench uses the backbones listed in Table 4. All metrics come from environment ground truth or the repository test harness rather than an LLM judge. Full prompts, configurations, and per-seed trajectories are in the supplementary material.
Resource Contest.
RC is a hidden-cap allocation game: workers probe unknown caps and the orchestrator allocates a unit budget across workers. The optimal round reward is , and we report cumulative reward and regret . Clipping feedback is generated by the environment and therefore provides a grounded signal. The four settings vary difficulty: easy (, caps , ); hard (, caps , ; tightly packed caps test allocation precision); many (, caps , ; larger search space); and drift (, caps until , then ; a moving optimum tests re-adaptation). Since , the oracle -reward is , , and on easy/hard/many respectively.
| Layout | Greedy | No memory | Free-form | Self-gated | Grounded SRMA |
|---|---|---|---|---|---|
| cramped_room | |||||
| asymmetric_advantages | |||||
| centre_pots |
| Setting | Oracle | -greedy | No memory | SRMA |
|---|---|---|---|---|
| easy | 120 | |||
| hard | 160 | |||
| many | 180 |
SRMA reaches – of oracle reward. Execution memory adds reward points on average and reduces mean regret from to (): grounded cap evidence that is fragmented without memory becomes a functional coordination channel for the orchestrator.
Overcooked coordination.
We use Overcooked (Carroll et al. 2019) with three two-agent layouts, a horizon of , and an exact BFS verifier . The verifier is deterministic and supplies the risk used for both SRMA and the drift study. The layouts stress complementary coordination demands: mutual blocking in a tight kitchen (cramped_room), role specialisation (asymmetric_advantages), and contention over shared pots (centre_pots).
Grounded SRMA is best on every layout (Table 1). Relative to the text-only self-gate, it raises score by , , and , and reaches the first delivery in , , and steps versus , , and for self-gating. The ordered improvement from no memory to free-form, self-gating, and grounded SRMA separates decomposition, memory, and grounding effects.
Grounding and gate quality.
A proposal is downstream harmful when it increases an independently evaluated oracle task risk, not necessarily the verifier score used by the gate. Table 3 shows that grounding sharply improves both selectivity and final risk.
| Method | Harmful | Helpful | Risk |
|---|---|---|---|
| No reflection | N/A | N/A | |
| Free-form | |||
| Self-gate | |||
| Grounded SRMA |
Grounded SRMA halves final risk relative to self-gating; the residual downstream-harmful rate measures verifier–oracle miscalibration rather than a violation of monotonicity in the verifier’s own risk.
Gate-level drift predicts held-out trajectories.
From gate events across five seeds, we use three complete seeds for calibration and hold out two entire trajectories. A trajectory-level bootstrap gives , , and , with . Without fitting trajectory-level parameters, the plug-in prediction tracks the held-out risks with Pearson’s and , implying decay near . Bootstrap lower bounds and yield a conservative envelope above the empirical mean risk at every recorded step—an empirical certificate on the observed range, not a claim about unobserved states.
Statistical resolution.
A one-shot stochastic verifier falsely accepts of worsening proposals (score ); fixed cuts this to (score ) at verifier calls, and the adaptive gate matches that reliability (, score ) with only calls (), supporting Proposition 4: grounding supplies information, confidence control supplies resolution.
Piecewise stationarity.
In RC drift, the optimal cap changes at while previously written text remains unchanged, testing the re-anchoring mechanism of Proposition 5. The re-anchored grounded gate detects the shift in rounds, switches to the new optimum in , and incurs post-shift regret, versus , , and for the grounded stale-anchor variant—a cut in switch time and in regret—while the text-only gate fails to detect the change within rounds (regret ). Grounding detects the shift, but re-anchoring is required to replace stale memory quickly.
End-to-end software repair.
We evaluate the complete bilevel system on all 500 SWE-bench instances, with the repository test harness as the grounded verifier (an instance counts as resolved only if its submitted patch passes the harness). Each worker is a mini-SWE-agent v2 instance; the bilevel system runs such workers over a shared repository and workboard for up to three coordination rounds per episode and submits the highest- patch, whereas the mini-SWE v2 row is a single mini-SWE-agent v2 worker with no orchestrator or shared memory. The Free-form MA row keeps the same multi-agent coordination but commits every proposed reflection ungated (no verifier check), isolating the effect of SRMA’s grounded gate. The public leaderboard row is an external reference, not a controlled ablation.
| System | Backbone | Rate |
|---|---|---|
| mini-SWE v2† | DeepSeek | |
| Bilevel SRMA† | DeepSeek | |
| Free-form MA† | Kimi K2.5 | |
| mini-SWE v2 (public) | Kimi K2.5 | |
| Bilevel SRMA | Kimi K2.5 |
On the Kimi K2.5 backbone the grounded gate is decisive: Bilevel SRMA resolves against for free-form (ungated) multi-agent reflection at matched backbone and budget, and exceeds the external public mini-SWE-agent v2 reference (). The controlled DeepSeek runs show the same direction ( vs. ), indicating that the gain comes from grounded, gated coordination rather than from raw model or compute.
5 Conclusion
We gave multi-agent LLM reflection a conditional, information-aware theory: bilevel coupling controls follower equilibrium slack, persistent harmful commitment creates free-form error floors, and no transcript-only gate can improve uniformly when the truth of a reflection depends on external state. SRMA supplies the missing grounding and converges exactly at order-tight geometric or polynomial rates, with confidence-gating and re-anchoring extensions; experiments on Resource Contest, Overcooked, and SWE-bench support the predicted coordination, grounding, and resolution mechanisms.
Limitations.
The guarantees are conditional: bounded coupling, finite action sets, verifier calibration, and non-degenerate corrective mass need not hold in open-ended agent tasks; drift parameters are validated only on observed trajectories; incomplete test suites guarantee monotonicity only for verifier risk, not true task utility; and re-anchoring gives per-segment convergence without a general switching-regret bound. Multi-agent coordination also spends substantial tokens before the final answer, and the Kimi result is compared with a public leaderboard run rather than a controlled method-only comparison. Future work should jointly optimise memory quality and budget-aware termination.
References
- Agyn: a multi-agent system for team-based autonomous software engineering. arXiv preprint arXiv:2602.01465. Cited by: §1, §2.
- Gradient convergence in gradient methods with errors. SIAM Journal on Optimization 10 (3), pp. 627–642. Cited by: §2.
- Stochastic approximation with two time scales. Systems & Control Letters 29 (5), pp. 291–294. Cited by: §2.
- Stochastic approximation: a dynamical systems viewpoint. Cambridge University Press. Cited by: §2.
- Flows and decompositions of games: harmonic and potential games. Mathematics of Operations Research 36 (3), pp. 474–503. Cited by: §2, §3.1, §3.1.
- On the utility of learning about humans for human-AI coordination. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: §4.
- Why do multi-agent LLM systems fail?. In Advances in Neural Information Processing Systems (NeurIPS), External Links: Link Cited by: §2.
- The price of stability of weighted congestion games. In International Colloquium on Automata, Languages, and Programming (ICALP), Cited by: §2, §3.1.
- On a stochastic approximation method. The Annals of Mathematical Statistics 25 (3), pp. 463–483. Cited by: §2.
- Multi-agent collaboration via evolving orchestration. In Advances in Neural Information Processing Systems (NeurIPS), Note: arXiv:2505.19591 Cited by: §1.
- Multiplicative drift analysis. Algorithmica 64 (4), pp. 673–697. Cited by: §2.
- Context length alone hurts LLM performance despite perfect retrieval. In Findings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), Note: arXiv:2510.05381 Cited by: §3.1.
- Hitting-time and occupation-time bounds implied by drift analysis with applications. Advances in Applied Probability 14 (3), pp. 502–525. Cited by: §2.
- A two-timescale stochastic algorithm framework for bilevel optimization: complexity analysis and application to actor-critic. SIAM Journal on Optimization 33 (1), pp. 147–180. Cited by: §2.
- MetaGPT: meta programming for a multi-agent collaborative framework. In International Conference on Learning Representations (ICLR), Cited by: §1, §2, §3.1.
- Automated design of agentic systems. In The Thirteenth International Conference on Learning Representations (ICLR), External Links: Link Cited by: §1.
- A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems. Cited by: §3.2.
- Survey of hallucination in natural language generation. ACM Computing Surveys 55 (12), pp. 1–38. Cited by: §3.2.
- SWE-bench: can language models resolve real-world GitHub issues?. In International Conference on Learning Representations (ICLR), Cited by: §1.
- Random combinatorial structures and randomized search heuristics. Ph.D. Thesis, Universität des Saarlandes. Cited by: §2.
- Linear convergence of gradient and proximal-gradient methods under the Polyak–Łojasiewicz condition. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases (ECML-PKDD), pp. 795–811. Cited by: §2.
- Tail bounds on hitting times of randomized search heuristics using variable drift analysis. Combinatorics, Probability and Computing 30 (4), pp. 550–569. Cited by: §2.
- Same task, more tokens: the impact of input length on the reasoning performance of large language models. In Association for Computational Linguistics (ACL), Cited by: §3.1.
- Lost in the middle: how language models use long contexts. In Transactions of the Association for Computational Linguistics (TACL), Cited by: §3.1.
- Select-then-decompose: from empirical analysis to adaptive selection strategy for task decomposition in large language models. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), Note: arXiv:2510.17922 Cited by: §3.1.
- AgentHallu: benchmarking automated hallucination attribution of LLM-based agents. arXiv preprint arXiv:2601.06818. Cited by: §2, §3.1.
- Self-refine: iterative refinement with self-feedback. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: §2.
- Markov chains and stochastic stability. 2nd edition, Cambridge University Press. Cited by: §2.
- Potential games. Games and Economic Behavior 14 (1), pp. 124–143. Cited by: §2, §3.1.
- LLM evaluators recognize and favor their own generations. arXiv preprint arXiv:2404.13076. Cited by: §2, §3.3.
- Scaling large language model-based multi-agent collaboration. In The Thirteenth International Conference on Learning Representations (ICLR), External Links: Link Cited by: §1.
- A stochastic approximation method. The Annals of Mathematical Statistics 22 (3), pp. 400–407. Cited by: §2.
- Reflexion: language agents with verbal reinforcement learning. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: §1, §2.
- Asymptotic solutions of weakly coupled stochastic teams with nonclassical information. IEEE Transactions on Automatic Control 37 (2), pp. 163–173. Cited by: §2, §3.1.
- EvoAgentX: an automated framework for evolving agentic workflows. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), System Demonstrations, pp. 643–655. External Links: Link Cited by: §1.
- AutoGen: enabling next-gen LLM applications via multi-agent conversations. In First Conference on Language Modeling (COLM), Cited by: §1, §2, §3.1.
- Council mode: a heterogeneous multi-agent consensus framework for reducing LLM hallucination and bias. arXiv preprint arXiv:2604.02923. Cited by: §2.
- A-MEM: agentic memory for LLM agents. In Advances in Neural Information Processing Systems (NeurIPS), External Links: Link Cited by: §1.
- AFlow: automating agentic workflow generation. In The Thirteenth International Conference on Learning Representations (ICLR), External Links: Link Cited by: §1.
- A survey on the memory mechanism of large language model-based agents. ACM Transactions on Information Systems 43 (6), pp. 155. External Links: Document, Link Cited by: §1.
- Judging llm-as-a-judge with mt-bench and chatbot arena. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: §2.
- Memento: fine-tuning LLM agents without fine-tuning LLMs. arXiv preprint arXiv:2508.16153. External Links: Link Cited by: §1.