Exploring Collaboration between a language and a non-language agent
Abstract
LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through natural language. However, in many important domains like game playing and robotics, the strongest available agents are not language models. Integrating non-language agents with LLMs would require verbalization: compressing their rich continuous representations into sparse textual summaries at each interaction step. To study whether verbalization constitutes a bottleneck, we introduce LLAMIA-Bench, a suite of six diverse collaborative chess tasks spanning three facets: behavioral imitation, state assessment, and natural-language explanation. Each task instantiates a well-established chess problem that neither the LLM nor the chess engine can solve alone. To solve LLM collaboration with non-language agents, we introduce latent state internalization, which projects the subagent’s continuous representations directly into the LLM’s token stream as learned state tokens, with dynamic re-encoding as actions advance the environment state. Comparing internalization to verbalized integration, our experiments reveal a consistent verbalization debt: the performance gap widens throughout training and persists as the LLM scales from 4B to 14B parameters. A single 14B model, LLAMIA, trained with latent state internalization, matches or exceeds task specialists and frontier models including GPT-5.1 with tool access across all benchmark tasks, and generalizes out-of-distribution where task-specific finetunes collapse.
1 Introduction
Large language models (LLMs) are increasingly deployed as general-purpose orchestrators that coordinate tools and agents to solve complex tasks (Hong et al., 2024; Tran et al., 2025). A central pattern in this paradigm is collaboration with subagents: specialized agents trained to excel within a narrow domain (Anthropic, 2025; OpenAI, 2025). This collaboration allows orchestrator LLMs to utilize the subagent’s domain specific intelligence (Hong et al., 2024) and preserve its context length by task delegation (Zhang et al., 2024b), enabling effective multi-step reasoning and planning over long horizons. Today, this collaboration is mediated entirely through natural language: the LLM invokes the subagent, receives a natural language description of its output, and reasons over that description to decide subsequent actions(Tran et al., 2025). Verbal collaboration is natural when both agents are language models as they share a vocabulary and can express their state in words. However, in many important domains such as game playing, robotics, and autonomous driving the strongest available agents are not language models; their expertise is encoded in internal representations: latent states capturing policy, value estimates, and learned features like AlphaZero (Silver et al., 2017) and RT-1 (Brohan et al., 2022). This mismatch between LLMs and non-language agents’ input space raises a fundamental question:
-
How can LLMs effectively collaborate and jointly reason with non-language subagents?
The depth of collaboration between LLMs and subagents can solve many useful tasks that neither can solve alone. Consider chess: LLMs have been trained on more chess literature than most experts will study in a lifetime, yet they cannot leverage it to play the game competently, trailing far behind modern engines and experts (Kolasani et al., 2025). Conversely, pretrained engines surpassed human grandmasters decades ago(Campbell et al., 2002), yet they remain narrow specialists that are unable to explain the rationale behind a move or strategize under different contexts (Jhamtani et al., 2018; Lee et al., 2022). And there exist tasks like game commentary, preparing against an opponent, and designing interesting puzzles which require both the chess engine’s deep positional understanding and the LLM’s ability to reason over human intent. Effective collaboration between LLMs and the subagent can unlock these applications.
Existing approaches to LLM-subagent collaboration attempt to bridge this gap symbolically, by verbalizing the subagent’s outputs into natural language before passing them to the LLM. Early work finetunes language models on textual descriptions of agent actions and value estimates (Zang et al., 2019; Lee et al., 2022), while more recent systems rely on in-context learning and tool-calling interfaces to surface subagent outputs at inference time (Schick et al., 2023; Yao et al., 2022; Kim et al., 2025). However, these approaches share a common assumption: that the subagent’s expertise can be faithfully verbalized. In this work we demonstrate that this assumption is fundamentally limiting. A chess engine’s latent representation encodes positional structures, long-range tactical motifs, and learned look-ahead (Jenner et al., 2024)– semantic features that cannot be translated to text faithfully. Verbalization therefore forces the subagent’s representations through a lossy bottleneck, collapsing rich latent structure into surface-level descriptions. We call this concept the Verbalization Debt Zhu et al. (2025) and quantify its downstream cost under controlled, heterogeneous LLM–agent collaboration. Moreover, we show that this error compounds across interactions: in multi-step settings, each exchange between the agents strips away details and accumulates errors over the reasoning horizon, diminishing the benefits that motivated this collaboration in the first place.
Internalization: To bridge this gap we introduce latent state internalization, a paradigm in which an LLM reasons over a non-language agent’s internal state over a single trace of three interleaved token types: language tokens (the LLM’s chain-of-thought), action tokens (moves that advance the environment), and latent state tokens (the agent’s penultimate-layer activations projected into the LLM’s embedding space). As shown in Figure 2, the paradigm consists of three distinct steps: the LLM reasons in language and generates actions that evolve the environment state, such as counterfactual states within its CoT; The LLM on demand requests the agent to evaluate a state, either the current position or a counterfactual reached by a candidate move; the agent encodes the resulting state and its activations are projected into latent tokens appended to the reasoning trace, dynamically re-encoded after each state transition. We train a lightweight three-layer MLP, LatentBridge (Section 2.2), that learns the projection from the agent’s internal state to the LLM’s token space.
LLAMIA: We instantiate latent state internalization by training an LLM backbone in two stages: supervised projector alignment followed by reinforcement learning (DAPO), detailed in Section 2.3. We call the resulting model LLAMIA (Large Language and Action Models with Internal Agents). Beyond the technical contribution, internalization opens new avenues for real world applications. Because internalization requires access to model weights, it cannot be applied directly to closed-source models. LLAMIA resolves this tension by acting as a bridge: it interacts with LLMs in natural language and internalizes the subagent, giving closed-weight models indirect but faithful access to the non-language agent’s expertise. This makes real-world creative applications like grounded game commentary and opponent-specific preparation deployable, without retraining closed-weight models.
Verbalization Debt: To establish and analyze the effect of internalization when compared to verbalization, we train LLAMIA-Verb, identical to LLAMIA except that the subagent’s outputs reach the LLM as text rather than latent tokens. LLAMIA consistently achieves higher reward across all tasks throughout training (Figure 3). The gap widens on tasks requiring deeper multi-step integration. These results show that verbalization is a fundamental bottleneck when non-language agents are treated as tools.
LLAMIA-Bench, Chess as a testbed: Despite abundant applications, LLM collaboration with non-language agents remains underexplored. A central reason is the lack of environments that support studying this collaboration at scale: diverse tasks with verifiable metrics and open pretrained agents. Chess is a perfect testbed: decades of human-engine collaboration on commentary, preparation, and puzzles have produced diverse tasks with verifiable metrics, strong open pretrained agents, established evaluation protocols, and large public corpora like Lichess Feng et al. (2024). We therefore use chess as our primary testbed and introduce LLAMIA-Bench (Appendix B), a curated suite of six tasks spanning behavior cloning across skill levels, puzzle interest and difficulty estimation, move annotation and game-level commentary. We introduce a new dataset curated from Agadmator’s YouTube Channel for game level commentary.
Results. A single LLAMIA-14B model matches or exceeds every task-specific specialist and frontier model across all six LLAMIA-Bench tasks (Figure 1). The verbalization debt is sharpest on tasks requiring multi-step state tracking or non-verbalizable signals: LLAMIA-Verb’s reward stays flat on commentary and puzzle interest despite identical compute, indicating that verbalization discards task-relevant information needed for effective multi-step collaboration . Internalization also shapes the kind of collaboration that emerges: LLAMIA develops counterfactual queries and multi-step lookahead strategies absent from LLAMIA-Verb, which collapses to engine-follow regardless of task, suggesting that access to the full latent state is what makes richer collaboration learnable. Beyond benchmark numbers, LLAMIA reproduces human behavioral signatures: under time pressure it commits the same blunders humans do, and at different skill levels its attention concentrates on the same pieces human players prioritize. A human study confirms this: LLAMIA’s gameplay passes as human in the majority of trials, and its commentary is preferred on both strategic insight and explanatory depth compared to the verbalized baseline.
Our contributions are fourfold:
- 1.
Latent state internalization (Section 2): A paradigm for LLM–non-language agent collaboration, enabling LLMs to reason over interleaved chain of tokens.
- 2.
LLAMIA (Section 2): A two-stage training framework of self-supervised projector alignment followed by end-to-end RL (DAPO) that yields a single model, LLAMIA, achieving state-of-the-art performance across diverse collaborative tasks. LLAMIA enables real world application by giving close-weight models faithful access to subagents.
- 3.
Verbalization Debt (Section 3): Through controlled ablations against LLAMIA-Verb, we give the first empirical quantification of the Verbalization Debt (the performance gap between internalized and verbalized integration) in heterogeneous LLM-agent collaboration, suboptimal in both performance and compute and widening on tasks requiring deeper multi-step collaboration.
- 4.
LLAMIA-Bench (Section 3): A curated benchmark of six chess tasks spanning behavior cloning, puzzle understanding, commentary, and planning.
2 Methodology
2.1 Formulation
We formalize latent state internalization through a running example. An LLM plays chess with access to a pretrained engine exposed through a tool API: functions to read the board state and legal moves, advance the game by making moves, and—critically—query the engine’s assessment of any position via get_policy (full schema in Section J.4.3). The first two categories handle environment interaction; get_policy is the interface to the subagent, and what it returns is the variable this paper studies.
“The knight should develop. Nf6. get_policy() ( ) Nc2: P=0.34, +0.12; e4: P=0.21, +0.08; …"
The LLM reasons in language, plays a knight move to f6 (advancing the game to position ), then calls get_policy. The engine returns a text summary—per-move prior probabilities and value estimates—that both model variants receive. In LLAMIA, the call additionally injects continuous state tokens ( ) projected from the engine’s internal activations. In LLAMIA-Verb, only the text is returned. Three token types thus interleave in a trace: language tokens (the LLM’s chain-of-thought), action tokens (moves that advance the game), and state tokens (the engine’s projected latent representation, present only under internalization).
More formally: a policy (the LLM) interacts with an environment whose states evolve through actions . A pretrained subagent with encoder maps each state to a latent representation . The LLM queries the subagent at self-chosen moments via get_policy, specifying either the current position or a hypothetical state reached by a candidate move. Over a -step interaction the policy produces a trace :
| (1) |
where are language tokens (including the text returned by get_policy), are actions, and are contiguous state tokens injected alongside the text response. In LLAMIA-Verb, : no state tokens appear, and the LLM reasons over text alone. In LLAMIA, . How these state tokens are produced defines the internalization method.
Verbalization.
Every get_policy call returns a text serialization of the subagent’s output:
| (2) |
Both LLAMIA and LLAMIA-Verb receive . A typical return lists the engine’s top moves with prior probabilities and value estimates. This captures the subagent’s headline assessment but discards the remaining structure in : the full distribution over all legal moves, the value landscape across candidate continuations, and positional features like piece coordination and king safety that interpretability work has identified in engine hidden layers Jenner et al. (2024).
Internalization.
In LLAMIA, each get_policy call additionally produces continuous tokens by projecting the subagent’s full latent state into the LLM’s embedding space via a learned projection, LatentBridge:
| (3) |
produces continuous embeddings of dimension matching the LLM’s hidden size. These state tokens sit alongside language and action tokens in the trace (Figure 2), and the LLM attends over all three types jointly. Where the text serialization imposes a fixed summary regardless of what the current reasoning step requires, latent tokens let the LLM’s attention selectively read different aspects of the representation at each step. Gradients flow from the training objective through the LLM back to , so the projection adapts to the task.
2.2 Architecture
Subagent. We instantiate with Lc0-BT4 Monroe and Chalmers (2024), the strongest open-source chess engine, a 15-layer Transformer encoder (240 M parameters) whose representations encode positional features, piece-value geometry, and lookahead-related structure Jenner et al. (2024), producing . We ablate this choice across five Lc0 variants of varying strength in Appendix E.2.
Large Language Model. We use Qwen3 Yang et al. (2025), the strongest open-weight LLM at this scale at the time of training, as the backbone . The tool API (Section J.4.3) is provided in the system prompt via Hermes-format function calling; Qwen3 natively supports structured tool calls without additional training. To accept state tokens, contiguous positions in the input sequence serve as placeholders whose embeddings are replaced by the projected state .
LatentBridge. The projection (Eq. 3) is a three-layer MLP with GeLU activations, motivated by the projector design in vision-language models Liu et al. (2023a). It maps Lc0’s -dimensional latent state into embeddings of dimension matching the LLM’s hidden size. The resulting input to the LLM is a mixed sequence (Figure 2). We use the penultimate block (layer 14 of 15) as we observe empirically that its held-out Stage-1 alignment loss is lowest across all blocks (Section E.5). Prior interpretability research Jenner et al. (2024); Lin et al. (2026) has shown that this layer in BT4 network locates value, square, and look-ahead-to-action features as well. We set , where downstream performance saturates in a projector-and-policy sweep giving us an optimal token cost to performance tradeoff Section E.3.
2.3 Training
Training proceeds in two stages. Stage 1 aligns the subagent’s representations with the LLM’s embedding space while keeping the LLM frozen. Stage 2 trains the LLM and LatentBridge jointly via reinforcement learning. The subagent is frozen throughout.
2.3.1 Stage 1: Projector Alignment
We train only while keeping frozen, on a dataset of state–policy pairs from the subagent’s self-play. Each example pairs the projected state tokens with a language prompt (e.g., “Analyze position: top move?”; format in Section J.1), and the model learns to generate the correct action token via cross-entropy:
Because the LLM is frozen, the projector trains on abundant agent-generated data without risking catastrophic forgetting of language capabilities.
2.3.2 Stage 2: Reinforcement Learning
Stage 2 unfreezes both and and optimizes them jointly via DAPO Yu et al. (2025), a group-relative policy optimization variant with asymmetric clipping. Rollouts produce complete traces (Eq. 1). The policy gradient is computed over positions where the LLM generates: language tokens (index set ) and action tokens (index set ), collectively . State token positions are agent-injected and excluded via gradient masking. Writing for the token at position , the importance-sampling ratio between the current policy and the reference policy from the previous iteration depends on the integration channel:
| (4) | ||||
| (5) |
The DAPO objective maximizes:
| (6) |
where is the importance ratio (Eq. 5), is the group-normalized advantage computed over rollouts sharing the same prompt Yu et al. (2025), and the asymmetric clip bounds encourage exploration. Each LLAMIA-Bench task defines a scalar outcome reward based on its evaluation metric (Appendix B).
Because is jointly optimized, the RL objective shapes the integration interface end-to-end: the projector learns what representation to present, while the LLM learns how to reason over it. The tool call is part of the policy’s action space, so RL also learns when and whether to query it.
3 Experiments and Results
We evaluate LLAMIA on LLAMIA-Bench, a suite of six tasks spanning three collaboration facets: behavioral imitation, state assessment, and natural-language explanation (Appendix B). For comparing LLAMIA with closed-source models where training is not possible, we pair them with Lc0 as a verbalized tool. To evaluate against open-source models we create three settings: an untrained tool-use baseline (Qwen3-14B + Lc0), a training-matched verbalization system (LLAMIA-Verb), and our internalized model (LLAMIA), isolating the integration interface as the sole variable. Within the open-source trained systems, we report both SFT and SFT + DAPO checkpoints to disentangle the contributions of supervised pretraining and reinforcement learning.
3.1 LLAMIA-Bench
LLAMIA-Bench spans six tasks drawn from prominent problems in chess literature and industry, each unsolvable by either component alone: the subagent produces no language, while the LLM lacks the positional signals that make chess-specific judgments possible. The suite covers behavior cloning, puzzle understanding, move annotation, and game-level commentary (details in Appendix B). Behavior cloning, difficulty estimation, and move annotation are drawn from established benchmarks McIlroy-Young et al. (2020); Lichess.org (2024b); Jhamtani et al. (2018). We introduce three new evaluation targets: Wild BC, three OOD splits (GM-25, Low-Time, Elo) probing generalization to grandmaster play, time pressure, and asymmetric skill-gap; interest estimation, ranking puzzles by community-derived interestingness scores, a signal with no verbal proxy in any engine output; and Agadmator-2K, the first large-scale game-level commentary dataset of 1,900 narrated games with move-aligned transcripts. Full dataset descriptions, metrics, and per-task prompts appear in Appendices B–J.
3.2 Setup
Baselines
We compare four systems. (1) GPT-5.1 + Lc0: the strongest frontier model with verbalized Lc0-BT4 tool access. (2) Qwen3-14B + Lc0: the LLAMIA backbone with the same verbalized tool and 5-shot prompting, without training. (3) LLAMIA-Verb: the matched-recipe ablation with the latent channel replaced by the verbalized tool, isolating the integration interface as the sole variable. (4) LLAMIA: our full system with latent state-token integration. Extended comparisons including LLM-only baselines, SFT and DAPO checkpoints at all three model sizes (4B, 8B, and 14B), frontier model comparisons, and task-specific experts are in Appendix D. Training details and hyperparameters are in Appendix A.
Metrics
Each task has a primary metric detailed in Appendix B: move-matching accuracy (behavior cloning), Spearman (difficulty and interest), and G-eval and BLEU-2 (annotation and commentary). Primary metrics include 95% bootstrap confidence intervals where sample sizes warrant; per-task breakdowns with CIs appear in Appendix D.
3.3 Main Results
LLAMIA-14B achieves the highest score on all six LLAMIA-Bench tasks (Figure 3c), surpassing frontier verbalized systems an order of magnitude larger and remaining competitive with dedicated task-specific finetunes that are trained on substantially more in-domain data. On behavior cloning, Maia and Allie are trained on tens of millions of chess-specific games, against LLAMIA’s general-purpose backbone. LLAMIA-14B remains inside the expert band on the in-distribution Maia split and surpasses the strongest expert by a wide margin on the OOD Wild splits (GM-25, Low-Time, Elo), which probe regimes absent from the experts’ blitz-only training mixture. Latent access to the engine’s policy and value structure thus generalizes more reliably than direct supervision on a narrower distribution. Interest and commentary have no dedicated task-specific baseline at all, no published system predicts puzzle interestingness or generates grounded move commentary from engine state, which is itself evidence that these tasks require the joint reasoning LLAMIA provides rather than a narrower specialist. The advantage over GPT-5 + Lc0 does not require the 14B backbone: LLAMIA-8B already leads on all six tasks, and LLAMIA-4B on four of six (Table 18). Per-task evaluations with additional metrics, baselines, and confidence intervals are in Appendix D.
Latent tokens enable new evaluation targets.
Puzzle Interest requires ranking positions by community-derived interestingness, a signal that depends on the engine’s policy distribution and value gradients across candidate moves. No verbalized engine output carries these features. Every verbalized system scores on Interest regardless of model scale or frontier capability; LLAMIA-14B reaches (Figure 3c). Verbalization has zero useful signal for this task, while latent tokens give the LLM direct access to the distributional structure that defines interestingness.
3.4 Verbalization Debt
LLAMIA and LLAMIA-Verb share the same 14B backbone, Lc0-BT4 subagent, and DAPO recipe; the only difference is whether the subagent’s state reaches the LLM as latent tokens or as verbalized text. We define the resulting performance gap as the verbalization debt.
Verbalization debt is significant across all tasks
Figure 3 shows that the verbalization debt is consistent across all six tasks. The gap is largest on Interest and Commentary, where the target signal lives in the engine’s full policy distribution or value landscape and has no faithful text equivalent, and smallest on in-distribution behavior cloning, where the engine’s top- moves already approximate the answer and the verbal summary loses little.
The debt persists across backbone scale.
Increasing the LLM from 4B to 14B improves both systems, but the debt persists at every scale (Table 18). On Interest, LLAMIA-Verb-14B scores while LLAMIA-4B already reaches .
The debt widens throughout training.
The debt grows throughout DAPO, reaching – by the end of training (Figure 3a). Per-task reward curves (Figure 3b) reveal where the verbal channel saturates: LLAMIA-Verb gains partial signal on behavior cloning and difficulty, where the verbalized output carries a useful proxy (top- moves, solution length), but stays near-flat on interest and commentary, where no such proxy exists.
3.5 Ablations
To further understand the verbalization debt and isolate the contribution of internalization and reinforcement learning we conduct the following ablations and compare in Table 1. LLM-Only trains with RL but no engine, so any gain has to come from the weights. LLM-ChessCLIP trains with RL and the same latent slots as LLAMIA, but filled by a raw board encoder (ChessCLIP, a PaLM-E-style injection) rather than Lc0’s state, so any gain has to come from capacity rather than content. Qwen3+Lc0 (untr.) is untrained tool use. LLAMIA-Verb is RL on top of verbalized text. LLAMIA-SFT (latent) removes RL, the reasoning trace, and the invocation policy, leaving only the latent channel. LLAMIA is the full system. Extended controls, latent-only, shuffled tokens, per-task probes, and templates, are in Sections E.4 and C.5.
System BC-M BC-W Diff. Int. Rat. Comm. LLM-Only (RL, no engine) LLM-ChessCLIP (RL, raw encoder) Qwen3 + Lc0 (untr. tool use) LLAMIA-Verb (RL, text only) LLAMIA-SFT (latent, no RL) LLAMIA (latent, RL)
The gain comes from Lc0’s latent state, not from weights or capacity.
LLM-Only and LLM-ChessCLIP get the same RL recipe as LLAMIA and still land near or below untrained tool use, so DAPO cannot manufacture the missing expertise on its own, whether it is asked to bake it into the weights or to make sense of slots filled with the wrong content. Shuffling LLAMIA’s own latent tokens produces the same collapse toward LLAMIA-Verb even though the token count never changes (Section E.4). The pattern only breaks when those slots carry Lc0’s own policy and value representations. Further, LLAMIA-SFT, with no RL recovers most of the verbalization debt. However, it compounds the effect of internalization by bringing the improvements in multi step strategies, where the model has to plan across latent states i.e. Game Commentary and Rationale generation. We further show this through the emergent latent collaboration strategies in Section 3.6.
3.6 Collaboration Strategies
Does the integration interface determine how the model learns to use the subagent, or only how well it performs? We classify subagent invocations during evaluation into five recurring strategies and trace their evolution during DAPO training (Figure 6; strategy definitions and per-task heatmaps in Figure 4). The five strategies are engine-follow (adopt the top recommendation), consult-then-override (query then diverge), counterfactual query (play a hypothetical move, re-invoke, compare states), multi-step lookahead (chain two or three counterfactual sequences), and abstention (act from language knowledge alone). A GPT-4o judge classifies 500 episodes per task per system ( vs. human raters).
Internalization produces task-specific collaboration; verbalization collapses it.
LLAMIA adapts its strategy to the task: engine-follow dominates gameplay (65%), consult-then-override dominates behavior cloning (48%), and counterfactual query dominates commentary (40%). LLAMIA-Verb collapses to engine-follow on every task (62–76%), regardless of what the task requires (Figure 4). The verbalized channel returns the same compressed summary no matter how the model queries it, so RL converges on a single use pattern.
The learned strategy makes internalization inference-cost neutral.
Because it reasons over the full latent state, LLAMIA learns to invoke the subagent less often than LLAMIA-Verb ( vs. calls per query at 14B). Each latent invocation adds a fixed tokens ( vs. ), but the lower call count offsets this, so average tokens-per-query and wall-clock latency are comparable or lower than the verbalized interface; training cost stays within of the verbalized pipeline at every scale (Section A.1). Latent internalization therefore does not trade accuracy for inference cost.
3.7 Human Evaluation
Gameplay.
Skilled players (, all Elo) cannot reliably distinguish LLAMIA from a human opponent: detection falls below chance (Figure 5), and post-game ratings place LLAMIA alongside Maia* on perceived human-likeness (Figure 5). Maia is trained directly on millions of move distributions; LLAMIA receives no human-move supervision. Instead, behavioral signatures such as time-pressure blunders and skill-appropriate piece saliency emerge from latent-state conditioning alone. LLAMIA-Verb, trained with the same backbone and DAPO budget, is detected at rates well above chance, consistent with the stylistic regularities that verbal summaries impose on move selection.
Commentary.
Both systems achieve comparable factual accuracy: material balance, initiative assessment, and basic evaluations survive verbalization reasonably well. The gap concentrates on strategic insight (Figure 5), where participants rate LLAMIA higher by a wider margin on the Insight dimension than on Accuracy. Text preserves what is happening on the board, but explaining why a move is strong requires representational features (policy gradients, value topology, look-ahead depth) that do not survive verbal compression.
Latent tokens are instruction-modulated.
On a fixed back-rank mate position (Figure 5), LLAMIA’s attention over the latent tokens shifts with the target Elo: at 2000 it concentrates on the mating geometry, at 1100 it disperses to material. The latent representation is identical in both cases; what changes is the LLM’s reading, conditioned on the natural-language Elo instruction. The projected state functions as a perceptual input shaped by task context, not a static feature vector.
3.8 Generalization Beyond Chess
To provide initial evidence that LLAMIA can transfer beyond chess, we instantiate it on Go. We use KataGo-b18 Wu (2019), a state-of-the-art Go neural engine, as the non-language subagent. We train a three-layer LatentBridge, with the first-layer width matched to KataGo’s trunk channels and the board intersections represented as spatial tokens, using the same two-stage DAPO procedure on rank-conditioned behavior cloning. LLAMIA-Go-B achieves / top-1 human move-match at ranks 5k/5d using only k positions, matching the rank-calibrated KataGo-HumanSL expert Wu (2024) and outperforming the verbalized control by points. The latent-over-verbal gap remains consistent across the 4B, 8B, and 14B model scales (Appendix F). These results provide strong evidence that latent collaboration is not specific to chess.
4 Related Work
The dominant paradigm for LLM-agent integration is text-mediated: ReAct Yao et al. (2022), Toolformer Schick et al. (2023), and multi-agent orchestrators like AutoGen Wu et al. (2023) and HuggingGPT Shen et al. (2024) all route communication through natural language. This is lossless when both parties are language models, but compresses the policy and value representations of pretrained neural agents into a few tokens. A parallel line moves reasoning itself into continuous representations to escape the bandwidth limit of discrete tokens Zhu et al. (2025), either through latent recurrence within one model (CoCoNut Hao et al. (2025)) or by interleaving latent and text tokens in a single reasoning stream (Token Assorted Su et al. (2025), latent tokens as extra computation Sun et al. (2025)); these operate on a model’s own hidden state, not a separate agent’s. Latent channels exist in multi-agent RL Sukhbaatar et al. (2016); Das et al. (2019); Foerster et al. (2016) and homogeneous LLMs Wang and others (2025), but assume jointly trained or homogeneous populations, not a frozen LLM with a frozen specialist. Cross-modal injection (PaLM-E Driess et al. (2023), RT-2 Brohan et al. (2023)) projects raw sensory observations, not a pretrained agent’s processed policy/value representations. None of these internalizes a non-language specialist’s latent state into an LLM (Table 26). On the domain side, neural chess engines encode rich positional structure in their activations Silver et al. (2017); Monroe and Chalmers (2024), as interpretability work confirms Jenner et al. (2024), yet prior LLM-chess work either trains task-specific models Jhamtani et al. (2018) or conditions on verbalized outputs Feng et al. (2024); none exposes the engine’s latent state to the LLM.
5 Conclusion
We introduced latent state internalization, which replaces verbalized LLM–agent communication with direct projection of the agent’s continuous representations into the LLM’s embedding space. LLAMIA-14B, trained via projector alignment followed by DAPO, matches or exceeds dedicated task finetunes across LLAMIA-Bench. The verbalization debt widens with interaction depth and on signals that resist text serialization (e.g., puzzle interest), and does not close with LLM scale or RL budget in our evaluated range, indicating verbalization is a structural bottleneck.
References
- Create custom subagents. Note: https://code.claude.com/docs/en/sub-agentsClaude Code Documentation. Accessed: 2026-05-02 Cited by: §1.
- RT-2: vision-language-action models transfer web knowledge to robotic control. External Links: 2307.15818, Link Cited by: Table 26, §4.
- Rt-1: robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817. Cited by: §1.
- Caissabase: a free chess database. Note: https://mattplayschess.com/free-large-db/Accessed: 2025 Cited by: Appendix B.
- Deep blue. Artificial intelligence 134 (1-2), pp. 57–83. Cited by: §1.
- Tarmac: targeted multi-agent communication. In International Conference on machine learning, pp. 1538–1546. Cited by: §4.
- PaLM-e: an embodied multimodal language model. External Links: 2303.03378, Link Cited by: Table 26, §4.
- Chessgpt: bridging policy learning and language modeling. Advances in Neural Information Processing Systems 36. Cited by: §C.5, §1, §4.
- Learning to communicate with deep multi-agent reinforcement learning. Advances in neural information processing systems 29. Cited by: §4.
- GameKnot: online chess. Note: https://gameknot.com/Accessed: 2024 Cited by: §I.2.
- Example of the glicko-2 system. Boston University 28, pp. 2012. Cited by: §B.2.3.
- Training large language models to reason in a continuous latent space. External Links: 2412.06769, Link Cited by: Table 26, §4.
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework. arXiv. Note: arXiv:2308.00352 [cs] External Links: Link, Document Cited by: §1.
- Evidence of learned look-ahead in a chess-playing neural network. arXiv preprint arXiv:2406.00877. Cited by: Appendix B, §E.5, §1, §2.1, §2.2, §2.2, §4.
- Learning to generate move-by-move commentary for chess games from large-scale social forum data. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), I. Gurevych and Y. Miyao (Eds.), Melbourne, Australia, pp. 1661–1671. External Links: Link, Document Cited by: §B.2.4, Table 10, Appendix B, Table 15, §1, §3.1, §4.
- Bridging the gap between expert and language models: concept-guided chess commentary generation and evaluation. External Links: 2410.20811, Link Cited by: Table 16, §1.
- LLM chess: benchmarking reasoning and instruction-following in llms through chess. External Links: 2512.01992, Link Cited by: §1.
- Improving chess commentaries by combining language models with symbolic reasoning engines. arXiv preprint arXiv:2212.08195. Cited by: §B.2.4, §1, §1.
- Lichess open database: puzzles. Note: https://database.lichess.org/Accessed: 2025 Cited by: Appendix B.
- Lichess open database: puzzles. Note: https://database.lichess.org/#puzzlesAccessed: 2025 Cited by: Table 10, Appendix B, §3.1.
- Tracing the thought of a grandmaster-level chess-playing transformer. arXiv preprint arXiv:2604.10158. Cited by: §E.5, §2.2.
- Visual instruction tuning. arXiv preprint arXiv:2304.08485. Cited by: §2.2.
- G-eval: nlg evaluation using gpt-4 with better human alignment. External Links: 2303.16634, Link Cited by: §B.2.4, §H.4.
- Aligning superhuman ai with human behavior: chess as a model system. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pp. 1677–1687. Cited by: §B.2.2, §B.2.2, §B.2.2, §B.3, Table 10, Appendix B, §D.1, Table 14, Table 14, Table 14, §3.1.
- Predicting chess puzzle difficulty with transformers. In 2024 IEEE International Conference on Big Data (BigData), pp. 8377–8384. Cited by: Table 13.
- Mastering chess with a transformer model. arXiv preprint arXiv:2409.12272. Cited by: §A.5, §2.2, §4.
- Subagents – Codex. Note: https://developers.openai.com/codex/subagentsOpenAI Developer Documentation. Accessed: 2026-05-02 Cited by: §1.
- Effective generative ai: the human-algorithm centaur. Harvard Data Science Review (Special Issue 5). Cited by: §H.5.
- Toolformer: language models can teach themselves to use tools. External Links: 2302.04761, Link Cited by: §1, §4.
- Hugginggpt: solving ai tasks with chatgpt and its friends in hugging face. Advances in Neural Information Processing Systems 36. Cited by: §4.
- Mastering the game of go without human knowledge. nature 550 (7676), pp. 354–359. Cited by: §1, §4.
- Token assorted: mixing latent and text tokens for improved language model reasoning. arXiv preprint arXiv:2502.03275. Cited by: Table 26, §4.
- Learning multiagent communication with backpropagation. External Links: 1605.07736, Link Cited by: §4.
- Enhancing latent computation in transformers with latent tokens. arXiv preprint arXiv:2505.12629. Cited by: Table 26, §4.
- Multi-Agent Collaboration Mechanisms: A Survey of LLMs. arXiv. Note: arXiv:2501.06322 [cs] External Links: Link, Document Cited by: §1.
- LatentMAS: pure latent collaboration for multi-agent systems. arXiv preprint arXiv:2511.20639. Cited by: Table 26, §4.
- Accelerating self-play learning in Go. arXiv preprint arXiv:1902.10565. Cited by: 1st item, §3.8.
- New human-like play and analysis (KataGo human SL network). Note: KataGo v1.15.0 release, https://github.com/lightvector/KataGo/releases/tag/v1.15.0 Cited by: Appendix F, §3.8.
- Autogen: enabling next-gen llm applications via multi-agent conversation framework. arXiv preprint arXiv:2308.08155. Cited by: §4.
- Qwen3 technical report. arXiv preprint arXiv:2505.09388. Cited by: Table 5, §C.1, §2.2.
- React: synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629. Cited by: §C.1, §1, §4.
- Dapo: an open-source llm reinforcement learning system at scale. arXiv preprint arXiv:2503.14476. Cited by: §A.4.2, §2.3.2, §2.3.2.
- Automated chess commentator powered by neural chess engine. arXiv preprint arXiv:1909.10413. Cited by: Appendix B, Table 13, Table 15, Table 18, §1.
- Human-aligned chess with a bit of search. External Links: 2410.03893, Link Cited by: Table 13, Table 13, Table 14, Table 14, Table 14, Table 18.
- Chain of Agents: Large Language Models Collaborating on Long-Context Tasks. arXiv. Note: arXiv:2406.02818 [cs] External Links: Link, Document Cited by: §1.
- A survey on latent reasoning. External Links: 2507.06203, Link Cited by: Table 26, §1, §4.
Appendix Table of Contents
Appendix A Implementation Details
A.1 Training and Inference Cost
We report training and inference cost for LLAMIA and LLAMIA-Verb, benchmarked on A100-80GB (4 nodes 8 = 32 GPUs).
Training (GPU-hours).
The LatentBridge and Stage-1 alignment add a small fixed overhead (roughly 1–2 GPU-hours at 14B), so total training cost stays within of the verbalized pipeline at every scale (Table 2).
| Backbone | LLAMIA | LLAMIA-Verb |
| 4B | 9.1 | 8.8 |
| 8B | 16.8 | 15.0 |
| 14B | 22.6 | 21.4 |
Inference (per query).
A verbalized call returns tokens; a LLAMIA call adds a fixed latent tokens ( total). However, LLAMIA learns to invoke the subagent less frequently, so the extra per-invocation cost is offset by fewer invocations, yielding comparable or lower average tokens-per-query and wall-clock latency (Table 3). The two interfaces invoke the subagent at different rates because they learn different collaboration strategies during DAPO (Section 3.6). Latent internalization thus does not increase average inference cost.
Backbone Interface Calls Tok/Call Tok/Query Latency 4B Verb 3.4 150 510 1.4 s LLAMIA 2.3 182 419 1.0 s 8B Verb 3.2 150 480 1.8 s LLAMIA 2.1 182 382 1.2 s 14B Verb 2.9 150 435 2.2 s LLAMIA 1.9 182 346 1.4 s
A.2 Libraries and Software
Package Version Role PyTorch 2.8 Training backend MegatronLM 0.15.0 Tensor-parallel training VERL 0.7.1 RL training framework vLLM 0.10.2 Rollout inference engine Ray 2.55.1 Distributed orchestration Transformers 4.56.2 Model loading & tokenization lc0 0.32.1 Chess specialist engine (UCI)
A.3 Model Architecture and LatentBridge
A.3.1 Backbone LLM
Model Params Hidden Layers Heads (KV) Context Qwen3-4B 4B 2,560 36 32 (8) 32,768 Qwen3-8B 8B 4,096 36 32 (8) 32,768 Qwen3-14B 14B 5,120 40 40 (8) 32,768
A.3.2 LatentBridge Projector
| Parameter | Chess (Lc0-BT4) |
| Input dimension | 1024 |
| Output tokens | 32 |
| MLP layers | 3 |
| Activation | GeLU |
| Source layer | 14 of 15 (penultimate) |
A.4 Training Configuration
Model parameters and optimizer state are partitioned across 4 nodes of 8xA100 GPUs using PyTorch FSDP, coordinated via Ray. The rollout vLLM instance and lc0 server fleet co-reside on the same GPUs, with 3 GiB per GPU reserved for the BT4 network.
A.4.1 Stage 1: Projector Alignment
Stage 1 trains alone with frozen on 5 M state–policy pairs drawn from lc0’s forward pass over the Lichess evaluation database. Each example pairs a FEN with lc0’s top moves and pawn evaluations; the model minimizes cross-entropy over verbalized policy output conditioned on and a fixed prompt template (Appendix J). Skipping Stage 1 causes training instability during the first 40% of Stage 2 (Appendix E).
| Hyperparameter | Value |
| LR () | 2e-4 |
| Batch size | 256 |
| Steps | 2 epochs (39K steps) |
| (state tokens) | 32 |
| Max sequence length | 32,768 |
A.4.2 Stage 2: Reinforcement Learning (DAPO)
Stage 2 fine-tunes both and jointly with DAPO Yu et al. (2025). The reward is a scalar outcome signal defined per task (Appendix B.1).
| Hyperparameter | Value |
| LR () | 1e-5 |
| LR () | 1e-6 |
| Batch size | 128 |
| Steps | 3,000 |
| KL coefficient | 0.01 |
| Clip | 0.2 / 0.28 |
| Group size | 8 |
| (state tokens) | 32 |
| Max sequence length | 32,768 |
| Rollout temperature | 1.0 |
| Eval temperature | 0.0 |
A.5 Chess Specialist: Lc0-BT4
With ScoreType=WDL_mu, lc0 reports scores as where is the WDL-mean utility from the side to move’s perspective Monroe and Chalmers (2024).
A.5.1 Engine Configuration
Network.
BT4-1024x15x32h-swa-614750011 1 https://storage.lczero.org/files/networks-contrib/BT4-1024x15x32h-swa-6147500-policytune-332.pb.gz: a 1024-dim, 15-layer, 32-head BT4 transformer saved at step 6,147,500 after SWA and policy-tuning. GPU footprint: GB in fp16 on A100.
Parameter Upstream default Our value Rationale Backend cuda-auto cuda-fp16 Explicit fp16 on Ampere WeightsFile <autodiscover> BT4 path Pinned network Threads 0 (backend) 2 Two CPU workers per GPU MinibatchSize 0 (backend) 128 Throughput sweet spot on A100 NNCacheSize 2,000,000 200,000 Cap host RAM VerboseMoveStats false true Required for /policy parser PolicyTemperature 1.36 1.36 Policy softmax (upstream default retained) ScoreType WDL_mu WDL_mu Score =
Appendix B Datasets & Benchmarks
This section describes the datasets and evaluation protocol for LLAMIA-Bench. We find four major themes in how AI systems collaborate with domain-expert agents: Behavioral imitation: reproducing human play at a target skill level, the problem behind bots like Play Magnus22 2 https://www.playmagnus.com and Maia (McIlroy-Young et al., 2020). State assessment: predicting human-aligned properties of game states such as difficulty and engagement, the core task in Lichess’s puzzle rating system and Chess.com’s adaptive training (Lichess.org, 2024b). Comparative: explaining why a position favors one side—identifying material or structural advantages and disadvantages—as required in single-position analysis (Jhamtani et al., 2018) and game-level commentary. Rationale: generating natural-language explanations for why a player made a specific move, from single-move annotation (Jhamtani et al., 2018; Zang et al., 2019) to the game-length narratives produced by channels like Agadmator and GothamChess. These four themes place progressively harder demands on the communication channel between LLM and subagent: from a single-position state query (behavioral imitation) to game-length narrative integration (commentary), and from signals with partial textual correlates like move quality to ones without, such as aesthetic interest (Section B.2.3). LLAMIA-Bench instantiates each as an evaluation task. Chess serves as the testbed because it offers all four at once: subagents whose internal representations are mapped by interpretability work (Jenner et al., 2024), public game databases at scale (Lichess.org, 2024a; Caissabase Contributors, 2024), established benchmarks with dedicated task finetune baselines, and decades of human–engine collaboration.
B.1 Dataset & Metrics
Table 10 summarizes dataset provenance; Table 11 lists evaluation metrics per task. Splits marked OOD are out-of-distribution: the test distribution is absent or shifted relative to Stage-2 training, so generalization must come from internalized representations rather than memorization. Where the original authors provide a fixed test split we use it; otherwise we sample a random held-out split. Detailed descriptions of each task follow in §B.2.
| Task / Source | Description | Train / Test |
| Behavior Cloning (§B.2.2) MAIA-KDD, Lichess McIlroy-Young et al. (2020) | Predict the move a human of a given Elo would play; 5 Elo buckets (1100–1900) | 12M / authors’ |
| In the wild (OOD) MAIA-KDD, Lichess (Ours) | Three OOD splits: GM-25 (top-25 GMs), Low-Time (clock pressure 10%), Elo Gap (500 pts) | 167K / — |
| Puzzle Understanding (§B.2.3) Lichess Puzzles Lichess.org (2024b) | Predict difficulty (Glicko-2) and interest (community votes) of tactical puzzles | 4M / 5K |
| Move Annotation (§B.2.4) Lichess Jhamtani et al. (2018) | Generate natural-language explanation for a single move across 5 semantic categories | 90K / authors’ |
| Game Commentary (§B.2.1) Agadmator YouTube (Ours) | Produce coherent multi-turn narrative spanning an entire game | 1.9K / 100 |
| Task | Metric(s) | RL Reward |
| Behavior Cloning | Move-match accuracy | Top-3 rank |
| Puzzle Understanding | Spearman (§B.2.3) | Normalized MAE |
| Difficulty | Spearman (§B.2.3) | Normalized MAE |
| Interest | Spearman (§B.2.3) | Normalized MAE |
| Solved (%) | Exact solution-line accuracy (§B.2.3); parity metric, not primary | — |
| Move Annotation | G-eval ; BLEU-2 (§B.2.4) | G-eval |
| Game Commentary | G-eval ; BLEU-2 (§B.2.1) | G-eval |
B.2 Detailed Task Descriptions
B.2.1 Game-Level Commentary
Game-level commentary requires a coherent, multi-turn narrative spanning an entire game—unlike move-level annotation, errors compound across the narrative, and the model must track evolving themes (initiative shifts, pawn-structure transformations, time trouble). We introduce Agadmator-2K, the first large-scale dataset for this task: 1,900 narrated games from Agadmator’s YouTube channel,33 3 https://www.youtube.com/@agadmator totaling approximately 500 hours.
Dataset construction.
Move-segmented commentary is unavailable from YouTube. We construct it in four steps: (i) transcripts are extracted via Whisper-v3-large; (ii) video timestamps are aligned to PGN move sequences using a sliding-window move-tracking buffer; (iii) GPT-4o labels which moves each transcript segment references, guided by the known PGN; (iv) segments are accepted only when the inferred move order matches the PGN exactly, discarding retries and out-of-order narration. The test set consists of 100 games, held out by ascending view count to minimize overlap with LLM pretraining corpora. This is a heuristic proxy for low contamination, not a guarantee; we discuss contamination further in §B.3.
Metrics.
We use the same G-eval framework as move annotation (§B.2.4), adapted for game-level commentary: each generated segment is scored on relevance, completeness, clarity, and fluency, with the judge grounded by Lc0-BT4 engine lines and the ground-truth transcript. We also use the BLEU-2 metric. The RL reward is G-eval.
B.2.2 Behavior Cloning
Behavior cloning measures whether LLAMIA can imitate human play conditioned on skill level or player identity. The evaluation metric across all splits is move-match accuracy: the fraction of positions where the model’s top-1 predicted move exactly matches the target player’s move. We follow the MAIA evaluation protocol (McIlroy-Young et al., 2020): Maia variants use best-of- sampling; LLAMIA uses a single forward pass.
The RL reward uses a softer signal: the top-3 rank of the target move in the model’s output distribution, normalized to . Top-1 exact match as a reward collapsed training—the signal was too sparse for most positions, yielding near-zero gradients throughout Stage 2. Rank within the top-3 provides a dense, monotone reward that penalizes misranking without requiring exact prediction, while remaining consistent with the evaluation objective.
Maia Benchmark (Elo Buckets).
We evaluate on the MAIA-KDD held-out test set (McIlroy-Young et al., 2020), stratified into five Elo buckets: 1100, 1300, 1500, 1700, and 1900. The test set is player–game disjoint from all training data. Each bucket is treated as an independent task; the aggregate BC score reported in the main paper is the unweighted average across buckets.
GM-25 (OOD).
GM-25 targets the top-25 rated grandmasters in FIDE history by peak rating.44 4 https://en.wikipedia.org/wiki/List_of_chess_players_by_peak_FIDE_rating Each grandmaster is a separate behavioral target. The largest available per-GM corpus is 4,641 games (Viktor Korchnoi),55 5 https://www.365chess.com/top-chess-players-games.php less than 3% of the data that per-GM Maia models require (McIlroy-Young et al., 2020). No Stage-2 training data is drawn from these GM corpora; generalization must come from internalized representations and cross-Elo behavioral transfer.
Low-Time (OOD).
Under severe clock pressure, players shift strategy regardless of position quality. We extract positions where either player’s remaining clock is below 10% of the initial time control, or where cumulative time usage differs by more than 50% between the two sides. Positions are stratified by game phase (opening, middlegame, endgame) and sampled equally across time controls, yielding 129,000 positions. Clock-context metadata is absent from Stage-2 training, making this an OOD split: the model must infer time-pressure effects from the position and move alone.
Elo Gap (OOD).
Players adapt their style when facing a large skill gap—weaker players take more risks, stronger players simplify. We filter Lichess Rapid and Classical games where the Elo difference exceeds 500 points, yielding 34,000 games (68,000 player-side instances). Extreme skill-gap matchups are rare in the Stage-2 training distribution; conditioning on opponent strength must emerge from contextual signals rather than memorization.
B.2.3 Puzzle Understanding
Puzzle understanding probes whether LLAMIA has internalized the subagent’s positional representations well enough to predict human-aligned properties of game states. Both sub-tasks draw from the same 4-million-puzzle Lichess corpus,66 6 https://database.lichess.org/lichess_db_puzzle.csv.zst which provides community-derived ground-truth labels for difficulty and engagement. We hold out a shared test set of 5,000 puzzles, stratified by difficulty (Glicko-2 quintiles), theme (tactical motif), and interest (score quintiles) to ensure uniform coverage across the label space. Evaluation uses Spearman between predicted and ground-truth values; the RL reward is a normalized mean-absolute-error penalty.
Difficulty Estimation.
Puzzle difficulty is operationalized via a Glicko-2 rating system (Glickman, 2012): each human solving attempt is treated as a rated match between solver and puzzle, and the Glicko-2 rating accumulated over all attempts serves as ground truth. The model receives the puzzle position and solution line, and predicts difficulty on a normalized scale. Evaluation uses Spearman between predicted values and ground-truth Glicko-2 ratings on the stratified 5,000-puzzle test split.
Interest Estimation.
Lichess assigns each puzzle an interestingness score (range: to ) computed from community upvotes and downvotes, weighted by solver performance. This signal has no straightforward textual correlate: a puzzle’s aesthetic appeal depends on motif rarity, surprise, and solution elegance—features encoded in the subagent’s positional representation but absent from any verbalized move list. The model predicts interest from the same input as difficulty; evaluation uses Spearman on the same stratified 5,000-puzzle test split. Interest estimation is the diagnostic task on LLAMIA-Bench: the non-verbalizable nature of the target signal means that all text-mediated systems collapse on this task (Section 3.4).
Puzzle Solving Accuracy.
We also report Solved (%): the fraction of test puzzles for which the model produces the complete correct solution line—every forced move in sequence—using policy-only decoding (single forward pass per position, no search). The model receives the initial puzzle FEN and outputs moves one at a time; a puzzle is marked solved only if all moves in the ground-truth solution are produced in the correct order.
This metric is excluded for GPT-5.1 + Lc0 (marked — in Table 17) because verbalized engine access makes it uninterpretable: a system that queries Lc0 at each puzzle position and forwards the top-ranked move would score near-perfect not by reasoning about the position but by delegating each step to the engine. The metric is informative only when the model must solve the puzzle from its own internalized representations without live tool queries. All other systems in Table 17 use policy-only decoding for this column.
Puzzle solving accuracy functions as a parity metric on LLAMIA-Bench: all systems with Lc0 access cluster in the 84–94% range, and LLAMIA’s improvement over LLAMIA-Verb is modest (2–3 pp). The metric confirms that engine-access systems are not deficient tactically; the differentiation between LLAMIA and LLAMIA-Verb arises in difficulty and interest prediction, not in puzzle-solving throughput.
B.2.4 Move Annotation
Move annotation evaluates LLAMIA’s ability to generate natural-language explanations for individual moves, conditioned on the board state and the move played. We follow the benchmark of Jhamtani et al. (2018): 90,000 Lichess games annotated in English, with annotations categorized into five semantic dimensions that span both explanation themes from §B. The rationale theme is instantiated by three dimensions—description (what the move does), quality (blunder, inaccuracy, good, best), and planning (lookahead and intent)—while the comparative theme is instantiated by two—context (positional advantages and disadvantages relative to prior or future moves) and comparative (alternative moves and why they were rejected). The standard benchmark provides the target move as input; we additionally evaluate zero-shot without this prior to test whether internalized representations can identify annotation-worthy moves.
Metrics.
BLEU-2 and perplexity (Lee et al., 2022) are evaluated per annotation category, following prior work. The primary metric is G-eval (Liu et al., 2023b): an LLM-as-judge framework in which GPT-4o rates each generated annotation on a 0–1 scale across four dimensions (relevance, accuracy, completeness, fluency). The judge receives the board FEN, the move in algebraic notation, and Lc0-BT4’s top-3 engine lines as grounding context, so its assessments are anchored in engine analysis rather than surface plausibility alone. Per-annotation G-eval scores are averaged across the four dimensions; the corpus-level score is the mean over all test annotations. G-eval also serves as the RL reward signal for this task.
B.3 Data Contamination Statement
All evaluation in LLAMIA-Bench is conditioned on board positions represented as FEN strings. We enforce a strict FEN-level disjointness guarantee: no FEN appearing in any test split co-occurs in any stage of training—projector pretraining (Stage 1), RL training (Stage 2), or the base LLM’s supervised fine-tuning data. Concretely, we collect the set of all FENs used across projector pretraining pairs and Stage-2 RL rollouts, and verify that the intersection with each test split is empty. For the Maia BC test set, this property is inherited from the player–game disjoint split of McIlroy-Young et al. (2020). For the puzzle understanding test split, the 5,000 held-out puzzles are sampled after removing all FENs present in the training pool. For Agadmator-2K, the 100 held-out games are additionally sorted by ascending view count as a heuristic to reduce overlap with LLM pretraining corpora, though we cannot verify disjointness with respect to closed-source pretraining data. For the OOD behavior-cloning splits (GM-25, Low-Time, Elo Gap), no Stage-2 training data is drawn from these distributions by construction; we further verify that no test FEN appears in the Stage-1 projector data.
Appendix C Baselines
C.1 Frontier LLMs with Verbalized Tools
To select the strongest frontier baseline, we evaluate five models—GPT-5.1 GPT-5.1, Claude Sonnet 4.5, Claude Opus 4.5, Gemini 2.5 Pro, and Qwen3-235B Yang et al. (2025)—each given access to Lc0-BT4 via a ReAct Yao et al. (2022) tool-calling loop. At each invocation the tool returns the top-5 moves with centipawn evaluations, win/draw/loss probabilities, and principal variations up to depth 20. All models share identical tool schemas, system prompts, and sampling parameters; the only variable is the LLM backbone. We sample 100 positions from each LLAMIA-Bench task and report the aggregate metric per task group.
Model BC Puzzle Annot. Comm. Avg. GPT-5.1 + Lc0 42.5 29.0 37.5 52.0 40.3 Claude Opus 4.5 + Lc0 39.6 27.4 36.3 50.3 38.4 Claude Sonnet 4.5 + Lc0 36.9 26.3 34.2 47.1 36.1 Gemini 2.5 Pro + Lc0 40.6 28.1 36.7 49.8 38.8 Qwen3-235B + Lc0 36.5 25.5 33.4 45.6 35.3
GPT-5.1 obtains the highest average across all task groups. We therefore use GPT-5.1 + Lc0 (Verb) as the frontier verbalized baseline throughout the paper.
C.2 LLAMIA-Verb
LLAMIA-Verb is the primary controlled ablation of LLAMIA. It receives the identical base model, training data, reward signals, and RL recipe (DAPO) as LLAMIA, but the subagent’s output is provided exclusively through verbalized tool responses: top- moves, centipawn evaluations, WDL probabilities, and principal variations rendered as text tokens. No LatentBridge projection is trained; the continuous latent tokens that LLAMIA receives are replaced by their textual equivalents.
We train LLAMIA-Verb at three scales—4B, 8B, and 14B—matching the corresponding LLAMIA checkpoints in base model architecture, training data, and total compute budget. This controlled setup isolates the contribution of latent state internalization from model family, data mix, reward shaping, and optimization, and directly tests the central claim that verbalization is a lossy bottleneck (Section 3).
C.3 LLAMIA (4B, 8B, 14B)
To test whether the verbalization debt is an artifact of scale rather than interface, we train LLAMIA at three scales—4B, 8B, and 14B—matching the corresponding LLAMIA-Verb checkpoints in base model architecture, training data, and total compute budget. All three LLAMIA checkpoints use the full latent-state internalization pipeline: LatentBridge projection of BT4 activations into continuous tokens, Stage 1 projector alignment, and Stage 2 end-to-end DAPO. Comparing LLAMIA-B against LLAMIA-Verb-B at each scale isolates the interface contribution (internalization vs. verbalization) independently of model capacity, and directly supports the claim that the performance gap is not closed by scaling the LLM (Section 3).
C.4 Dedicated task finetunes
For each LLAMIA-Bench task we compare against the strongest published or reproducible task-specific model. Table 13 lists the expert per task alongside its training paradigm and data scale. These models represent the performance ceiling achievable with task-specific architectures and, in several cases, substantially more training data than LLAMIA receives. Tasks marked — have no established prior expert; LLAMIA-Bench introduces them as new evaluation targets.
Task / Split Task Expert Method Data Scale Behavior Cloning Maia (Elo buckets) Allie Zhang et al. (2024a) SL+Search 93M games GM-25 / Low-Time / Elo Gap Allie Zhang et al. (2024a) SL+Search 93M games Puzzle Understanding Difficulty Estimation Miłosz and Kapusta (2024) SFT 4M Interest Estimation — — — Move Annotation Move Annotation SCC Zang et al. (2019) SL 90K games
C.5 Lightweight Probes, Templates, and Alternative Injections
To attribute LLAMIA’s gains, we add four controls beyond the dedicated task experts above; headline numbers appear in Table 1.
Probe on frozen BT4.
A -layer MLP (–, M params) is trained end-to-end on frozen BT4 activations (layer 14/15 residual stream), one head per task, for behavior cloning, difficulty, and interest. It measures how much the latent state gives up when decoded by a lightweight predictor rather than an LLM: it scores below even text-only GPT-5 on every task (BC-MAIA , Difficulty , Interest ), so the engine’s penultimate state is not directly decodable into these human-aligned targets.
Template baselines.
For move annotation and commentary we fill a fixed three-line template directly from the engine’s output—move quality (centipawn loss vs. the engine’s best), the engine’s preferred line and evaluation, and the best alternative:
-
Move quality: Nf3 loses 40cp vs. best (inaccuracy).
Best line: engine prefers Nc3 Nf6 d4, +0.6 (W/D/L 48/40/12).
Alternative: Bb5 slightly weaker, dropping to +0.3.
The deterministic template reaches only BLEU-2 / G-eval; rewriting it with GPT-5 improves marginally ( / ) and remains below the verbalized tool ( / ) and far below LLAMIA ( / ). Presenting engine statistics in more natural language is not the source of LLAMIA’s gains.
LLM-ChessCLIP (PaLM-E-style injection).
Following the representation-injection paradigm of PaLM-E, we replace the engine’s latent state with embeddings from ChessCLIP Feng et al. (2024) while keeping the injection mechanism fixed, isolating what is injected (a raw board encoder vs. a pretrained agent’s processed policy/value state). It recovers only a fraction of LLAMIA’s improvement, indicating the benefit is specific to the agent’s internal state, not any learned board representation.
LLM-Only.
The same backbone is post-trained (SFT + RL) on the identical chess and commentary data with no engine access, testing whether the expertise can be absorbed into weights. It trails even untrained tool use (per-scale numbers in Appendix D).
Appendix D Extended Results
The five tasks in LLAMIA-Bench probe different regimes of LLM–agent collaboration, varying in horizon, evaluation metrics, and the strength of task-specific baselines. Here we discuss the extended per-task evaluation of LLAMIA with more metrics, baselines, and detailed analysis of LLAMIA’s task specific behavior.
D.1 Behavior Cloning
Behavior Cloning asks the system to predict the move a human at a given skill level would play rather than the optimal move, given the position and the target Elo rating.
We report two splits. The first is the MAIA Test split: five Elo buckets (, , , , ) drawn from Lichess blitz, matching the protocol of McIlroy-Young et al. (2020) and the in-distribution setting for the published BC experts. The second is Wild, three out-of-distribution splits we introduce. GM-25: contains top-grandmaster games Low-Time: contains positions played with under 30 seconds remaining, players often play different move when they or opponents are under time pressure. Elo: contains games with large rating gaps between the two players leading to different attacking or defensive strategies.
Model #Train Maia Benchmark Elo Buckets LLAMIA-Bench (Wild) games 1100 1300 1500 1700 1900 Avg. GM-25 Low-Time Elo Avg. Task-Specific Expert Stockfish (d15) McIlroy-Young et al. (2020) – 36 38 39 40 41 39 53 27 32 37 Maia McIlroy-Young et al. (2020) 12 M/bkt 51 52 53 53 52 52 38 45 44 42 Allie-Policy Zhang et al. (2024a) 93 M 51 53 54 56 57 54 43 45 44 44 Allie-Adaptive-Search Zhang et al. (2024a) 93 M 52 54 56 57 58 55 44 45 45 45 Frontier Baselines (5-shot) GPT-5 (text only) – 24 27 29 30 29 28 22 20 24 22 GPT-5 + Lc0 – 42 44 45 47 46 45 40 39 41 40 Qwen3-14B + Lc0 – 36 38 40 41 39 39 34 31 35 33 Verbalized, SFT LLAMIA-Verb-4B 20 K 38 40 41 41 39 32 30 34 LLAMIA-Verb-8B 20 K 41 42 43 44 42 35 33 37 LLAMIA-Verb-14B 20 K 42 44 45 45 43 37 35 39 Verbalized, DAPO LLAMIA-Verb-4B 20 K 39 41 42 43 41 34 32 36 LLAMIA-Verb-8B 20 K 42 44 45 46 44 37 35 39 LLAMIA-Verb-14B 20 K 43 45 46 47 44 39 37 41 Latent, SFT LLAMIA-4B 20 K 46 48 49 50 47 40 38 44 LLAMIA-8B 20 K 48 50 51 52 49 43 41 47 LLAMIA-14B 20 K 50 51 52 53 51 45 43 49 Latent, DAPO LLAMIA-4B 20 K 48 50 51 52 49 46 45 50 LLAMIA-8B 20 K 50 51 53 54 51 51 49 54 LLAMIA-14B 20 K 51 53 53 55 52 57 54 60
D.2 Move Annotation & Game Commentary
Move Annotation and Game Commentary are language generation tasks that require the system to generate explanations or commentate on moves played by a player or game segment between two players. To explain these move sequences the system must understand the gameplay strategies. Both tasks require the system to understand the position and require counterfactual exploration to explain move choices calculated by the players.
D.2.1 Move Annotation
requires the agent to deduce the intent or rationale behind a player’s move made in a given position– what is this move trying to achieve?– given the context including their previously played moves, Elo (skill levels), and time remaining.
Model Planning Comparative Avg. Task-Specific FT SCC Zang et al. (2019) Frontier Baselines (5-shot) GPT-5 (text only) GPT-5 + Lc0 Qwen3-14B + Lc0 Verbalized, SFT LLAMIA-Verb-4B LLAMIA-Verb-8B LLAMIA-Verb-14B Verbalized, DAPO LLAMIA-Verb-4B LLAMIA-Verb-8B LLAMIA-Verb-14B Latent, SFT LLAMIA-4B LLAMIA-8B LLAMIA-14B Latent, DAPO LLAMIA-4B LLAMIA-8B LLAMIA-14B
D.2.2 Game commentary
requires the agent to produce a coherent natural-language narrative spanning an entire game ( moves), explaining strategic plans, critical turning points, and tactical sequences as they unfold given each position, the move sequence, and the players’ Elo. Unlike Move Annotation, which targets a single position, the system must integrate positional understanding with multi-step counterfactual reasoning to decide which moves merit elaboration and which are routine, and to maintain a coherent storyline across the game.
Model G-eval BLEU-2 Frontier Baselines (5-shot) GPT-5 (text only) GPT-5 + Lc0 Qwen3-14B + Lc0 Verbalized, SFT LLAMIA-Verb-4B LLAMIA-Verb-8B LLAMIA-Verb-14B Verbalized, DAPO LLAMIA-Verb-4B LLAMIA-Verb-8B LLAMIA-Verb-14B Latent, SFT LLAMIA-4B LLAMIA-8B LLAMIA-14B Latent, DAPO LLAMIA-4B LLAMIA-8B LLAMIA-14B
We train all baselines and LLAMIA on both tasks and evaluate by G-eval on relevance, completeness, clarity, and fluency (the same G-eval also acts as the DAPO reward, Section B.2.1) We also report BLEU-2 scores alongside G-eval for game level commentary, to make an LLM as a judge free metric.
Importance calibration in qualitative outputs.
Beyond the aggregate scores, the systems differ qualitatively in how they allocate explanation depth across moves. LLM-Only (Qwen3-14B+Lc0) and GPT-5 produce near-uniform-length commentary: every move receives 2–3 sentences regardless of whether it is a routine development move or a critical sacrifice. LLAMIA-Verb-DAPO partially corrects this by elaborating on large-eval-swing moves, but its pacing tracks eval magnitude rather than positional significance. It over-emphasises – eval drifts that human commentators ignore as “human moves” (the position is essentially unchanged in character even though the number moved), and it misses sacrifices and quiet winners that signal interesting moments without producing large eval changes. LLAMIA-DAPO modulates depth by latent-token change patterns directly: moves where the positional features (king-safety, pawn-structure, piece-coordination) shift discontinuously receive paragraph-scale analysis, while routine moves receive a single clause.
D.3 Puzzle Understanding
Puzzle Understanding evaluates whether a system can use the agent’s state to judge a tactical position rather than only select the engine move. We report two single-position ranking tasks. Difficulty asks the system to order puzzles by empirical Lichess solve difficulty. Interest asks it to order positions by the fraction of users who mark them interesting. Solving is included only as a parity check: engine-access systems solve most puzzles, so the discriminative metrics are the two Spearman correlations.
D.3.1 Difficulty and Interest prediction
Difficulty has a partial verbal proxy in solution length, which appears in the principal variation (PV). Interest has no comparable text proxy in the standard verbalized Lc0 output: it tests whether non-PV signals in the agent state help predict which positions humans mark as interesting. The matched LLAMIA-Verb and LLAMIA rows in Table 17 therefore test how much of the agent state survives verbalization at fixed position set, subagent, backbone scale, and optimization recipe.
Model Solved (%) Difficulty Interest Frontier Baselines (5-shot) GPT-5 (text only) – GPT-5 + Lc0 XX.X Qwen3-14B + Lc0 84.0 Verbalized, SFT LLAMIA-Verb-4B 84.0 LLAMIA-Verb-8B 85.5 LLAMIA-Verb-14B 86.5 Verbalized, DAPO LLAMIA-Verb-4B 86.0 LLAMIA-Verb-8B 87.5 LLAMIA-Verb-14B 88.5 Latent, SFT LLAMIA-4B 86.0 LLAMIA-8B 88.0 LLAMIA-14B 89.5 Latent, DAPO LLAMIA-4B 88.0 LLAMIA-8B 90.0 LLAMIA-14B 91.5
D.4 Full LLAMIA-Bench Results Table
System Behavior Cloning Puzzle Understanding Rationale Prediction Game Commentary MAIA Wild Difficulty Interest Task-Specific Expert Allie-Adaptive-Search Zhang et al. (2024a) 55 45 – – – – SCC Zang et al. (2019) – – – – – Frontier Baselines (5-shot) GPT-5 (text only) 28 22 GPT-5 + Lc0 45 40 Qwen3-14B + Lc0 39 33 Verbalized, DAPO LLAMIA-Verb-4B LLAMIA-Verb-8B LLAMIA-Verb-14B Latent, DAPO LLAMIA-4B LLAMIA-8B LLAMIA-14B
Verbalization is lossy, and the loss is task-specific.
At fixed backbone (Qwen3-14B), fixed subagent (Lc0-BT4), and fixed training recipe (DAPO), replacing the verbalized text channel with latent tokens lifts every column. The relative gain follows the landscape mechanism: largest on Interest (, a jump), where the discriminative signal lives in the policy distribution and the value gradients across candidate moves, none of which the verbalized output carries; large on Commentary (), where positional mechanisms compound across the game; moderate on Difficulty (), where solution length is a partial verbal proxy but leaves a residual signal over motif type and distractor sharpness; smaller on Rationale (), where PV-derived inference covers much of Planning; and smallest on Behavior Cloning ( on MAIA, on Wild), where the engine’s top- moves cover most of what humans actually play. The matched-recipe LLAMIA-Verb-4B,8B,14B rows isolate this ordering from confounds: the only variable that changes across the LLAMIA-VerbLLAMIA boundary is the integration interface.
Internalization is perceptual; agency is multi-step.
The gain from latent tokens splits into two components that surface in different task families. On single-step prediction (Interest, Difficulty, BC), the latent advantage is mostly perceptual: latent SFT alone closes most of the verb–latent gap, and DAPO adds little on top. The per-task tables make this concrete: LLAMIA-SFT-14B already reaches on Interest (vs. verbal DAPO at ), and adding RL lifts it only to . On multi-step tasks (Annotation, Commentary), DAPO becomes load-bearing because RL discovers collaboration strategies that are structurally unproductive under verbalization: counterfactual queries (play the alternative move, re-invoke the subagent on the resulting position, compare the two latent states), feature reading (attend to specific king-safety or pawn-structure activations to ground an explanation), and narrative pacing across 30+ moves.
Appendix E Ablations
This section isolates individual components of the LLAMIA pipeline. All ablations use Qwen3-14B as the backbone and Lc0-BT4 as the subagent unless stated otherwise. Metrics are averaged across all seven LLAMIA-Bench tasks unless a specific task is noted.
E.1 Emergent Collaboration Agency
Verbalization gives the model answers: “best move: e4, eval: +0.3.” Internalization gives the model perception: a 32-token encoding of the engine’s full representational state. Under verbalization, the model’s agency is over the questions—when to invoke, whether to follow. Under internalization, the agency extends to the reading—what to attend to in the latent state, how to interpret it for the current task, how to compose perceptions across invocations.
Figure 4 decomposes this difference. Five collaboration patterns are identified—engine-follow (adopt the top recommendation), consult-then-override (query then diverge), counterfactual query (play a hypothetical move, invoke, undo), multi-step lookahead (chain 2–3 counterfactual sequences), and abstention (act from language knowledge alone)—and the figure shows how each system allocates across them per task. The internalized model reads the engine differently depending on purpose; the verbalized model treats it as an answering machine. Figure 6 traces how both metrics evolve during training.
LLAMIA and LLAMIA-Verb have identical harness
LLAMIA and LLAMIA-Verb share the same system prompt, tool catalogue, reward function, backbone, and DAPO hyperparameters (Section J.2); no term in the reward and no curriculum stage targets counterfactual querying, lookahead, or any other pattern. The only instruction present in both prompts is to make selective, strategic use of the expensive get_policy call. The divergence in Figure 4 is therefore attributable to the integration interface, not to prompting or reward shaping.
E.2 Agent Size and Playing Strength
This ablation asks whether internalization gains are tied to BT4 specifically or generalize across agent architectures and capacities. We draw the agent pool from five additional Lc0 networks spanning roughly 2000–2600 Elo, covering three architectural families—convolutional SE-ResNets (T72, T78), standard transformers (T80, T82), and big transformers (BT3)—so that architecture is separated from raw capacity. All Elos are measured without search (single forward pass, policy-only decoding) via the gauntlet protocol : LLAMIA internalizes the agent’s single-forward-pass representation through the LatentBridge, so the no-search rating reflects the information actually available to internalization—tree-search budget is not distilled into the latent state. Table 19 lists the networks, with BT4 included as the primary agent for reference.
Network Architecture Params Elo (no search) T72 SE-ResNet, 25620 40M 2,010 T78 SE-ResNet, 38420 95M 2,180 T80 Transformer, 7681524h 109M 2,250 T82 Transformer, 7681524h 109M 2,292 BT3 Transformer, 7681524h 160M 2,510 BT4 Transformer, 10241532h 240M 2,810
We replace BT4 (240M params, 2810 Elo without search, 3300 with 1000-node MCTS) with progressively weaker networks from this pool. The projector is retrained from scratch for each agent; the RL recipe is identical.
Agent Elo (no search) BC Comm. Avg. T72 2000 44 52 38.7 T80 2200 49 62 46.3 T82 2292 51 66 48.9 BT3 2500 53 71 51.7 BT4 2810 56 75 56.9
Stronger agents monotonically improve LLAMIA across the representative tasks and the overall average, indicating that the projected representation preserves capability-relevant information. The trend does not rely on BT4 alone: T72 and T80 follow the same ordering on behavior cloning and commentary under the identical training recipe.
E.2.1 LLM Backbone and Agent Strength
Figure 7 extends the agent-size ablation to two LLM backbones (4B and 14B) across the same five agents, now spanning both SE-ResNet and Transformer architectures (Table 19). Performance increases monotonically with both LLM capacity and agent strength on all representative tasks.
The 4B-to-14B improvement is 2.4–3.1 larger for Transformer-architecture agents (T80, BT3, BT4) than for SE-ResNet agents (T72, T78). The ratio is largest on Interest () and Commentary (), the two tasks most dependent on reading the agent’s internal representation, and smallest on Behavior Cloning (). On Interest, the 4B scores for the strongest SE-ResNet agent (T78, 40) and the weakest Transformer agent (T80, 41) are nearly identical, yet their 14B scores diverge sharply (46 vs. 55). The Transformer representation carries signal that a 4B backbone cannot exploit but a 14B can. Three confounds prevent a causal claim: (i) architecture (attention-based representations may align with LLM attention more naturally), (ii) scale (Transformer agents in our set are also larger), and (iii) projector compatibility (the MLP LatentBridge may be better suited to projecting transformer features). Controlled experiments that vary architecture at matched parameter count are future work, but the pattern raises a practical question: does subagent architecture matter for internalization beyond raw agent strength?
E.3 Projection Token Count
We vary the number of state tokens injected per <invoke> call. Each configuration retrains both the projector and the RL policy from scratch. Increasing provides more bandwidth for the projector to encode the agent’s state but adds proportionally to the LLM’s context length per invocation.
BC Comm. Avg. Tokens/episode 4 48 58 45.1 680 8 51 65 49.7 720 16 54 73 54.9 790 32 56 75 56.9 920 64 56 75 56.7 1180
Performance increases monotonically from to and changes little at . The largest gains occur between and , suggesting that most useful signal is captured by the lower-bandwidth settings. We use throughout the paper, as it achieves the highest average score with lower context overhead than .
E.4 Interface Ablations: Latent-only and Shuffled Tokens
This section expands the interface controls summarized in Table 1. All rows use the B backbone, Lc0-BT4 subagent, and the DAPO recipe; only the integration interface changes.
Latent-only.
We retrain LLAMIA with the latent tokens only, removing the verbalized output, so the LLM sees only the latent tokens. Table 22 reports all six tasks. Latent-only nearly matches the full system everywhere; the small residual gap is largest on behavior cloning, consistent with the verbalized text supplying the top- move surface the latent state already encodes. Without the returned move, the LLM occasionally loses board tracking, which is why we retain the verbalized output.
System BC-MAIA BC-Wild Diff. Int. Rat. Comm. LLAMIA-Verb Latent-only LLAMIA (text+lat.)
Shuffled latent tokens.
To test whether the gain is merely extra embedding capacity, we retrain LLAMIA with shuffle- noise: of the latent tokens are swapped with the same-index tokens from random data points. Table 23 shows that shuffling more tokens degrades performance monotonically toward LLAMIA-Verb even though the model still receives embeddings, so added capacity and sequence length do not explain the gains. Degradation is fastest on Interest (no verbal proxy) and slowest on behavior cloning (top- proxy already in the text).
System Int. Comm. Diff. BC-MAIA LLAMIA (shuffle-0) shuffle-4 shuffle-8 LLAMIA-Verb
E.5 Layer Selection
LatentBridge reads the penultimate block (layer 14 of 15) of Lc0-BT4. We chose this empirically via the Stage-1 alignment loss: Stage 1 trains only LatentBridge to predict the engine’s move from the projected state while the LLM stays frozen, so its held-out cross-entropy measures how much decodable structure a layer exposes without the expensive Stage-2 RL run. Ablating every layer, layer 14 gave the lowest loss. To characterize this directly, we froze BT4 and trained lightweight linear probes (bilinear for moves) on the activations at every block, reading out four targets that stand in for our harder tasks: the played move, the best move two plies ahead, puzzle difficulty, and tactical-motif presence (Table 24). Blocks 12–14 are within noise of each other and jointly best, validating the Stage-1 choice; this matches the only interpretability study on this exact BT4 network, which locates value, source/target-square, and look-ahead-to-action features in block 14 Lin et al. (2026), and the late-block look-ahead structure reported for earlier Lc0 networks Jenner et al. (2024).
Layer BC move Look-ahead 2-ply Diff. Tactical 1 2 10 0.02 50 2 3 14 0.03 52 3 4 20 0.04 55 4 5 28 0.06 58 5 7 37 0.07 61 6 8 47 0.08 63 7 10 57 0.10 65 8 11 66 0.11 67 9 12 74 0.12 69 10 13 82 0.13 70 11 13 88 0.14 71 12 14 92 0.14 72 13 14 91 0.15 72 14 (ours) 14 89 0.15 71 15 (heads) 13 82 0.13 66
Appendix F Generalization to Go
To test whether latent state internalization transfers beyond chess, we instantiate LLAMIA on Go, keeping the recipe fixed and changing only what the specialist and the task require.
Setup.
- •
Frozen specialist: KataGo b18c384nbt Wu (2019), used frozen exactly as Lc0-BT4 is in chess.
- •
Extraction point: the shared trunk output—the activation map after KataGo’s final trunk normalization, immediately before the policy, value, and ownership heads.
- •
LatentBridge: the same three-layer adapter; we treat KataGo’s board intersections as spatial tokens, and only the first-layer input width changes to match KataGo’s trunk channels.
- •
Training: the identical two-stage recipe—Stage-1 projector alignment on (state, KataGo-move) pairs, then Stage-2 DAPO for behavior cloning.
Task and baselines.
We instantiate the direct Go analog of Behavior Cloning-Maia: predicting the move a human of a given rank plays, not the strongest move. The rank-matched reference expert is KataGo-HumanSL Wu (2024), a single net conditioned on KGS rank (the Go analog of Maia). The verbalized control (LLAMIA-Verb-Go) exposes KataGo’s top moves as text.
System (Go BC) rank 5k rank 5d Qwen3-4B (text only, no engine) 8 12 Qwen3-8B (text only, no engine) 11 16 Qwen3-14B (text only, no engine) 13 19 LLAMIA-Verb-Go-4B 31 34 LLAMIA-Verb-Go-8B 34 37 LLAMIA-Verb-Go-14B 36 39 LLAMIA-Go-4B (latent, ours) 40 42 LLAMIA-Go-8B (latent, ours) 43 46 LLAMIA-Go-14B (latent, ours) 48 50 KataGo-HumanSL (rank-calibrated) 46 50
As in chess, LLAMIA-Go leads LLAMIA-Verb-Go at 4B, 8B, and 14B, and LLAMIA-Go-8B already surpasses the verbalized 14B system, indicating the advantage comes from the latent state rather than backbone scale. These initial results suggest the recipe transfers beyond chess; extending to additional Go tasks that mirror the collaborative chess tasks is future work.
Appendix G Positioning vs. Latent-Space Work
Table 26 expands the Related Work discussion. The closest prior work either studies communication among homogeneous language models (LLM-to-LLM) or converts a non-language agent’s output back into text before the LLM consumes it. LLAMIA differs in internalizing a heterogeneous, non-language agent’s processed latent state directly into the LLM’s reasoning trace.
Prior work Communication medium Latent link to non-lang. agent? Latent reasoning survey Zhu et al. (2025) Latent reasoning within one model’s hidden state × CoCoNut Hao et al. (2025) Latent recurrence (same LLM) × Token Assorted Su et al. (2025) Latent tokens interleaved with text (same LLM) × Latent Tokens Sun et al. (2025) Extra latent tokens (same LLM) × LatentMAS Wang and others (2025) Latent hidden-state comms among homogeneous LLMs × PaLM-E Driess et al. (2023), RT-2 Brohan et al. (2023) Raw-observation embeddings × LLAMIA (ours) Pretrained agent’s latent state internalized into the LLM trace
Appendix H Human Evaluation Studies
Automated metrics measure textual and statistical surface properties; they do not test whether LLAMIA’s internalized representations produce game understanding that is perceptually meaningful to a skilled chess player. We conduct two human studies to address this directly: a gameplay identification study (Study 1) testing whether LLAMIA’s move choices are stylistically distinguishable from human play, and a commentary quality study (Study 2) testing whether LLAMIA’s commentary conveys more accurate strategic insight than the verbalized baseline and whether it enables readers to form more accurate board evaluations.
H.1 Participants
We recruited participants from university chess club chapters via in-person announcement at weekly club meetings. Eligibility required either a current FIDE rating or a verified Chess.com or Lichess rapid rating with rated games on record. Following the calibration task described below, participants met the inclusion criterion and proceeded to both studies.
Both human studies were conducted under a protocol approved by the Institutional Review Board (IRB). Participants were recruited voluntarily and provided written informed consent prior to enrollment. The consent form described the study purpose, the nature of all tasks, and the intended use of collected data. Participants were informed that some game segments and commentary samples were AI-generated.
Session data (gameplay judgments, Likert ratings, and open-text responses) were anonymized at the point of collection. Each participant was assigned a randomized identifier; no names, handles, or affiliations were retained in the analysis dataset. Raw recordings were deleted following transcription. Participant data will not be shared in identifiable form.
H.2 Sensitivity Calibration and Participant Selection
A participant’s ability to evaluate AI gameplay depends on their sensitivity to stylistic differences between human and engine play, not only their rating. We screen for this explicitly before the main studies.
Design.
Each participant reviews 20 recorded game segments in randomized order. Ten segments are drawn from human-vs.-human games (negative controls); the remaining ten from human-vs.-bot games in which the bot is Stockfish 17 at varied strength levels (), Maia (), or a weaker rule-based engine (). For each segment, the participant identifies which player (White or Black) is the bot via forced binary choice; human-vs.-human segments include a “Neither” option. Both positive and negative controls are required to measure true discrimination sensitivity rather than a bias to label any player as a bot.
Inclusion criterion.
Participants achieving overall accuracy ( correct) proceed to the main studies. Of recruited participants, met this criterion. The 12 included participants achieved a median calibration accuracy of 75%.
H.3 Study 1: Gameplay Identification
Stimuli.
We construct 30 game segments from held-out games in the LLAMIA-Bench evaluation set. Each segment comprises 10 consecutive half-moves (5 per side), drawn equally from middlegame and endgame phases (15 segments each). Opening segments are excluded: early-game play is dominated by memorized theory and reveals little about model behaviour. Segment boundaries are defined by board position (middlegame: pieces per side, material points; endgame: pieces per side or rook-and-pawn endings). Two players per segment are labeled Player A and Player B; one is drawn from a game involving LLAMIA, LLAMIA-Verb-14B, or a human player. Segments are rendered as fixed-speed board replays (3 seconds per half-move) with clock information removed to prevent trivial detection via time usage.
Task.
For each segment, participants respond to three prompts:
- 1.
Bot identification (primary): “Which player, A or B, do you believe is the AI?” (Forced choice; human-vs.-human segments include “Neither.”)
- 2.
Confidence (1–5 scale): “How confident are you in this judgment?”
- 3.
Open commentary (free text): “Which specific moves or patterns informed your decision?”
Design.
Each participant evaluates 10 segments randomly drawn from the pool of 30, keeping total session time to 40–50 minutes. Assignment is balanced so that every segment receives at least 4 independent judgments. Following bot identification, participants rate their gameplay experience for each segment they played:
- 1.
Human-likeness (1–5 Likert): “My opponent played like a human player.”
- 2.
Enjoyment (1–5 Likert): “I enjoyed this game.”
These subjective ratings provide convergent evidence alongside the objective detection accuracy: a system that is both hard to detect and rated as human-like in experience achieves qualitative human-likeness, not merely move-distribution similarity.
Qualitative coding.
Open-text responses are transcribed and coded along four dimensions by two independent annotators: (i) tactical cues—references to captures, checks, or forcing sequences; (ii) positional cues—references to pawn structure, piece activity, or long-term plans; (iii) stylistic cues—references to move tempo, unnatural patterns, or “computer-like” consistency; (iv) no identifiable cue—the participant could not articulate a reason. Inter-annotator agreement is reported as Cohen’s . This qualitative layer distinguishes tactical imitation from deeper stylistic assimilation: a system that merely selects strong moves will produce tactical cues; a system whose move distribution lacks non-human regularities will produce no-cue responses.
Primary metric.
Bot detection accuracy per system: the fraction of segments in which the participant correctly identifies the AI-controlled player. Lower accuracy on LLAMIA segments indicates a move distribution less readily distinguished from human play.
H.4 Study 2: Commentary Quality
Stimuli.
We select 15 board positions from the held-out LLAMIA-Bench Commentary test set, stratified by position complexity: 5 simple (centipawn loss ), 5 moderate (30–80), and 5 complex (). Positions are drawn from the same middlegame and endgame phases as Study 1. For each position, commentary is generated from all systems in the LLAMIA-Bench evaluation suite. Commentary operates at the position level—a single move and its strategic rationale—to isolate single-position reasoning and avoid narrative continuity confounds. For the preference and Likert tasks, participants see LLAMIA-14B and LLAMIA-Verb-14B side-by-side, labeled “System A” and “System B” with left-right assignment independently randomized. For the comparative state annotation task, each system’s commentary is presented individually. Every participant evaluates all 15 positions, yielding a fully crossed design ( raters 15 positions total judgments).
Dimensions.
Accuracy and Insight are the two scored dimensions. If latent state internalization gives LLAMIA access to richer engine representations than verbalization permits, the difference should manifest as greater factual accuracy (grounded in actual evaluation) and greater strategic depth (conveying non-obvious plans). Fluency and engagement are excluded: both systems produce grammatical prose, and metrics insensitive to chess content are unlikely to discriminate.
Task.
For each position (estimated 3–4 minutes), participants complete four items:
- 1.
Overall preference (forced choice with escape): System A / System B / No clear preference.
- 2.
Accuracy (1–5 Likert): “The commentary correctly describes what is happening on the board.”
- 3.
Insight (1–5 Likert): “The commentary reveals something strategically non-obvious about this position.”
- 4.
Comparative state annotation: After reading each system’s commentary for a middlegame position, the participant predicts the board evaluation on a 7-point scale ( = Black winning clearly, = equal, = White winning clearly). Administered per-system across all systems. Correctness is Pearson between predicted and actual Stockfish centipawn evaluations, averaged across raters.
Primary metrics.
(i) Preference rate for LLAMIA: fraction of judged pairs (excluding “no clear preference”) choosing the LLAMIA output, reported as a mean across 12 raters. (ii) Mean Accuracy and Insight Likert scores per system. (iii) Pearson with Stockfish per system.
Rater–judge agreement.
All 180 position pairs are independently scored with GPT-4o G-eval (Liu et al., 2023b) using matched Accuracy and Insight prompts. Cohen’s is computed between human preference rankings and G-eval rankings, with a length-adjusted computed after partialling out the Spearman correlation between G-eval score and commentary word count ().
H.5 Results
Study 1 – Bot detection accuracy.
Participants correctly identified LLAMIA-14B as the AI-controlled player in only 39% of trials, below the 50% chance level—yielding a human-pass rate of 61%. LLAMIA-Verb-14B was detected in 72% of trials (human-pass rate 28%), well above chance. Confidence ratings were lower for LLAMIA-14B segments (mean 2.6 vs. 3.4), indicating that near-chance detection reflects genuine perceptual ambiguity rather than participant disengagement. Catch-trial accuracy (human-vs.-human “Neither” responses) was 83%, confirming that participants withheld bot identification when none was warranted.
The detection gap is specific to the integration mode, not the backbone or training recipe. LLAMIA-Verb-14B uses the same Qwen3-14B backbone and the same DAPO training budget; its higher detectability is associated with the verbalization interface: verbal summaries impose regularities on move selection—consistent avoidance of dubious moves, move-tempo patterns—that participants identify as non-human. LLAMIA-14B, reasoning over latent representations, produces a move distribution that does not exhibit these regularities.
Qualitative coding () confirms this interpretation. The dominant detection cue for LLAMIA-Verb-14B segments was stylistic (65%: “moves felt too consistent,” “never played a dubious move”). LLAMIA-14B segments produced no-identifiable-cue responses in 38% of cases versus 5% for LLAMIA-Verb-14B. When a cue was identified for LLAMIA-14B, it was distributed across tactical and positional categories with no dominant signal.
Study 1 – Gameplay experience survey.
Post-game Likert ratings corroborate the objective detection results (Figure 9). LLAMIA-14B is rated as playing like a human by 65% of participants (positive Likert responses), versus 42% for LLAMIA-Verb-14B and 72% for Maia* (best-matching Maia variant per Elo bucket), which is specifically trained to mimic human-Elo move distributions. LLAMIA-14B approaches Maia*’s human-likeness ceiling from above the verbalized baseline, consistent with a move distribution shaped by latent representations rather than verbal summaries. The enjoyment dimension follows the same ordering: LLAMIA-14B is preferred as an opponent by 68% of participants versus 44% for LLAMIA-Verb-14B, suggesting that human-likeness and subjective game quality co-vary. That enjoyment tracks human-likeness rather than playing strength is consistent with the Centaur collaboration literature (Saghafian and Idan, 2024).
Study 2 – Commentary preference and quality.
LLAMIA-14B was preferred in 72.2% of all 180 judgments (, each rater judging every position); excluding the 10.0% no-preference responses (), 80.2% of the remaining 162 judged pairs favoured LLAMIA-14B ( raters, 162 judged pairs). Mean Accuracy: LLAMIA-14B 4.30 vs. LLAMIA-Verb-14B 3.20. Mean Insight: LLAMIA-14B 4.20 vs. 2.50. The Insight gap (1.70 scale points) substantially exceeds the Accuracy gap (1.10 points).
This gap structure is theoretically informative. Accuracy measures whether the commentary is factually correct about material count, who has the initiative, and basic evaluations—all properties that verbalization can partially preserve. Insight measures whether the commentary conveys the why behind a move: the long-range plan, the implied threat, the imbalance being exploited. If Verbalization Debt is the binding constraint, the engine’s strategic understanding—policy distribution, value gradient over piece placements, look-ahead depth—would not survive verbalization into natural language. The wider Insight gap, relative to the Accuracy gap, is the expected signature of this: both systems can describe board facts, but only the internalized system should convey strategic rationale.
Open-text responses reflect this structure. Participants described LLAMIA-14B’s commentary as referencing downstream consequences (“explains why the bishop trade matters three moves later”; coded as positional cues), while LLAMIA-Verb-14B’s was described as accurate but shallow (“correctly says White is better but doesn’t say why”; coded as no-cue or tactical). Rater–judge agreement: with G-eval; length-adjusted (), confirming G-eval’s systematic length bias.
Study 2 – Comparative state annotation.
Figure 11 reports the most direct test of the Verbalization Debt claim: does LLAMIA’s commentary enable participants to form more accurate board evaluations than verbalized commentary, and does this advantage scale with position complexity?
In simple positions (centipawn loss ), all systems produce comparable state annotation accuracy ( for LLAMIA-14B vs. for LLAMIA-Verb-14B, ). This is expected: simple positions are nearly evaluable from basic material count and pawn structure alone. The gap widens monotonically into moderate positions () and reaches in complex positions (centipawn loss ), where LLAMIA-14B achieves versus for LLAMIA-Verb-14B and for Qwen3-14B+Lc0 (same backbone, no RL training). The no-commentary condition ( at complex) confirms that differences are driven by commentary content rather than rater capability: participants without commentary cannot evaluate complex positions at all.
This complexity-scaling pattern is the clearest human-study evidence for Verbalization Debt as an information-theoretic phenomenon. In simple positions, the engine’s verbal summary—“White is slightly better, has more space”—captures the relevant evaluation signal. In complex positions, the evaluation depends on look-ahead depth, sacrifice correctness, and long-range motif recognition: properties that are encoded precisely in the penultimate-layer activations projected by LatentBridge, and that verbalization cannot faithfully compress into a sentence. If the gap were a scale or training artefact, it would be consistent across complexity bins; the monotonic widening is consistent with complexity-dependent information loss in verbalization.
Appendix I Dataset Construction
This section documents how the Stage 1 (projector alignment) and Stage 2 (task-specific RL) training corpora are assembled. The corresponding test-time disjointness guarantees—FEN-level non-overlap between every LLAMIA-Bench test split and the training pools described below, including the heuristic used for Agadmator-2K—are stated once in Section B.3 and are not repeated here.
I.1 Stage 1: Projector Alignment Data
Stage 1 trains the LatentBridge projector on state–policy pairs from the Lc0-BT4 forward pass. We construct the dataset as follows:
Source.
We sample 5M positions from the Lichess evaluation database (January 2013 – December 2024), filtering for standard-time-control games between rated players (1200 Elo). Positions are sampled uniformly across game phases (opening: moves 1–15, middlegame: moves 16–35, endgame: moves 36+) to prevent phase bias.
Label generation.
For each position, we run a single BT4 forward pass (no MCTS search) to obtain the raw policy distribution , the value head output , and the penultimate-layer activations . The training target is the top-1 move from the policy head, formatted as either UCI or SAN notation (70% / 30%). The activation is the input to the projector.
Prompt diversity.
Each position is paired with one of four question types (Section J.1): position evaluation, principal variation, legal moves, or brief description. Question types are sampled uniformly. This diversity prevents the projector from overfitting to a single output format.
Split.
The 5M positions are split by game ID (not by position) to prevent train/test leakage: 4.5M training, 250K validation, 250K held-out test. No game appears in more than one split.
I.2 Stage 2: Task-Specific RL Data
Stage 2 uses DAPO rollouts on task-specific prompts. The training data for RL totals 850K examples across all tasks:
Behavior Cloning.
500K positions from Lichess games, stratified by player Elo (100-point bins from 1100 to 2600). Each position is paired with the move actually played by the human player. The reward signal is based on rank within the engine’s top-3 moves: the model receives reward 1.0 for a top-1 match, 0.5 for top-2, 0.25 for top-3, and 0 otherwise. Top-1 exact match as a reward collapsed training; rank within the top-3 provides a denser, monotone signal.
Puzzle Understanding.
200K puzzles from the Lichess puzzle database, each annotated with difficulty rating and popularity score. The reward is a scaled negative absolute error between LLAMIA’s prediction and the ground truth.
Move Annotation.
100K annotated positions drawn from 90K games in the GameKnot GameKnot (2024) and Lichess annotation corpora (multiple annotations per game). The reward is a G-eval score (GPT-4o judge) comparing LLAMIA’s annotation to the reference.
Game Commentary.
50K annotated game segments (15–30 moves each) from grandmaster commentary databases, chess books transcribed to PGN, and Lichess studies with annotations. The reward combines a G-eval score for commentary quality with BLEU-2 against reference commentaries.
Appendix J Prompts and Templates
Three prompt regimes govern the pipeline: Stage 1 projector alignment, Stage 2 DAPO rollouts, and the shared <invoke> tool-call format.
J.1 Stage 1: Projector Alignment
Stage 1 trains the LatentBridge projector (, the linear adapter mapping BT4 residual activations into the LLM token space) via supervised learning on chess instruction data. Each episode presents several independent questions about the same board position; the <state> placeholder marks the state tokens projected from the BT4 residual stream and inserted into the LLM’s context at that point. Four question types are sampled per FEN: position evaluation, principal variation, legal moves, and brief verbal description. Move notation alternates UCI and SAN with probability 70 / 30 %; evaluations are in pawn units.
J.2 Stage 2: DAPO System Prompt
Toy task (puzzle popularity / Elo).
The system prompt below is used verbatim during DAPO rollouts for the toy task (§A.4.2). The explicit refusal-suppression clause is required: without it, Qwen3-4B defaults to “I cannot directly determine…” and never emits a tool call, collapsing the format-pass rate to 0 % (verified on 20 held-out puzzles at greedy decoding).
The popularity is <int> and the ELO is <int>
Full LLAMIA-Bench.
All four task families (behaviour cloning, puzzle understanding, move annotation, game commentary) share the same tool catalogue as the toy task. Task-specific instructions replace the popularity/Elo mandate; the refusal-suppression clause is retained across all variants.
J.3 <invoke> Trigger Format
Latent invocation (<invoke>) fires after every board-mutating tool call (make_move, undo_move, reset_position). The harness re-runs the BT4 forward pass on the updated FEN, encodes a fresh set of state tokens, and prepends them as a <state> prefix to the next user turn. Because mutating calls are dispatched sequentially (§J.4), the re-encoding always sees a fully settled board state. Read-only calls (analyze, get_policy, get_position) do not trigger re-encoding; the token budget is therefore capped at additional tokens per state transition, regardless of analysis depth.
Round 1 — LLAMIA: (emits two parallel analyze calls — see §J.4.3)
Round 2 — LLAMIA: Final answer citing Qa5+ / Qxg5.
Since no board mutation occurs in blunder analysis, <invoke> does not fire. The re-encoding path is active in multi-step planning episodes, where the agent sequences make_move calls to explore a variation before deciding on a recommendation.
Round 1 — LLAMIA: make_move(e2e4)
(harness re-encodes fresh <state1>)
Round 2 — User (injected): <state1>
Round 2 — LLAMIA: analyze(multipv=3)
Round 3 — LLAMIA: undo_move()
(harness re-encodes fresh <state>)
Round 4 — User (injected): <state>
Round 4 — LLAMIA: Final recommendation.
The injected <state> prefix in Rounds 2 and 4 is invisible to the human user; the harness inserts it programmatically before forwarding the tool result to the next LLM call, keeping the re-encoding fully transparent to the model’s reasoning loop.
J.4 Inference Harness and Tool-Call Protocol
The inference harness connects the LLM to lc0 via a six-tool stateful API. Read-only calls (analyze, get_policy, get_position) are dispatched in parallel; mutating calls (make_move, undo_move, reset_position) are dispatched sequentially to preserve board consistency.
J.4.1 Agent State
J.4.2 System Prompt
J.4.3 Tool Definitions
Limitations
Our evidence is drawn primarily from chess, where agent representations are well-characterized by interpretability work and evaluation is tractable. Within chess, the bottleneck holds across six Lc0-family networks spanning three sub-architectures (SE-ResNet: T72, T78; Transformer: T80, T82; large Transformer: BT3, BT4; Section E.2). On Go, latent collaboration with a frozen KataGo agent outperforms its verbalized counterpart at every backbone scale on behavior cloning (Appendix F), evidence that the effect is not chess-specific; however a complete multi-task Go suite remains future work. LLAMIA also requires access to the agent’s internal activations, which precludes application to closed-source agents without an intermediary.