arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.00355v2 [cs.AI] 29 Sep 2026

Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

Jungseob Lee    Seongtae Hong    Dongyub Jude Lee    Chanjun Park Affiliation: Korea University  Zoom Communications  Soongsil University    Jaehyung Seo    Sugyeong Eo ††thanks: Corresponding authors. Affiliation: Konkuk University  Yonsei University{omanma1928,ghdchlwls123,limhseok}@korea.ac.kr  jude.lee@zoom.uschanjun.park@ssu.ac.kr  seojae777@konkuk.ac.kr  s.eo@yonsei.ac.kr    Heuiseok Lim11footnotemark: 1
Abstract

Speculative decoding accelerates generation without changing its output, but on vision-language models (VLMs) a self-reinforcing cycle holds it back. Because an autoregressive drafter pays a sequential pass for each drafted token, it must stay small and can ill afford to attend to the image at each pass. Prior work therefore compresses or hides the image, leaving the drafter weakest on the text the image determines. We present GLANCE, a one-pass block drafter that breaks this cycle on an unmodified VLM target. Its block-diffusion head drafts a whole block in one forward pass over the target’s already fused vision-language states, reading the multimodal context once, however deep the draft. The target verifies a wide candidate tree in one pass and commits exactly its greedy output. In one production engine at a fixed round budget, GLANCE decodes up to 3.05 times faster than autoregressive decoding and outpaces the production EAGLE3-VL head on average and by about 11% on grounded tasks. An entropy law explains when drafting pays, predicting the longest accepted blocks on grounded tasks, where the target’s next-token entropy is lowest. Our code is available at https://github.com/js-lee-AI/GLANCE.

1 Introduction

Vision-language models (VLMs) increasingly produce long text about an image. They describe photographs, read scanned documents, and pull values off charts (Bai et al., 2025a; Bai et al., 2025b). Generation remains autoregressive, so every output token needs a full forward pass of the target model, and at small batch sizes this decode loop is bound by memory bandwidth rather than compute. The visual encoder, in contrast, runs once during prefill, and in our workloads the decode loop takes most of the end-to-end time. Faster grounded generation therefore has to come from the decode loop.

Refer to caption
Figure 1: Draftability of grounded and open generation. On a document (top), the image pins the answer, entropy is low, and one draft pass commits the entire block. Open captioning (bottom) admits many continuations, entropy is high, and the same pass commits only a short prefix.

Speculative decoding shortens this loop without changing the output. A cheap drafter proposes several future tokens, and the target verifies them in one forward pass, committing the longest prefix it agrees with (Leviathan et al., 2023; Chen et al., 2023). Modern drafters are lightweight heads that read the target’s own hidden features and propose a tree of candidates (Cai et al., 2024; Li et al., 2024b; Li et al., 2024a; Li et al., 2025; Miao et al., 2024), and vision-language serving stacks now ship such heads for VLM targets (AQ-MedAI, 2026).

Adapting this recipe to VLMs, however, has led into a cycle. The drafter is autoregressive, so a candidate kk tokens deep costs kk sequential draft passes, and the drafter is kept small so that those passes stay cheap. Each pass also attends over the whole multimodal context, in which image tokens usually outnumber text tokens, and a drafter that looks at the image pays for it again at every step. Heads that keep the full context, such as the production EAGLE3-VL head (AQ-MedAI, 2026), therefore keep their drafts shallow, and methods built for VLMs compress, prune, or hide the image tokens from the drafter (Gagrani et al., 2024; Kang et al., 2025; Huang et al., 2025a; Xie et al., 2026). In both cases drafts stay short, and in the second the drafter also loses the image evidence that grounded text depends on.

On grounded workloads, however, the image makes drafting work. When a VLM answers a question about a document, much of its output already appears on the page, and the next token is often nearly deterministic and frequently copied verbatim, as Figure 1 illustrates. We make this precise with an entropy law, under which the expected accepted length falls with the target’s next-token entropy, and grounded tasks concentrate at the low-entropy end.

Refer to caption
Figure 2: Overview of GLANCE. The block head (top) reads the frozen target’s fused vision-language states at five layers, never raw visual tokens. In one round (bottom), one draft pass fills the block, the top prefixes form a candidate tree, and one target pass verifies it.

To exploit this predictability, we propose long and wide candidate sets in a single draft pass. We present GLANCE (grounded block drafting with one-pass candidate expansion). A block-diffusion head (Chen et al., 2026; Arriola et al., 2025) reads the frozen target’s fused vision-language states and, in a single forward pass, produces a distribution for every position of a block. The highest-scoring prefixes form a wide candidate tree, and the target verifies all of them in one pass. Because the head has already paid for the whole block, width costs no further draft passes (Ringel and Romano, 2026; Zhang et al., 2026c). The committed text is exactly the target’s greedy output, which we prove and also verify in fp32 on every audited prompt. One-pass block drafting has so far been demonstrated only on text-only language models, while lossless drafters built for VLMs have remained autoregressive, and GLANCE is the first one-pass block drafter that is lossless on an unmodified VLM target.

Inside one production engine at a fixed round budget, GLANCE decodes up to 3.05×3.05\times faster than autoregressive decoding on chart question answering and outpaces the production EAGLE3-VL head on average and by about 11%11\% on every grounded task, using one draft pass where that head uses eight. Trained on the same corpus and schedule as an EAGLE-3 head and run under the same tree, it decodes 4.54.5 to 25.7%25.7\% faster on all five tasks, and it keeps its lead when retrained on the corpus of ViSpec, the strongest published VLM drafter. Moreover, the form of the entropy law carries over to other VLM targets and to speech, code, and chat.

This paper makes three contributions.

  • •

    We diagnose why speculative decoding has underdelivered on VLMs. Autoregressive drafting and reduced vision access reinforce each other, and together they forgo the long verbatim runs of grounded generation, which a one-pass drafter collects in a single pass.

  • •

    We introduce GLANCE, the first one-pass block drafter that is lossless on an unmodified VLM target. Its head drafts a whole block in one pass from the target’s fused vision-language states, a wide tree drawn from that pass is verified in one target pass, and we build it into SGLang, where it runs against the production head under one engine and one round budget.

  • •

    We establish an entropy law of draftability and test it with a mismatched-image intervention and a grounded-tail analysis on five tasks, three VLM targets, and four settings beyond vision.

2 Related Work

Draft heads and tree verification.

Speculative decoding accepts drafted tokens with a rejection rule that provably preserves the target distribution (Leviathan et al., 2023; Chen et al., 2023). The practical line drafts from the target’s own hidden features with a lightweight head. Medusa attaches independent heads for each future offset (Cai et al., 2024), Hydra makes them sequentially dependent (Ankner et al., 2024), and EAGLE drafts autoregressively at the feature level over dynamic candidate trees (Li et al., 2024b; Li et al., 2024a), which EAGLE-3 extends with training-time test and multi-layer feature fusion (Li et al., 2025). Many candidates can be verified at once by packing a prefix tree into a single ancestor-masked target pass (Miao et al., 2024), and later work studies tree shape (Chen et al., 2024b; Wang et al., 2025a), theory, and benchmarking (Yin et al., 2024; Huang et al., 2025b; Xia et al., 2024).

Speculative decoding for VLMs.

The first study of speculative decoding for multimodal models found a text-only drafter to be a strong baseline (Gagrani et al., 2024). Subsequent work follows two routes. One shrinks what the drafter reads, by compressing image tokens into a few adaptor embeddings (Kang et al., 2025), pruning visual tokens (Huang et al., 2025a), pruning video tokens under verifier guidance (Ji et al., 2025), or hiding them from the drafter altogether (Xie et al., 2026). The other gives a small drafter its own cheap path to vision, through multimodal distillation (Ganesan et al., 2025), dynamic draft trees (Huo et al., 2025), or refined target features fused by cross-attention (Hu et al., 2025). DREAM uses entropy to weight that fusion inside an autoregressive drafter, whereas we use entropy to predict acceptance and draft in one pass. Semi-autoregressive multimodal drafting relaxes the one-token-a-pass constraint but gives up exactness (Wang et al., 2025b), and a benchmark now covers the setting (Shen et al., 2026). All of these methods keep an autoregressive drafter and engineer its access to vision, whereas GLANCE inherits vision from the target and drafts in one pass.

Block drafting and diffusion drafters.

A parallel line removes the sequential bottleneck inside the drafter. It includes lookahead embeddings (Monea et al., 2023), Jacobi decoding (Fu et al., 2024), block-parallel discrete diffusion language models (Nie et al., 2025; Arriola et al., 2025), diffusion drafters verified by an autoregressive target (Christopher et al., 2025; Li et al., 2026a; Cheng et al., 2025), and DFlash, a block-diffusion head on the frozen target’s hidden states (Chen et al., 2026). DDTree and CaDDTree show that one-pass drafting makes tree width cheap for text-only targets (Ringel and Romano, 2026; Zhang et al., 2026c), and later text-only work builds on such drafters (Zhang et al., 2026a; Wu et al., 2026b; Zhang et al., 2026d; Wang et al., 2026; Li et al., 2026c; Kwon et al., 2026). Block drafting has reached VLMs only by converting the target itself into a self-speculating diffusion model (Wu et al., 2026a; Zhang et al., 2026b), which changes the model’s output. GLANCE instead keeps the target frozen and its output exact.

3 GLANCE

Figure 2 shows one decoding round of GLANCE. A block head drafts a whole block of future tokens in one forward pass, and the target verifies a wide tree of those candidates in one forward pass and commits its own greedy prefix. Only the head is trained, and the target stays frozen.

3.1 Speculative Decoding and Acceptance

Let p(⋅∣x)p(\cdot\mid x) be the frozen target’s next-token distribution given the multimodal prefix xx, which holds the text tokens and the encoded image. Greedy decoding emits arg​maxw⁡p​(w∣x)\argmax_{w}p(w\mid x), one token for each target forward pass. Speculative decoding instead proceeds in rounds (Leviathan et al., 2023; Chen et al., 2023). Given the committed prefix and a pending root token bb, a drafter proposes candidate continuations, and the target commits the longest prefix that matches its own greedy continuation Y=(Y1,Y2,…)Y=(Y_{1},Y_{2},\dots) after bb, where Yk=arg​maxwp(w∣x∘b∘Y1:k−1)Y_{k}=\argmax_{w}p(w\mid x\circ b\circ Y_{1:k-1}). When aa nonroot tokens are accepted, the round commits a+1a+1 tokens, namely the accepted prefix and one more token read off the last verified distribution. We report the mean acceptance length

τ=𝔼⁡[a+1]=𝔼⁡[a]+1,\tau\;=\;\mathbb{E}[a+1]\;=\;\mathbb{E}[a]+1, (1)

so autoregressive decoding has τ=1\tau{=}1. For systems that log 𝔼⁡[a]\mathbb{E}[a] instead, we add the committed token so that all rows share one scale. A larger τ\tau means fewer target passes for each generated token, and the wall-clock speedup is τ\tau discounted by the cost of drafting and verification.

3.2 One-Pass Drafting on Fused States

GLANCE drafts with a block-diffusion head (Chen et al., 2026) of 55 layers, 1.051.05B parameters, and block size B=16B{=}16, attached to a frozen Qwen3-VL-8B target. The first block position holds the pending root bb and the other B−1B-1 positions are masked, so one forward pass returns a marginal over the vocabulary at every offset j=1,…,Lj=1,\dots,L with L=B−1=15L=B-1=15. The head never reads raw visual tokens. Its cross-attention keys and values are the target’s hidden states at five layers spread over the depth of the stack, {1,9,17,25,33}\{1,9,17,25,33\} (Li et al., 2025). By the time these states exist, the target has already merged the image into its text representation, so the drafter inherits visual grounding without an encoder, an adaptor, or compressed image tokens of its own.

Assumption 1 (One-pass marginal interface).

Conditioned on the cached prefix xx and the pending root bb, the drafter returns one marginal qj(⋅∣x,b)q_{j}(\cdot\mid x,b) at every offset j=1,…,Lj=1,\dots,L in a single forward pass, before any token of the block is committed. A candidate prefix is ranked by its plug-in score π^(y1:ℓ)=∏j≤ℓqj(yj∣x,b)\hat{\pi}(y_{1:\ell})=\prod_{j\leq\ell}q_{j}(y_{j}\mid x,b).

An autoregressive drafter pays for depth with sequential passes, each attending over the full multimodal context, and the production EAGLE3-VL head takes three passes a round in its released configuration and eight in the production-engine comparison below. The block head pays one pass for the whole block, and the image enters that pass only through states the target has already computed.

3.3 Wide-Tree Verification and Exactness

Definition 1 (Candidate tree and accept rule).

A candidate tree SS is a prefix-closed set of nonempty strings after the root bb, of depth at most LL and budget N=|S|N=|S|. Packed into one ancestor-masked target pass, the row of a depth-dd node y1:dy_{1:d} yields exactly p(⋅∣x∘b∘y1:d)p(\cdot\mid x\circ b\circ y_{1:d}). With YY the target’s greedy continuation, the accepted length is A(S)=max{k≥0:Y1:j∈Sfor all j≤k}A(S)=\max\{k\geq 0:Y_{1:j}\in S\ \text{for all }j\leq k\}. The walk commits Y1:A⁡(S)Y_{1:A(S)} and emits the next target-greedy token from the last verified row as the new root.

The tree builder keeps the NN prefixes with the highest plug-in score π^\hat{\pi}, and the kept set is prefix-closed because no prefix scores below its extensions (Ringel and Romano, 2026; Zhang et al., 2026c). The tree is packed into one ancestor-masked target pass (Miao et al., 2024), the decoder walks it along the target’s greedy tokens, and Appendix B gives the complete round as pseudocode. Growing NN adds verifier work inside that single pass but no draft passes. We use N=63N{=}63 in our Hugging Face implementation and a 3131-node tree in the production engine, where together with the root GLANCE verifies the same 3232 draft tokens a round as the production head.

Theorem 1 (Lossless greedy equivalence, informal).

Assume the target’s argmax is unique at every visited prefix and its top-22 logit gap there exceeds the numerical difference between packed and unpacked evaluation. Then, for any candidate tree, the walk of Definition 1 commits exactly the target’s greedy autoregressive sequence, token for token.

We call a run bitwise identical to greedy decoding when its committed token ids equal those of the target’s own greedy decoding at every position. A weak head can only shorten the accepted prefixes, and no fallback path is needed, because every committed token is the argmax of a target row. We nevertheless measure it instead of assuming it, and the full statement and proof are in Appendix A. In fp32, GLANCE is bitwise identical to greedy decoding on all 6060 audited prompts. In bf16, the target does not reproduce its own greedy output across two runs even without any drafter, because kernel choice perturbs near-ties, and GLANCE reproduces that output at least as often as a second run of the target does. Exactness is therefore decided in fp32, and Appendix C reports both audits for every system.

3.4 Training

Only the block head is trained, while the vision tower and the target stay frozen. The training rows are the target’s own greedy generations on 8,0008{,}000 COCO-Caption2017 and 8,0008{,}000 TextVQA prompts. The objective is top-11 coverage of the target’s token at every block offset, with offset jj weighted by 0.8j0.8^{j}, and a reveal scheme exposes a random prefix of the block so that the head learns every reveal length. The head is initialized from a released text-only block head for Qwen3-8B (Chen et al., 2026) and trained for one epoch on one GPU. The recipe contains no document, infographic, or chart data. The full training card is in Appendix C.

4 An Entropy Law of Draftability

Which workloads does a one-pass drafter serve best? We answer with one scalar, the target’s next-token entropy HH at the round root. Entropy has been used to gate or stop drafting (Agrawal et al., 2024; Tong et al., 2026; Mahmoud, 2026), and we give its relation to acceptance a closed form that the experiments then test.

We model the matches along a block as a survival process. The accepted run continues at offset jj with probability pm,jp_{\mathrm{m},j}, and we approximate these hazards by a single match probability pm​(H)p_{\mathrm{m}}(H) that varies slowly with the offset. Taking the match logit to be affine in the entropy gives

logit⁡pm​(H)\displaystyle\operatorname{logit}p_{\mathrm{m}}(H) =b0−b1​H,\displaystyle=b_{0}-b_{1}H, (2)
𝔼⁡[a∣H]\displaystyle\mathbb{E}[a\mid H] ≈∑ℓ=1Lpmℓ=pm​(1−pmL)1−pm,\displaystyle\approx\;\sum_{\ell=1}^{L}p_{\mathrm{m}}^{\ell}=\frac{p_{\mathrm{m}}\,(1-p_{\mathrm{m}}^{L})}{1-p_{\mathrm{m}}},

with block horizon L=15L{=}15. Here b0b_{0} is the zero-entropy intercept and b1>0b_{1}>0 the entropy slope, both fitted on each task’s round log. The truncation matters only near H=0H{=}0, since the untruncated mean pm/(1−pm)p_{\mathrm{m}}/(1-p_{\mathrm{m}}) exceeds the truncated one by the relative amount pmLp_{\mathrm{m}}^{L}, which is below 4%4\% whenever pm≤0.8p_{\mathrm{m}}\leq 0.8. Appendix A derives Equation (2), grounds its monotone part in Fano’s inequality, and bounds how much round-level variance any entropy-only predictor can explain.

Equation (2) describes the top-11 path of a draft. A wider tree adds siblings at every depth and so raises acceptance at every entropy, and we measure this gain to be a nearly constant factor across tasks.

Two predictions follow. First, 𝔼⁡[a∣H]\mathbb{E}[a\mid H] strictly decreases in HH, hence any workload that systematically lowers the target’s entropy yields longer accepted blocks. Grounded generation is such a workload, because it copies determinate strings such as glyphs, numbers, and answer spans from the image. Second, the law has a ceiling that grounded generation breaks.

Theorem 2 (Grounded tail, informal).

Equation (2) caps the near-certain mean at eb0e^{b_{0}}. If near-zero-entropy rounds enter, with probability π\pi, a verbatim-copy state whose matches persist, the mean as H→0H\to 0 becomes π​L+(1−π)​M\pi L+(1-\pi)M, where MM is the mean of the ordinary state, and it exceeds eb0e^{b_{0}} once π\pi passes an explicit threshold.

The copy regime is where one-pass drafting gains the most, because a verbatim run of length ℓ\ell costs an autoregressive drafter ℓ\ell sequential passes and a block drafter one. Appendix A states the two-state model behind Theorem 2 and tests it on prompt-disjoint splits.

EAGLE3-VL, eight draft passes GLANCE, one draft pass
Task H¯\bar{H} AR ms/tok τ↑\tau\ \uparrow ms/tok ↓\downarrow speedup ↑\uparrow τ↑\tau\ \uparrow ms/tok ↓\downarrow speedup ↑\uparrow GLANCE faster by
Higher-entropy tasks
Captioning 0.470.47 24.7224.72 3.51\mathbf{3.51} 11.54\mathbf{11.54} 2.14×\mathbf{2.14\times} 2.922.92 13.1013.10 1.89×1.89\times −11.9%-11.9\%
TextVQA 0.390.39 24.3424.34 4.31\mathbf{4.31} 9.56\mathbf{9.56} 2.55×\mathbf{2.55\times} 3.573.57 10.7610.76 2.26×2.26\times −11.2%-11.2\%
Lower-entropy tasks
InfographicVQA 0.310.31 24.5824.58 3.513.51 11.9911.99 2.05×2.05\times 3.65\mathbf{3.65} 10.84\mathbf{10.84} 2.27×\mathbf{2.27\times} +10.6%\mathbf{+10.6\%}
DocVQA 0.150.15 24.9124.91 3.683.68 11.6711.67 2.14×2.14\times 3.91\mathbf{3.91} 10.51\mathbf{10.51} 2.37×\mathbf{2.37\times} +11.0%\mathbf{+11.0\%}
ChartQA 0.190.19 24.1824.18 4.504.50 8.798.79 2.75×2.75\times 4.69\mathbf{4.69} 7.93\mathbf{7.93} 3.05×\mathbf{3.05\times} +10.8%\mathbf{+10.8\%}
Geometric mean 2.31×2.31\times 2.34×\mathbf{2.34\times} +1.3%\mathbf{+1.3\%}
Table 1: GLANCE against the production EAGLE3-VL head in SGLang, with engine, GPU, and a tree of 3232 draft tokens shared, filled in eight passes by EAGLE3-VL and one by GLANCE. H¯\bar{H} is the target’s mean entropy in nats. Blue marks GLANCE.

5 Experiments

5.1 Setup

Models and tasks.

The target is Qwen3-VL-8B-Instruct (Bai et al., 2025a), kept frozen. We evaluate five tasks ordered from open-ended to grounded, namely COCO captioning (Lin et al., 2014), TextVQA (Singh et al., 2019), InfographicVQA (InfoVQA) (Mathew et al., 2022), DocVQA (Mathew et al., 2021), and ChartQA (Masry et al., 2022). We call the last three, whose answers are read off a document or chart, the grounded tasks, and they are the three tasks of lowest mean target entropy in Table 1. The rest of the standard VLM suite is multiple-choice or single-word and leaves no decode loop to shorten. Captioning and TextVQA prompts for GLANCE come from the test split of each source, which is disjoint from its training rows, and the other three tasks appear in none of the primary head’s training data.

Baselines.

Following the EAGLE series (Li et al., 2024b; Li et al., 2024a; Li et al., 2025), we compare against released methods that preserve the target’s output. These are n-gram prompt lookup, classic two-model speculation with image-conditioned and text-only drafts, the production EAGLE3-VL head (AQ-MedAI, 2026), and ViSpec (Kang et al., 2025) together with the EAGLE-2 and Medusa baselines of its codebase. No published VLM drafter releases a head for our target, so we train the ViSpec-codebase heads ourselves with the official recipe.

Protocol.

Decoding is batch one and greedy, with up to 256256 new tokens. The acceptance length τ\tau counts tokens rather than time and thus compares systems across engines. Wall-clock speedup is the ratio of the mean decode time for a token, autoregressive over speculative, with both arms measured in one engine on one GPU, and we never compare speedups across engines. Prompt counts, hardware, and results under sampling at temperature 11 are in Appendix C.

acceptance length τ↑\tau\ \uparrow
Method (draft passes a round) Params Captioning TextVQA InfoVQA DocVQA ChartQA Lossless
Qwen3-VL-8B, training-free and two-model drafting
n-gram lookup (0) 00 1.291.29 2.492.49 2.532.53 3.303.30 2.572.57 ✓\checkmark
Classic SD, Qwen3-VL-4B (8) 4.44.4B 3.53\mathbf{3.53} 3.79\mathbf{3.79} 3.95\mathbf{3.95} 4.39\mathbf{4.39} 4.47\mathbf{4.47} ✓\checkmark
Classic SD, Qwen3-1.7B text-only (8) 2.02.0B 1.491.49 1.421.42 1.901.90 1.681.68 2.122.12 ✓\checkmark
Qwen3-VL-8B, trained draft heads
EAGLE3-VL, production (5) 0.400.40B 2.132.13 2.662.66 2.482.48 2.562.56 2.922.92 ✓\checkmark
EAGLE-2, ViSpec codebase (3) 0.230.23B 2.412.41 2.452.45 2.382.38 2.542.54 2.892.89 ✓†\checkmark^{\dagger}
ViSpec, official recipe (3) 0.310.31B 2.452.45 2.432.43 2.452.45 2.462.46 2.952.95 ✓†\checkmark^{\dagger}
Medusa, ViSpec codebase (1) 0.080.08B 1.511.51 1.521.52 1.471.47 1.531.53 1.611.61 ✓†\checkmark^{\dagger}
GLANCE (1) 1.051.05B 3.09\mathbf{3.09} 3.46\mathbf{3.46} 3.75\mathbf{3.75} 3.76\mathbf{3.76} 5.12\mathbf{5.12} ✓†\checkmark^{\dagger}
Qwen3-VL-8B, matched training with one corpus, target, batch, schedule, and framework
EAGLE-3 head, depth-3 chain (3) 0.400.40B 1.621.62 1.591.59 1.551.55 1.581.58 1.871.87 ✓\checkmark
GLANCE, budget-63 tree (1) 1.051.05B 4.05\mathbf{4.05} 3.97\mathbf{3.97} 3.90\mathbf{3.90} 4.13\mathbf{4.13} 7.44\mathbf{7.44} ✓†\checkmark^{\dagger}
Qwen3-VL-8B, GLANCE retrained on ViSpec’s corpus
GLANCE, 1 epoch (1) 1.051.05B 3.023.02 3.443.44 3.753.75 3.783.78 5.025.02 ✓†\checkmark^{\dagger}
GLANCE, 21 epochs (1) 1.051.05B 3.043.04 3.493.49 3.783.78 3.853.85 5.075.07 ✓†\checkmark^{\dagger}
ViSpec’s home target Qwen2.5-VL-7B
ViSpec, released head (3) 0.350.35B 3.343.34 3.26\mathbf{3.26} 3.213.21 3.043.04 3.593.59 ✓†\checkmark^{\dagger}
GLANCE (1) 1.231.23B 4.00\mathbf{4.00} 2.612.61 3.47\mathbf{3.47} 3.12\mathbf{3.12} 4.72\mathbf{4.72} ✓†\checkmark^{\dagger}
Table 2: Acceptance length of released drafters and of GLANCE, each at its own operating point, GLANCE with the budget-6363 tree. Params counts drafter weights. Under Lossless, ✓\checkmark marks exact acceptance and †\dagger output audited bitwise identical to greedy decoding. Blue marks GLANCE.

5.2 Head-to-Head in a Production Engine

Table 1 shows GLANCE and the production EAGLE3-VL head inside SGLang 0.5.60.5.6 on one RTX A6000 in bf16, with both CUDA-graph captured and both verifying a tree of 3232 draft tokens a round, EAGLE3-VL its own top-88 dynamic tree and GLANCE its prefix tree. Engine, GPU, and round budget are thus shared, and the one structural difference left is how the draft tokens are produced, in eight sequential passes by EAGLE3-VL and in one by GLANCE. This eight-pass tree decodes 1818 to 33%33\% faster than EAGLE3-VL’s released configuration of three passes over four draft tokens on every task. On the three lower-entropy tasks, GLANCE is the faster system by 10.610.6 to 11.0%11.0\% and reaches 3.05×3.05\times the speed of autoregressive decoding on ChartQA. Every paired bootstrap interval excludes zero, GLANCE is faster on at least 8181 of the 101101 prompts of each of these tasks, and it also leads on the five-task geometric mean, by 1.3%1.3\% with a 95%95\% interval from 0.30.3 to 2.2%2.2\%. None of the three tasks appears in GLANCE’s training data. The lead on these three tasks holds for two further training seeds and at temperature 11.

On captioning and TextVQA, the two tasks with the highest mean entropy H¯\bar{H}, the eight-pass head leads. The product-of-marginals ranking of Assumption 1 accounts for this split. It is accurate when the tokens of a block are nearly determined by the image, whereas on free-running text an autoregressive drafter stays coherent by construction. Consistently, GLANCE accepts longer blocks on every lower-entropy task than on either higher-entropy task, whereas the production head places TextVQA above both DocVQA and InfographicVQA. The production head’s lead on these two tasks also comes from its training, on 400400K ALLaVA-4V samples (AQ-MedAI, 2026) against our 1616K. With an EAGLE-3 head and a GLANCE head trained on one corpus and run under the same tree, GLANCE is faster on all five tasks, by 4.5%4.5\% on captioning, 14.9%14.9\% on TextVQA, and 16.116.1 to 25.7%25.7\% on the three grounded tasks, as Table 4 shows. On full-resolution InfographicVQA prompts, the margin holds at 6.9%6.9\% with 22K tokens of context and at 5.5%5.5\% with about 6.76.7K, and both 95%95\% intervals lie above zero. Details, including the prompt-level intervals, are in Appendix E.

Figure 3: Sources of GLANCE’s acceptance on Qwen3-VL-8B. (a) The identical head run as a width-1 chain and as a budget-63 tree. (b) The head reading the target’s fused states or a text-only language model’s states, at budget 31. (c) Speedup against the verifier budget NN.

5.3 Acceptance Against Released Drafters

Table 2 reports the acceptance length of every drafter released for these targets, each at its own operating point. Among trained heads, GLANCE accepts the longest blocks on all five tasks. Classic two-model speculation with a 44B draft accepts more on four of the five tasks, but a draft half the size of the target costs about half a target pass for each drafted token, so it slows decoding on every task, as Appendix F shows. The ViSpec codebase isolates what its vision adaptor adds. Trained without the adaptor, the head is exactly the EAGLE-2 drafter, and the two heads differ by at most 0.080.08 in acceptance length on any task, so the target’s fused states, which both heads read, already carry what the adaptor’s compressed image tokens add.

Four controls separate the architecture from its training. First, trained from scratch with one 2626K-row corpus, target, global batch, schedule, and framework (Li et al., 2026b), GLANCE decodes faster than an EAGLE-3 head on all five tasks under the shared tree of Section 5.2, where that head’s acceptance length is 2.82.8 to 3.53.5. This corpus, unlike our primary recipe, contains document and chart rows. Second, retrained on ViSpec’s own 6868K-row corpus, at one epoch and at ViSpec’s twenty-one, GLANCE stays within one percent of its acceptance under our recipe, pooled over the five tasks, and its lead over the ViSpec family therefore holds on ViSpec’s own data and schedule. Third, on ViSpec’s home target Qwen2.5-VL-7B, a GLANCE head trained on that target accepts longer blocks than the released ViSpec head on four of the five tasks. Fourth, the released text head that GLANCE starts from, run unchanged, trails the production head on all five tasks of Table 1, by 15%15\% on the geometric mean, and our training lengthens its accepted blocks by 1717 to 21%21\%, which turns that deficit into GLANCE’s lead. Details of all four controls are in Appendix C.

Figure 4: Tests of the entropy law on Qwen3-VL-8B. (a) Accepted length against the target’s next-token entropy, by entropy decile. (b) Acceptance gain from the correct over a mismatched image, with 95%95\% bootstrap intervals. (c) The fitted law’s cap eb0e^{b_{0}} against the measured lowest-decile mean.

5.4 Where the Acceptance Comes From

Tree against chain.

GLANCE’s head has 2.6×2.6\times the parameters of EAGLE3-VL’s, so a larger head might explain the gains. Panel (a) of Figure 3 measures what the tree adds by running the identical head as a width-11 chain. The budget-6363 tree accepts between 1.451.45 and 1.49×1.49\times the chain’s length on every task, a nearly constant factor, as the law anticipates. The tree also decodes 1.36×1.36\times faster than the chain in wall-clock time, net of verifying a wider tree. In SGLang, a GLANCE round is 44 to 7%7\% shorter than an EAGLE3-VL round on every task, so the larger head drafts all fifteen offsets in one pass for less than the small head pays for eight.

Vision through fused states.

Panel (b) replaces the target’s fused states with those of a text-only Qwen3-8B, which never sees the image and whose states the head was initialized on. Acceptance falls on every task, and it falls most on grounded ones, since the text-only states retain 97%97\% of GLANCE’s acceptance on captioning but only 80%80\% on DocVQA. Zeroing the conditioning states collapses τ\tau to 1.111.11, which shows that the head drafts from the target’s context.

Width saturates early.

Panel (c) sweeps the verifier budget. Speedup rises from N=15N{=}15 to 3131 on every task and then flattens. Width pays the most on ChartQA, the task with the longest accepted blocks. Appendix E tabulates the sweep and breaks down the cost of a round.

5.5 Testing the Entropy Law

Fit across tasks.

Panel (a) of Figure 4 shows GLANCE’s accepted length falling with entropy on all five tasks, with grounded tasks accepting more at nearly every entropy level. Pooled over the 7,7517{,}751 rounds of captioning, TextVQA, and DocVQA, the law has b0=0.96b_{0}{=}0.96 and b1=0.63b_{1}{=}0.63, and the slope steepens with grounding, from 0.460.46 on captioning to 1.711.71 on ChartQA. Isotonic fits explain at least 96%96\% of the variance of the ten entropy-decile means on every task. The fits are listed in Appendix G.

Image and entropy.

To test whether the image itself is responsible, we intervene on the image alone. The same prompts are decoded with the correct and with a mismatched image along a shared teacher-forced trajectory. Panel (b) shows that the correct image lengthens accepted blocks on every task, and more so as grounding increases. The correct image also lowers the target’s mean entropy along that trajectory, from 0.720.72 to 0.260.26 nats on DocVQA, and once entropy is held fixed the remaining image effect is no longer positive. Together with the fused-state ablation, this places the image’s contribution in the entropy channel.

The grounded tail.

Panel (c) compares the cap eb0e^{b_{0}} of Theorem 2 with the measured mean of aa in each task’s lowest-entropy decile. Every floor lies above its cap, and the excess grows with grounding, reaching 5.235.23 against 3.413.41 on DocVQA. On ChartQA, 11.5%11.5\% of all rounds accept eight or more tokens.

Transfer.

The law keeps its form under a change of drafter, target, and modality, with a positive slope in every case. Two-model drafts on Qwen2.5-VL-7B and Qwen2-VL-7B fit at b1=0.67b_{1}{=}0.67 and 0.400.40, and beyond vision the slope stays positive on speech recognition, chat, code, and a non-Qwen backbone. These targets all feed vision as a token prefix, and Appendix G reports the fits and a cross-attention target.

6 Conclusion

We presented GLANCE, a lossless one-pass block drafter for frozen vision-language models. Its block-diffusion head reads vision through the target’s fused states and fills a whole block in one pass, which suits the long verbatim runs of grounded generation. A wide tree verified in one target pass turns that block into accepted length, and the output stays exactly the target’s greedy decoding. An entropy law explains when this pays and carries over across targets and modalities.

Limitations

Our measurements are at batch one, the regime in which decoding is bound by memory bandwidth and speculative decoding is deployed for latency. We do not measure serving at large batch sizes, which changes the cost of verification. The exactness guarantee concerns greedy decoding. Under sampling, the walk draws each token from the verified row of its node and descends while the draw is a child in the tree. Every committed token is then a draw from the target’s own distribution, and the output distribution is preserved while byte identity is not defined. Our characterization of draftability is stated for targets that feed vision as a token prefix, which covers the Qwen-VL family studied here. Finally, we study single-image tasks whose outputs span tens to hundreds of tokens, since tasks with single-token answers leave no decode loop to shorten.

Ethical Considerations

This work accelerates inference of existing vision-language models without changing their greedy outputs, so it introduces no new model behavior. Risks of the underlying models, such as biased or incorrect descriptions of images, carry over unchanged and are neither amplified nor mitigated. All datasets are public research benchmarks used under their licenses, and no human subjects or personal data are involved. Faster decoding lowers the energy spent on each generated token.

References

  • Agrawal et al. (2024) Sudhanshu Agrawal, Wonseok Jeon, and Mingu Lee. 2024. AdaEDL: Early draft stopping for speculative decoding of large language models via an entropy-based lower bound on token acceptance probability. Preprint, arXiv:2410.18351.
  • Ankner et al. (2024) Zachary Ankner, Rishab Parthasarathy, Aniruddha Nrusimha, Christopher Rinard, Jonathan Ragan-Kelley, and William Brandon. 2024. Hydra: Sequentially-dependent draft heads for Medusa decoding. In Conference on Language Modeling (COLM). ArXiv:2402.05109.
  • AQ-MedAI (2026) AQ-MedAI. 2026. Qwen3-VL-8B-Instruct-eagle3. https://huggingface.co/AQ-MedAI/Qwen3-VL-8B-Instruct-eagle3. Hugging Face model checkpoint.
  • Arriola et al. (2025) Marianne Arriola, Aaron Gokaslan, Justin T. Chiu, Zhihan Yang, Zhixuan Qi, Jiaqi Han, Subham Sekhar Sahoo, and Volodymyr Kuleshov. 2025. Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models. In International Conference on Learning Representations.
  • Bai et al. (2025a) Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, and 1 others. 2025a. Qwen3-VL Technical Report. arXiv preprint arXiv:2511.21631.
  • Bai et al. (2025b) Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, and 1 others. 2025b. Qwen2.5-VL Technical Report. arXiv preprint arXiv:2502.13923.
  • Cai et al. (2024) Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D Lee, Deming Chen, and Tri Dao. 2024. Medusa: Simple LLM inference acceleration framework with multiple decoding heads. In International Conference on Machine Learning.
  • Chen et al. (2023) Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper. 2023. Accelerating large language model decoding with speculative sampling. arXiv preprint arXiv:2302.01318.
  • Chen et al. (2024a) Guiming Hardy Chen, Shunian Chen, Ruifei Zhang, Junying Chen, Xiangbo Wu, Zhiyi Zhang, Zhihong Chen, Jianquan Li, Xiang Wan, and Benyou Wang. 2024a. ALLaVA: Harnessing GPT4V-synthesized data for lite vision-language models. Preprint, arXiv:2402.11684.
  • Chen et al. (2026) Jian Chen, Yesheng Liang, and Zhijian Liu. 2026. DFlash: Block Diffusion for Flash Speculative Decoding. In International Conference on Machine Learning.
  • Chen et al. (2024b) Zhuoming Chen, Avner May, Ruslan Svirschevski, Yuhsun Huang, Max Ryabinin, Zhihao Jia, and Beidi Chen. 2024b. Sequoia: Scalable and robust speculative decoding. In Advances in Neural Information Processing Systems. ArXiv:2402.12374.
  • Cheng et al. (2025) Zicong Cheng, Guo-Wei Yang, Jia Li, Zhijie Deng, Meng-Hao Guo, and Shi-Min Hu. 2025. DEER: Draft with diffusion, verify with autoregressive models. Preprint, arXiv:2512.15176.
  • Christopher et al. (2025) Jacob K Christopher, Brian R Bartoldson, Tal Ben-Nun, Michael Cardei, Bhavya Kailkhura, and Ferdinando Fioretto. 2025. Speculative diffusion decoding: Accelerating language generation through diffusion. In Proceedings of the North American Chapter of the Association for Computational Linguistics (NAACL). ArXiv:2408.05636.
  • Fu et al. (2024) Yichao Fu, Peter Bailis, Ion Stoica, and Hao Zhang. 2024. Break the sequential dependency of LLM inference using lookahead decoding. In Proceedings of the International Conference on Machine Learning (ICML). ArXiv:2402.02057.
  • Gagrani et al. (2024) Mukul Gagrani, Raghavv Goel, Wonseok Jeon, Junyoung Park, Mingu Lee, and Christopher Lott. 2024. On speculative decoding for multimodal large language models. arXiv preprint arXiv:2404.08856. ELVM @ CVPR 2024.
  • Ganesan et al. (2025) Mugilan Ganesan, Shane Segal, Ankur Aggarwal, Nish Sinnadurai, Sean Lie, and Vithursan Thangarasa. 2025. MASSV: Multimodal adaptation and self-data distillation for speculative decoding of vision-language models. In Findings of EMNLP.
  • Hu et al. (2025) Yunhai Hu, Tianhua Xia, Zining Liu, Rahul Raman, Xingyu Liu, Bo Bao, Eric Sather, Vithursan Thangarasa, and Sai Qian Zhang. 2025. DREAM: Drafting with refined target features and entropy-adaptive cross-attention fusion for multimodal speculative decoding. In Advances in Neural Information Processing Systems. ArXiv:2505.19201.
  • Huang et al. (2025a) Haiduo Huang, Fuwei Yang, Zhenhua Liu, Xuanwu Yin, Dong Li, Pengju Ren, and Emad Barsoum. 2025a. SpecVLM: Fast speculative decoding in vision-language models. Preprint, arXiv:2509.11815.
  • Huang et al. (2025b) Kaixuan Huang, Xudong Guo, and Mengdi Wang. 2025b. SpecDec++: Boosting speculative decoding via adaptive candidate lengths. In Conference on Language Modeling (COLM).
  • Huo et al. (2025) Mingxiao Huo, Jiayi Zhang, Hewei Wang, Jinfeng Xu, Zheyu Chen, Huilin Tai, and Yijun Chen. 2025. Spec-LLaVA: Accelerating vision-language models with dynamic tree-based speculative decoding. arXiv preprint arXiv:2509.11961.
  • Ji et al. (2025) Yicheng Ji, Jun Zhang, Heming Xia, Jinpeng Chen, Lidan Shou, Gang Chen, and Huan Li. 2025. SpecVLM: Enhancing speculative decoding of video llms via verifier-guided token pruning. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP). ArXiv:2508.16201.
  • Kang et al. (2025) Jialiang Kang, Han Shu, Wenshuo Li, Yingjie Zhai, and Xinghao Chen. 2025. ViSpec: Accelerating vision-language models with vision-aware speculative decoding. In Advances in Neural Information Processing Systems.
  • Kwon et al. (2026) Young D. Kwon, Miles Williams, Rui Li, Alexandros Kouris, and Stylianos I. Venieris. 2026. WhiFlash: Accelerating speculative decoding with token-level cross-paradigm routing. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP). ArXiv:2606.07710.
  • Leviathan et al. (2023) Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023. Fast inference from transformers via speculative decoding. In International Conference on Machine Learning.
  • Li et al. (2026a) Guanghao Li, Zhihui Fu, Min Fang, Qibin Zhao, Ming Tang, Chun Yuan, and Jun Wang. 2026a. DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding. In Findings of ACL.
  • Li et al. (2026b) Shenggui Li, Chao Wang, Yikai Zhu, Yubo Wang, Fan Yin, Shuai Shi, Yefei Chen, Xiaomin Dong, Qiaoling Chen, Jin Pan, and 1 others. 2026b. SpecForge: A flexible and efficient open-source training framework for speculative decoding. arXiv preprint arXiv:2603.18567.
  • Li et al. (2026c) Tianyi Li, Yaxin Luo, Xinyi Shang, and Zhiqiang Shen. 2026c. DARTree: Speculative diffusion decoding with autoregressive draft trees. Preprint, arXiv:2608.13524.
  • Li et al. (2024a) Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. 2024a. EAGLE-2: Faster inference of language models with dynamic draft trees. In Empirical Methods in Natural Language Processing.
  • Li et al. (2024b) Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. 2024b. EAGLE: Speculative sampling requires rethinking feature uncertainty. In International Conference on Machine Learning.
  • Li et al. (2025) Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. 2025. EAGLE-3: Scaling up inference acceleration of large language models via training-time test. In Advances in Neural Information Processing Systems. ArXiv:2503.01840.
  • Lin et al. (2014) Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. In ECCV.
  • Mahmoud (2026) Saif Mahmoud. 2026. Acceptance dynamics across cognitive domains in speculative decoding. Preprint, arXiv:2604.14682.
  • Masry et al. (2022) Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. 2022. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. In Findings of ACL.
  • Mathew et al. (2022) Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawahar. 2022. InfographicVQA. In WACV.
  • Mathew et al. (2021) Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. 2021. DocVQA: A dataset for VQA on document images. In WACV.
  • Miao et al. (2024) Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, Chunan Shi, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia. 2024. SpecInfer: Accelerating large language model serving with tree-based speculative inference and verification. In ASPLOS.
  • Monea et al. (2023) Giovanni Monea, Armand Joulin, and Edouard Grave. 2023. PaSS: Parallel speculative sampling. arXiv preprint arXiv:2311.13581. NeurIPS 2023 Workshop on Efficient Natural Language and Speech Processing.
  • Nie et al. (2025) Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, and Chongxuan Li. 2025. Large Language Diffusion Models. In Advances in Neural Information Processing Systems.
  • Ringel and Romano (2026) Liran Ringel and Yaniv Romano. 2026. Accelerating Speculative Decoding with Block Diffusion Draft Trees. In Conference on Language Modeling (COLM).
  • Shen et al. (2026) Hui Shen, Xin Wang, Ping Zhang, Yunta Hsieh, Qi Han, Zhongwei Wan, Ziheng Zhang, Jingxuan Zhang, Jing Xiong, Ziyuan Liu, Yifan Zhang, Hangrui Cao, Chenyang Zhao, and Mi Zhang. 2026. MMSpec: Benchmarking speculative decoding for vision-language models. Preprint, arXiv:2603.14989.
  • Singh et al. (2019) Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. 2019. Towards VQA models that can read. In CVPR.
  • Tong et al. (2026) Yujia Tong, Tian Zhang, Yunyang Wan, Kaiwei Lin, Jingling Yuan, and Chuang Hu. 2026. SAGE: Accelerating vision-language models via entropy-guided adaptive speculative decoding. Preprint, arXiv:2602.00523.
  • Wang et al. (2025a) Jikai Wang, Yi Su, Juntao Li, Qingrong Xia, Zi Ye, Xinyu Duan, Zhefeng Wang, and Min Zhang. 2025a. OPT-Tree: Speculative decoding with adaptive draft tree structure. Transactions of the Association for Computational Linguistics (TACL). ArXiv:2406.17276.
  • Wang et al. (2026) Zheng Wang, Davis Wertheimer, Yu Chin Fabian Lim, Mudhakar Srivatsa, Raghu K. Ganti, Minjia Zhang, and Naigang Wang. 2026. xPress: Parallel refinement for diffusion drafters in speculative decoding. Preprint, arXiv:2608.02438.
  • Wang et al. (2025b) Zihua Wang, Ruibo Li, Haozhe Du, Joey Tianyi Zhou, Yu Zhang, and Xu Yang. 2025b. SpecFLASH: A latent-guided semi-autoregressive speculative decoding framework for efficient multimodal generation. Preprint, arXiv:2505.12728.
  • Wu et al. (2026a) Chengyue Wu, Shiyi Lan, Yonggan Fu, Sensen Gao, Jin Wang, Jincheng Yu, Jose M. Alvarez, Pavlo Molchanov, Ping Luo, Song Han, Ligeng Zhu, and Enze Xie. 2026a. Fast-dVLM: Efficient block-diffusion vlm via direct conversion from autoregressive vlm. arXiv preprint arXiv:2604.06832.
  • Wu et al. (2026b) Tianyu Wu, Yu Yao, Zhenting Qi, Han Zheng, Zhuohan Wang, Haoran Ma, Lawrence Liao, Himabindu Lakkaraju, Ju Li, and Yilun Du. 2026b. D-PACE: Dynamic position-aware cross-entropy for parallel speculative drafting. Preprint, arXiv:2605.18810.
  • Xia et al. (2024) Heming Xia, Zhe Yang, Qingxiu Dong, Peiyi Wang, Yongqi Li, Tao Ge, Tianyu Liu, Wenjie Li, and Zhifang Sui. 2024. Unlocking efficiency in large language model inference: A comprehensive survey of speculative decoding. In Findings of the Association for Computational Linguistics (ACL Findings). Spec-Bench.
  • Xie et al. (2026) Zhinan Xie, Peisong Wang, Shuang Qiu, and Jian Cheng. 2026. HiViS: Hiding visual tokens from the drafter for speculative decoding in vision-language models. In CVPR Findings.
  • Yin et al. (2024) Ming Yin, Minshuo Chen, Kaixuan Huang, and Mengdi Wang. 2024. A theoretical perspective for speculative decoding algorithm. In Advances in Neural Information Processing Systems (NeurIPS). ArXiv:2411.00841.
  • Zhang et al. (2026a) Jiebin Zhang, Zhenghan Yu, Song Liu, Eugene J. Yu, Zheng Li, Dawei Zhu, Jiangshan Duo, Weimin Xiong, Yifan Song, Guanghua Yu, and 1 others. 2026a. DFlare: Scaling up draft capacity for block diffusion speculative decoding. Preprint, arXiv:2606.02091.
  • Zhang et al. (2026b) Kewei Zhang, Jin Wang, Sensen Gao, Chengyue Wu, Yulong Cao, Songyang Han, Boris Ivanovic, Langechuan Liu, Marco Pavone, Song Han, Daquan Zhou, and Enze Xie. 2026b. Fast-dDrive: Efficient block-diffusion VLM for autonomous driving. Preprint, arXiv:2605.23163.
  • Zhang et al. (2026c) Shuai Zhang, Huachuan Qiu, Hongliang He, and Yong Dai. 2026c. Cost-Aware Diffusion Draft Trees for Speculative Decoding. Preprint, arXiv:2606.01813.
  • Zhang et al. (2026d) Yaojie Zhang, Linfeng Zhang, Bin Cui, and Xupeng Miao. 2026d. DFlow: Enabling verifier information flow in block diffusion speculative decoding. Preprint, arXiv:2609.06498.

Appendix A Theoretical Statements and Proofs

A.1 Losslessness

Theorem 1 (full).

For the decoder of Definition 1, assume (float-exactness) that at every visited prefix x∘b∘y1:kx\circ b\circ y_{1:k} the target arg​max\argmax is unique and the top logit gap γk:=ℓ(1)−ℓ(2)>0\gamma_{k}:=\ell_{(1)}-\ell_{(2)}>0 exceeds the packed-versus-unpacked logit perturbation. Then for any candidate tree the committed token sequence equals the target’s autoregressive greedy sequence exactly, and losslessness can fail only at a margin γk\gamma_{k} below that perturbation, that is, at a near-tie.

Proof.

Single round. By Definition 1 the verifier row at any prefix x∘b∘y1:kx\circ b\circ y_{1:k} reproduces p(⋅∣x∘b∘y1:k)p(\cdot\mid x\circ b\circ y_{1:k}) exactly, since ancestor-only masking and token packing change the attention layout, not the conditioning of a row. Let Yk+1Y_{k+1} be the target greedy token there. If y1:k∘Yk+1∈Sy_{1:k}\circ Y_{k+1}\in S, the walk descends and commits Yk+1Y_{k+1}, and otherwise it stops and emits Yk+1Y_{k+1} as the next root. Either way the token produced at that position is Yk+1Y_{k+1}, and the tree affects only how many tokens the round commits.

Induction over rounds. The emitted stream concatenates accepted paths and corrections, and each round’s correction is the next round’s root. Every produced token is the target greedy token at its own prefix, so induction on the output position gives the committed sequence Y1Y2⋯Y_{1}Y_{2}\cdots bitwise, independent of all budget choices.

Tie sensitivity. The only step that can break is the equality of the packed and the autoregressive arg​max\argmax. It fails only when γk\gamma_{k} falls below the packed-versus-unpacked logit perturbation, that is, at a near-tie. In fp32 the perturbation is below every observed margin, and the fp32 audit shows zero divergences. ∎

A.2 Derivation of the Law

Assumption 2 (Offset survival with marginal hazards).

Conditioned on the round, or equivalently on HH, let the drafter’s top-11 token match the target greedy token at offset jj with probability pm,jp_{\mathrm{m},j}, and let aa be the leading run of matches before the first miss, capped at LL. We assume the survival factorizes into the marginal hazards,

ℙ⁡[a≥ℓ∣H]\displaystyle\mathbb{P}[a\geq\ell\mid H] =∏j≤ℓpm,j,\displaystyle=\prod_{j\leq\ell}p_{\mathrm{m},j}, (3)
𝔼⁡[a∣H]\displaystyle\mathbb{E}[a\mid H] =∑ℓ=1L∏j≤ℓpm,j.\displaystyle=\sum_{\ell=1}^{L}\prod_{j\leq\ell}p_{\mathrm{m},j}.

This is an assumption, not a consequence of the chain rule, since it replaces the conditional hazards ℙ[matchj∣match<j,H]\mathbb{P}[\mathrm{match}_{j}\mid\mathrm{match}_{<j},H] by the marginals pm,jp_{\mathrm{m},j}.

Lemma 1 (Slowly varying hazard gives a truncated-geometric mean).

Under Assumption 2 and the further approximation pm,j≈pmp_{\mathrm{m},j}\approx p_{\mathrm{m}} for all j≤Lj\leq L, Equation (3) reduces to 𝔼⁡[a∣H]≈∑ℓ=1Lpmℓ=pm​(1−pmL)/(1−pm)\mathbb{E}[a\mid H]\approx\sum_{\ell=1}^{L}p_{\mathrm{m}}^{\ell}=p_{\mathrm{m}}(1-p_{\mathrm{m}}^{L})/(1-p_{\mathrm{m}}), which tends to pm/(1−pm)p_{\mathrm{m}}/(1-p_{\mathrm{m}}) as L→∞L\to\infty. The untruncated form is an upper bound for every finite LL, with relative truncation error pmLp_{\mathrm{m}}^{L}.

Proof.

Setting pm,j≡pmp_{\mathrm{m},j}\equiv p_{\mathrm{m}} in Equation (3) gives the finite geometric series. Its truncated tail is

∑ℓ>Lpmℓ=pmL⋅pm1−pm,\sum_{\ell>L}p_{\mathrm{m}}^{\ell}=p_{\mathrm{m}}^{L}\cdot\frac{p_{\mathrm{m}}}{1-p_{\mathrm{m}}},

so the relative truncation error is pmLp_{\mathrm{m}}^{L}. At L=15L{=}15, the error pmLp_{\mathrm{m}}^{L} is 0.0020.002, 0.0350.035, and 0.2060.206 at pm=0.65p_{\mathrm{m}}=0.65, 0.800.80, and 0.900.90, so the closed form is tight except near the entropy floor.

Two errors are folded in, the truncation tail and the slowly varying step itself. The step overstates 𝔼⁡[a∣H]\mathbb{E}[a\mid H] when pm,jp_{\mathrm{m},j} decays in jj. It understates the mean at extremely low entropy when one pmp_{\mathrm{m}} is fitted across a range that contains a sharp spike at low HH. The supremum of the DocVQA fit is 3.343.34 at b0=1.226b_{0}{=}1.226 and L=15L{=}15, below the measured DocVQA bottom-decile mean of 5.235.23, which Theorem 2 explains.

Substituting pm​(H)=σ⁡(b0−b1​H)p_{\mathrm{m}}(H)=\sigma(b_{0}-b_{1}H) gives Equation (2), which we use as an operational characterization of the conditional mean, not as an exact survival identity. ∎

Proposition 1 (Entropy controls the match, and the affine logit is a surrogate).

Let pmp_{\mathrm{m}} be the probability that the drafter’s top-11 token at the root equals the target greedy token, and HH the target’s root entropy. (i) The target’s own top-11 error e⋆=1−maxw⁡p⁡(w∣x,b)e^{\star}=1-\max_{w}p(w\mid x,b) obeys Fano’s inequality

H≤ℍb​(e⋆)+e⋆​log⁡(|Σ|−1),H\leq\mathbb{H}_{b}(e^{\star})+e^{\star}\log(|\Sigma|-1),

so larger HH forces larger e⋆e^{\star}, and H→0H\to 0 forces maxw⁡p→1\max_{w}p\to 1. (ii) Modeling the drafter’s logit error as logistic with scale ss gives pm=σ⁡(Δ/s)p_{\mathrm{m}}=\sigma(\Delta/s) for the target’s top-22 margin Δ⁡(H)\Delta(H), which decreases in HH by (i). (iii) A first-order expansion of Δ\Delta in HH gives logit⁡pm≈b0−b1​H\operatorname{logit}p_{\mathrm{m}}\approx b_{0}-b_{1}H with b1=−Δ′(H0)/s>0b_{1}=-\Delta^{\prime}(H_{0})/s>0. Part (i) predicts sign and monotonicity, and parts (ii) and (iii) are modeling steps that select the functional form.

Proof.
  1. (i)

    View the target’s top-11 token as an estimator of a draw W∼pW\sim p. Fano’s inequality bounds HH by an increasing function of the error e⋆e^{\star} on [0,1−1/|Σ|][0,1-1/|\Sigma|], so e⋆e^{\star} grows with HH and vanishes as H→0H\to 0.

  2. (ii)

    The drafter preserves the target’s ordering if and only if the perturbed margin stays positive. A difference of logistics is logistic, so

    logit⁡pm=Δ/s,\operatorname{logit}p_{\mathrm{m}}=\Delta/s,

    and on a shell of fixed entropy Δ\Delta decreases in HH by (i).

  3. (iii)

    This is the first-order expansion. Its adequacy over the empirical window is what the curve R2R^{2} measures.

A fuller derivation under a temperature family, in which the target’s logits are a fixed shape scaled by an inverse temperature, shows that logit⁡p(1)tgt​(H)\operatorname{logit}p^{\mathrm{tgt}}_{(1)}(H) is affine to R2=0.84R^{2}=0.84 to 0.900.90 over the resolvable window H∈[0.05,1.2]H\in[0.05,1.2] for Zipf, geometric-gap, and near-binary logit shapes, with self-confidence slopes κ≈3.2\kappa\approx 3.2 to 3.93.9.

This reduces the match slope to a single drafter-noise scale ss through b1=κ/sb_{1}=\kappa/s. Matching the measured offset-11 slopes gives s=3.4s=3.4, 7.77.7, and 3.83.8 on captioning, TextVQA, and DocVQA. The affine form is exactly a local linearization, and since the binary closed form is provably curved globally, the law is stated over the empirical window only. ∎

Corollary (Grounded is more draftable).

Under the untruncated form of Equation (2),

∂Hpm\displaystyle\partial_{H}p_{\mathrm{m}} =−b1​pm​(1−pm),\displaystyle=-b_{1}p_{\mathrm{m}}(1-p_{\mathrm{m}}),
∂H𝔼⁡[a∣H]\displaystyle\partial_{H}\mathbb{E}[a\mid H] =∂Hpm(1−pm)2\displaystyle=\frac{\partial_{H}p_{\mathrm{m}}}{(1-p_{\mathrm{m}})^{2}}
=−b1​𝔼​[a∣H]<0,\displaystyle=-b_{1}\mathbb{E}[a\mid H]<0,

so 𝔼⁡[a∣H]\mathbb{E}[a\mid H] is strictly decreasing in HH, and the truncated form is strictly increasing in pmp_{\mathrm{m}} and hence decreasing in HH as well. Any regime with systematically lower next-token entropy therefore has a longer expected accepted block. Grounded and OCR generation copies determinate glyph and answer strings from the image and thereby collapses the target’s entropy. Grounded spans are therefore more draftable than open text, and the mean accepted length of a task is ordered by grounding.

Proof.

The derivative computation is displayed in the statement and uses only b1>0b_{1}>0 and pm∈(0,1)p_{\mathrm{m}}\in(0,1). The ordering claim uses only the measured monotone decrease of aa in HH, not the fitted closed form. The measured round logs order the mean accepted length by grounding, at 1.881.88 on captioning, 2.122.12 on TextVQA, and 3.083.08 on DocVQA, and the InfoVQA and ChartQA logs extend the ordering to 2.672.67 and 3.713.71, the latter under ChartQA’s templated prompts. The gap between DocVQA and captioning, +1.20+1.20, has a 95%95\% bootstrap interval of [1.07,1.33][1.07,1.33] and a Mann-Whitney p=8.7×10−68p=8.7\times 10^{-68}, and the grounded task carries both the higher b0b_{0} and the steeper b1b_{1}, as Table 12 lists. ∎

A.3 The Noise Ceiling

Proposition 2 (Entropy-only noise ceiling).

Among all σ⁡(H)\sigma(H)-measurable predictors g⁡(H)g(H) of aa, the largest attainable coefficient of determination is R¯2=Var⁡(𝔼⁡[a∣H])/Var⁡(a)=1−𝔼⁡[Var⁡(a∣H)]/Var⁡(a)\bar{R}^{2}=\mathrm{Var}(\mathbb{E}[a\mid H])/\mathrm{Var}(a)=1-\mathbb{E}[\mathrm{Var}(a\mid H)]/\mathrm{Var}(a), attained by the conditional mean. For a truncated-geometric aa the within-round variance is of order 𝔼​[a∣H]2\mathbb{E}[a\mid H]^{2}, so R¯2\bar{R}^{2} is intrinsically small regardless of the predictor’s quality.

Proof.

For any gg, orthogonality of the conditional mean gives

𝔼⁡[(a−g⁡(H))2]\displaystyle\mathbb{E}[(a-g(H))^{2}] =𝔼⁡[Var⁡(a∣H)]\displaystyle=\mathbb{E}[\mathrm{Var}(a\mid H)]
+𝔼⁡[(𝔼⁡[a∣H]−g⁡(H))2],\displaystyle+\mathbb{E}[(\mathbb{E}[a\mid H]-g(H))^{2}],

which is minimized at g=𝔼⁡[a∣H]g=\mathbb{E}[a\mid H], and the law of total variance gives the stated ratio. For the untruncated geometric, with μ=𝔼⁡[a∣H]\mu=\mathbb{E}[a\mid H],

Var⁡(a∣H)=μ⁡(1+μ)≥μ2,\mathrm{Var}(a\mid H)=\mu(1+\mu)\geq\mu^{2},

so the within-round variance dominates unless Var⁡(μ⁡(H))\mathrm{Var}(\mu(H)) is comparably large. A nonparametric estimate of the between-to-total variance over 5050 bins gives ceilings of 0.1100.110, 0.1440.144, and 0.1880.188 on captioning, TextVQA, and DocVQA. The law’s round-level R2R^{2} values of 0.0970.097, 0.0910.091, and 0.0740.074 therefore realize 89%89\%, 63%63\%, and 39%39\% of what any entropy-only predictor could reach, and Table 12 lists the same fractions for InfoVQA and ChartQA. This is why the curve R2R^{2} is the appropriate score. ∎

A.4 The Grounded Tail

Assumption 3 (Two-state determinate or uncertain survival).

Conditioned on the round, the matches along a block follow a two-state hidden Markov chain over sj∈{D,U}s_{j}\in\{D,U\}, with entry distribution (π,1−π)(\pi,1-\pi), match probability ρs\rho_{s} in state ss, transition matrix TT upon a match, the run ending at the first miss, and a cap at LL. State DD models verbatim copying (ρD→1\rho_{D}\to 1, sticky tD​D>0t_{DD}>0), and UU is the ordinary regime with a moderate hazard. The parameters governing H→0H\to 0 are estimated in regime, from near-zero-entropy rounds only, rather than extrapolated from the affine fit. The geometric law of Lemma 1 is the degenerate case ρD=ρU\rho_{D}=\rho_{U}.

Theorem 2 (full).

Let M⁡(p)=∑ℓ=1LpℓM(p)=\sum_{\ell=1}^{L}p^{\ell}. (A) Under the single-hazard law of Equation (2), supH≥0𝔼⁡[a∣H]=M⁡(σ⁡(b0))≤eb0\sup_{H\geq 0}\mathbb{E}[a\mid H]=M(\sigma(b_{0}))\leq e^{b_{0}}, a rigorous cap that needs no fit. (B) Under Assumption 3, with 𝐯=(π,1−π)\mathbf{v}=(\pi,1-\pi), D𝝆=diag⁡(ρD,ρU)D_{\boldsymbol{\rho}}=\operatorname{diag}(\rho_{D},\rho_{U}), and K=T​D𝝆K=TD_{\boldsymbol{\rho}},

ℙ⁡[a≥ℓ∣H]\displaystyle\mathbb{P}[a\geq\ell\mid H] =𝐯​D𝝆​Kℓ−1​𝟏,\displaystyle=\mathbf{v}D_{\boldsymbol{\rho}}K^{\ell-1}\mathbf{1}, (4)
𝔼⁡[a∣H]\displaystyle\mathbb{E}[a\mid H] =𝐯​D𝝆​(∑m=0L−1Km)​𝟏,\displaystyle=\mathbf{v}D_{\boldsymbol{\rho}}\Big(\textstyle\sum_{m=0}^{L-1}K^{m}\Big)\mathbf{1},

whose effective hazard hℓh_{\ell} is not constant. Taking ρD,tD​D→1\rho_{D},t_{DD}\to 1 gives 𝔼​[a∣H]H→0→π​L+(1−π)​M​(ρU)\mathbb{E}[a\mid H]^{H\to 0}\to\pi L+(1-\pi)M(\rho_{U}), which exceeds eb0e^{b_{0}} exactly when π>π⋆:=(eb0−M⁡(ρU))/(L−M⁡(ρU))\pi>\pi^{\star}:=(e^{b_{0}}-M(\rho_{U}))/(L-M(\rho_{U})). This threshold lies in (0,1)(0,1) because M⁡(ρU)<eb0<LM(\rho_{U})<e^{b_{0}}<L, where on DocVQA L=15L{=}15 and eb0=3.41e^{b_{0}}{=}3.41. (C) With in-regime maximum-likelihood parameters, Equation (4) recovers the measured DocVQA floor mean of 5.235.23, against the geometric cap of 3.413.41. Proposition 3 gives the falsifiable validation.

Proof.
  1. (A)

    MM is strictly increasing on (0,1)(0,1), and pm​(H)p_{\mathrm{m}}(H) is maximized as H→0H\to 0, where it equals σ⁡(b0)\sigma(b_{0}). Then M⁡(p)≤p/(1−p)M(p)\leq p/(1-p) gives the bound eb0e^{b_{0}}.

  2. (B)

    Starting from 𝐯\mathbf{v}, the joint event of a match at offset 11 followed by state ss has mass (𝐯​D𝝆)s(\mathbf{v}D_{\boldsymbol{\rho}})_{s}, and each accepted offset multiplies by K=T​D𝝆K=TD_{\boldsymbol{\rho}}, so

    ℙ⁡[a≥ℓ∣H]=𝐯​D𝝆​Kℓ−1​𝟏,\mathbb{P}[a\geq\ell\mid H]=\mathbf{v}D_{\boldsymbol{\rho}}K^{\ell-1}\mathbf{1},

    and the tail-sum identity gives the mean. Survival is a nonnegative combination of the two eigenmodes of KK, so the hazard hℓh_{\ell} moves monotonically toward λmax​(K)\lambda_{\max}(K) and is not constant unless ρD=ρU\rho_{D}=\rho_{U}. Its first value π​ρD+(1−π)​ρU\pi\rho_{D}+(1-\pi)\rho_{U} can approach 11 and exceed σ⁡(b0)\sigma(b_{0}), which the constant-hazard family cannot do. As ρD,tD​D→1\rho_{D},t_{DD}\to 1,

    ℙ[a≥ℓ]→πfor all ℓ≤L,\mathbb{P}[a\geq\ell]\to\pi\quad\text{for all }\ell\leq L,

    which gives the stated excess over the cap.

  3. (C)

    This is the numerical instantiation, which recovers the measured floor mean against the geometric cap.∎

Proposition 3 (Prompt-disjoint held-out falsification).

Take the 122122 bottom-entropy-decile DocVQA rounds, from 4747 prompts, and split them by prompt into disjoint sets. (i) On the full floor the geometric run length is rejected, with Pearson χ2=105.3\chi^{2}=105.3, df=7\mathrm{df}=7, p=8.7×10−20p=8.7\times 10^{-20}, while the likelihood ratio decisively favors the two-state form, with 2​Δ​ℓ=86.52\Delta\ell=86.5, df=4\mathrm{df}=4, p=7.2×10−18p=7.2\times 10^{-18}. (ii) Frozen on one prompt set and scored on the disjoint one, the two-state model attains a smaller held-out χ2\chi^{2} than the geometric in 66 of 77 tests, over both split directions and five cross-validation folds, with a mean held-out χ2\chi^{2} of 4.44.4 against 11.911.9.

Proof.

The proposition records the outcome of the stated estimation procedure, with the construction and statistics as listed. ∎

Appendix B Decoding Round

Algorithm 1 gives one decoding round of GLANCE. The draft pass and the verify pass are the only two forward passes of a round, and the tree budget NN enters only the size of the verify pass.

Algorithm 1 One decoding round of GLANCE
1: cached prefix xx with the target’s fused states, pending root bb, budget NN, horizon LL
2: (q1,…,qL)←BlockHead​(x,b)(q_{1},\dots,q_{L})\leftarrow\textsc{BlockHead}(x,b) ⊳\triangleright one draft pass
3: S←S\leftarrow the NN prefixes y1:ℓy_{1:\ell} of highest ∏j≤ℓqj​(yj)\prod_{j\leq\ell}q_{j}(y_{j})
4: pack bb and SS with an ancestor mask
5: rows ←Target​(x,packed tree)\leftarrow\textsc{Target}(x,\text{packed tree}) ⊳\triangleright one verify pass
6: v←bv\leftarrow b; out←[]\text{out}\leftarrow[\,]
7: loop
8:   t←arg​maxw⁡rows​[v]​(w)t\leftarrow\argmax_{w}\text{rows}[v](w) ⊳\triangleright target greedy token
9:   if child v∘t∈Sv\circ t\in S then
10:    append tt to out; v←v∘tv\leftarrow v\circ t
11:   else
12:    return out, new root tt
13:   end if
14: end loop

Appendix C Implementation Details

Primary drafter.

The primary drafter is a block-diffusion head. It reads the target’s fused hidden states and proposes the whole block in one forward pass, with the vision tower and the target decoder frozen. We use the checkpoint at the end of its one-epoch run, and Table 3 lists its configuration. The remaining constants are an intermediate size of 1228812288, a block horizon L=15L{=}15, an image long side of 768768 pixels, AdamW betas of 0.90.9 and 0.950.95 with no weight decay, gradient clipping at 1.01.0, seed 00, one A6000-class GPU, and bf16 weights of 2.12.1 GB.

Setting Value
Architecture
layers 55
parameters 1.051.05B
block size BB 1616
hidden size 40964096
attention heads 32×12832\times 128
kept target layers {1,9,17,25,33}\{1,9,17,25,33\}
Objective
loss coverage CE, offset weight 0.8j0.8^{j}
reveal variable-mask block prefix
Data
source target greedy, 256256 new tokens
prompts 80008000 COCO ++ 80008000 TextVQA
rows 1616K teacher-forced
Optimization
optimizer, lr AdamW, 4×10−54\times 10^{-5}
epochs, accumulation 11, 88
max sequence length 20482048
warm start Qwen3-8B DFlash head
Table 3: Training card of the primary Qwen3-VL-8B drafter.

Evaluation prompts.

The drafter trains on COCO-Caption2017 val rows 100100 to 80998099 and on TextVQA validation rows 00 to 80058005. The home-target head trains on the first 80008000 rows of the same two splits. Every GLANCE row of the two main tables therefore takes its captioning and TextVQA prompts from the test split of each source, with no overlap with any training row, namely 100100 prompts in each task on Qwen3-VL-8B and 200200 captioning and 120120 TextVQA prompts on the home target, the counts of ViSpec’s own evaluation. In Table 1, all three systems decode these same test prompts. The other baselines of Table 2 draw their prompts from the validation rows on Qwen3-VL-8B, and the released ViSpec head runs its native loaders on its home target. InfographicVQA, DocVQA, and ChartQA appear in no training corpus of the primary head. The analyses of Figures 3 and 4 compare conditions of one head on shared prompts and draw their captioning and TextVQA prompts from the validation rows. Every system of the two main tables on Qwen3-VL-8B, apart from the ViSpec-codebase heads with their official loaders, uses the same prompt form for each task. Captioning prompts ask the target to describe the image in detail, and TextVQA prompts are the dataset questions as written. The DocVQA prompt reads “Read the document image and answer the question. Quote the exact text from the document, then briefly explain.” before the question, and the InfographicVQA and ChartQA prompts ask in the same way for the exact text, numbers, values, or labels the target reads, followed by an explanation, so that these tasks also leave a decode loop to shorten.

Second and third targets.

The Qwen2.5-VL-7B head is trained from scratch with the OCR-augmented recipe on that target’s own greedy generations, with 80008000 caption, 80008000 TextVQA, and 10,18010{,}180 OCR rows from DocVQA, ChartQA, and InfoVQA prompts disjoint from the evaluation prompts, 26,18026{,}180 rows in total, for seven epochs at global batch 4848. It is evaluated with the same budget-6363 tree and fp32 audit as the primary target, with captioning and TextVQA prompts from the test splits. Prompt counts match the ViSpec row task by task, 200200 for captioning and DocVQA and 120120 for TextVQA, which is where ViSpec’s staged image set ends, and the fp32 audit finds zero divergences at eight prompts in each task. The Llama-3.2-11B-Vision adapter is warm-started from the released Llama-3.1-8B DFlash head, whose held-out top-11 accuracy at the first offset climbs from 0.710.71 to 0.780.78.

Matched-training comparison.

Both architectures are trained from scratch under one protocol, with the same 26,19626{,}196-row OCR-augmented corpus, the same frozen target, global batch 2424, three epochs, and SpecForge, and with each method’s default learning rate scaled by the square root of the batch. Each drafter then runs at its standard operating point, ours with a budget-6363 tree and the EAGLE-3 head as a depth-33 chain, greedy, with 6464 new tokens, in one Hugging Face implementation on 256256 held-out prompts. Pooled over the five tasks the acceptance ratio is 2.732.73, at 4.374.37 against 1.601.60, and a depth-77 chain adds 0.010.01 to the EAGLE-3 head. The ALLaVA replication retrains both heads on an ALLaVA-Instruct corpus (Chen et al., 2024a) under the same protocol and gives 2.042.04 against 1.291.29 pooled over the five tasks. Timed in the same implementation on one A100, with 2020 prompts in each task and 128128 new tokens, GLANCE decodes at 2.362.36 to 2.59×2.59\times autoregressive decoding against 1.131.13 to 1.17×1.17\times for the EAGLE-3 head.

The same two heads also run in SGLang under the protocol of Table 1, with the tree of 3232 draft tokens, 101101 prompts in each task after a warm-up prompt, and 256256 new tokens. Each task’s two arms decode back to back on one RTX A6000. The corpus holds DocVQA and InfographicVQA validation rows and ChartQA test rows, so these three tasks take their prompts from the DocVQA and InfographicVQA test splits and the ChartQA validation split under the corpus’s prompt template, and captioning and TextVQA use the test prompts of Table 1. Table 4 reports the result. GLANCE is faster on 7373 to 100100 of the 101101 prompts of each task.

EAGLE-3 head GLANCE
Task τ↑\tau\ \uparrow ms/tok ↓\downarrow τ↑\tau\ \uparrow ms/tok ↓\downarrow GLANCE faster by
Captioning 3.293.29 12.4112.41 3.223.22 11.87\mathbf{11.87} +4.5%\mathbf{+4.5\%} [3.3,5.8][3.3,5.8]
TextVQA 2.852.85 14.3614.36 3.11\mathbf{3.11} 12.50\mathbf{12.50} +14.9%\mathbf{+14.9\%} [12.7,17.2][12.7,17.2]
InfographicVQA 2.782.78 15.0615.06 3.06\mathbf{3.06} 12.96\mathbf{12.96} +16.1%\mathbf{+16.1\%} [14.0,18.3][14.0,18.3]
DocVQA 3.003.00 15.1915.19 3.45\mathbf{3.45} 12.40\mathbf{12.40} +22.5%\mathbf{+22.5\%} [19.0,26.8][19.0,26.8]
ChartQA 3.493.49 11.4311.43 4.13\mathbf{4.13} 9.10\mathbf{9.10} +25.7%\mathbf{+25.7\%} [23.4,28.0][23.4,28.0]
Table 4: Matched-training heads in SGLang under the shared tree of 3232 draft tokens. Brackets give paired bootstrap 95%95\% intervals, and blue marks GLANCE.

The released text head without our training.

GLANCE starts from the released block head for Qwen3-8B (Chen et al., 2026), which has the same configuration and weight shapes. Run unchanged on Qwen3-VL-8B under the protocol of Table 1, with EAGLE3-VL decoding the same prompts back to back on one RTX A6000, it is slower than EAGLE3-VL on every task, and every paired bootstrap interval lies below zero, as Table 5 shows. EAGLE3-VL is the faster arm on at least 6666 of the 101101 prompts of each task. Our training lengthens the head’s accepted blocks by 1717 to 21%21\%.

τ↑\tau\ \uparrow
Task EAGLE3-VL Released GLANCE Released faster by
Captioning 3.513.51 2.402.40 2.922.92 −26.4%-26.4\% [−27.4,−25.3][-27.4,-25.3]
TextVQA 4.314.31 2.962.96 3.573.57 −26.2%-26.2\% [−27.9,−24.4][-27.9,-24.4]
InfographicVQA 3.513.51 3.073.07 3.653.65 −7.3%-7.3\% [−9.2,−5.4][-9.2,-5.4]
DocVQA 3.683.68 3.343.34 3.913.91 −4.4%-4.4\% [−7.0,−1.8][-7.0,-1.8]
ChartQA 4.504.50 3.913.91 4.694.69 −7.8%-7.8\% [−9.6,−5.8][-9.6,-5.8]
Geometric mean −15.0%-15.0\% [−15.8,−14.1][-15.8,-14.1]
Table 5: The released text head for Qwen3-8B, run unchanged on Qwen3-VL-8B in SGLang under the shared tree of 3232 draft tokens, against EAGLE3-VL on the same prompts. Margins follow Table 1, with paired bootstrap 95%95\% intervals. The GLANCE column repeats Table 1. Blue marks GLANCE.

ViSpec’s corpus and training length.

The training rows are the 6868K LLaVA-Pretrain questions and images regenerated under ViSpec’s own data pipeline, with the same source file, shuffle seed, prompt template, pixel bounds, and sampling temperature, and with responses produced by the frozen target, since self-generation is part of our recipe. The head then trains with the published card unchanged, once for one epoch and once for twenty-one, the latter continuing from the one-epoch checkpoint at the same constant learning rate and global batch of eight rows. Our own corpus receives the same treatment. All arms are evaluated with the budget-6363 tree, branching 88, 100100 prompts in each task, 256256 new tokens, and the fp32 audit at eight prompts in each task, back to back on one otherwise idle GPU. Figure 5 evaluates every checkpoint of both runs. Acceptance saturates within about five epochs on either corpus, and twenty-one epochs add 3.4%3.4\% on our corpus and 0.9%0.9\% on ViSpec’s.

Figure 5: Change in pooled acceptance against the one-epoch head, for every checkpoint of the two 2121-epoch runs, read with the budget-6363 tree at 5050 prompts in each task.

ViSpec-codebase heads on our target.

These heads are trained on Qwen3-VL-8B-Instruct with ViSpec’s official codebase and two-stage schedule, namely text-only drafter pre-training on the authors’ ShareGPT split to convergence followed by multimodal fine-tuning on target-generated data, and Medusa is trained through the same codebase. They decode with the codebase’s tree search and acceptance rule under its official tree of 3030 tokens at depth 33 with top-88 branching. We port the codebase to Qwen3-VL-8B so that the target runs its reference forward pass, with multimodal rotary positions and deepstack visual features. Each head decodes the 100100 prompts of each task that the GLANCE row uses, with the same 896896-pixel images, greedy, batch one, and up to 256256 new tokens. The EAGLE-2 row is this recipe without the vision adaptor, and the ViSpec row is the same schedule with the adaptor, both trained from the same stage-one checkpoint, data, and seed, so the gap between the rows is the adaptor’s contribution alone.

ViSpec on its home target.

The released ViSpec head for Qwen2.5-VL-7B-Instruct runs with its native loaders and its default tree of 3030 tokens at depth 33, with top-88 branching and two query heads, greedy, batch one, and up to 10241024 new tokens. ViSpec logs the bonus-excluded length, which we raise by the committed token.

Other protocol details.

Tables 1 and 2 run the same GLANCE head, the first in SGLang with images at native resolution and a tree of 3232 draft tokens, the second in our Hugging Face implementation at 896896 pixels with the budget-6363 tree, so their acceptance lengths differ. In Table 1 the three arms of each task decode back to back on one RTX A6000, and GLANCE’s draft pass replays from a CUDA graph over a static key-value cache, with one graph for each number of tokens committed in the previous round. The mean entropy H¯\bar{H} of Table 1 is the target’s mean root entropy over the rounds of the law fits in Section 5.5. EAGLE3-VL, the production AQ-MedAI checkpoint for this target, and the n-gram baseline run in vLLM 0.22.10.22.1, with 2020 prompts in each task and the warm-up sample excluded, and EAGLE3-VL runs as a chain of 55, which Table 7 sweeps from 33 to 1515. Prompt lookup proposes a draft only when an n-gram match exists, so its acceptance length is the length of the blocks it does propose. The classic two-model rows are the endpoints of the three-draft sweep of Table 11. The mismatched-image probe uses 6060 prompts in each task, with the three image conditions scored on a shared teacher-forced trajectory and prompt-clustered bootstrap intervals. Speedup is the ratio of mean decode time for a token over prompts.

Output equivalence.

Our implementation decodes 2020 prompts in each of three tasks, and each system’s generated tokens are compared position by position with the target’s own greedy decoding of the same prompts, in the same arithmetic. Table 6 reports the result. Decoding the target twice with no drafter shows that bf16 identity measures kernel arithmetic rather than any drafter, so the guarantee is decided in fp32, where both the drafter-free control and GLANCE reach 6060 of 6060. The ViSpec-codebase heads and the released ViSpec head, compared the same way on 2121 prompts in each task with up to 512512 new tokens, reproduce all 6363 in fp32, as exact acceptance requires.

System Arithmetic Identical
Our implementation
autoregressive, no drafter bf16 25/6025/60
autoregressive, no drafter fp32 60/6060/60
GLANCE, budget-6363 tree bf16 27/6027/60
GLANCE, budget-6363 tree fp32 𝟔𝟎/𝟔𝟎\mathbf{60/60}
GLANCE, matched-training head bf16 33/6033/60
EAGLE-3 head, matched training bf16 19/6019/60
ViSpec codebase
EAGLE-2, ViSpec codebase fp32 63/6363/63
ViSpec, official recipe fp32 63/6363/63
Medusa, ViSpec codebase fp32 63/6363/63
ViSpec, released head, home target fp32 63/6363/63
Table 6: Output equivalence with the target’s greedy decoding, with both arms in the stated arithmetic. Blue marks GLANCE.

Sampling.

At temperature 11, with a top-pp of 11, no top-kk cut, and the protocol of Table 1 otherwise unchanged, GLANCE is faster than EAGLE3-VL on the three lower-entropy tasks, by 7.5%7.5\% on InfographicVQA, 14.9%14.9\% on DocVQA, and 9.9%9.9\% on ChartQA. Every paired bootstrap interval lies above zero, GLANCE is the faster arm on at least 7373 of the 101101 prompts of each of these tasks, and the geometric-mean margin over the five tasks is +2.0%+2.0\%, with a 95%95\% interval from 0.80.8 to 3.2%3.2\%. Both arms verify with SGLang’s tree sampling, which commits draws from the target’s own distribution.

acceptance length τ\tau ms/token
Chain length Cap. TVQA Doc Cap. TVQA Doc
33 2.17\mathbf{2.17} 2.472.47 2.362.36 6.776.77 6.786.78 7.467.46
55 2.102.10 2.65\mathbf{2.65} 2.56\mathbf{2.56} 7.547.54 6.936.93 7.627.62
77 1.971.97 2.432.43 2.502.50 8.638.63 8.428.42 8.168.16
1010 1.881.88 2.242.24 2.372.37 9.939.93 9.329.32 9.199.19
1515 1.801.80 2.232.23 2.332.33 11.9611.96 10.6810.68 10.6810.68
Table 7: EAGLE3-VL chain-length sweep in vLLM, end-to-end times, 1919 prompts in each task after the warm-up sample. Table 2 uses length 55.

Appendix D Reproducibility

Section 3 and Appendix C specify the drafter’s architecture, training data, objective, and hyperparameters, and Appendix B gives one decoding round as pseudocode. Appendix A contains the complete proofs of all theoretical statements. Every measurement states its prompt count, and Appendix E lists each wall-clock comparison with its engine and sample size. Our decoding code, exactness audit, and law-fitting code are available at https://github.com/js-lee-AI/GLANCE, and we will release the drafter heads, the training code, and the round-level logs behind the entropy-law fits.

Appendix E Additional Wall-Clock Results

Where the cost sits.

Across the five tasks of Table 1, prefill including the visual encoder takes 33 to 21%21\% of the end-to-end time of autoregressive decoding, the most on DocVQA, whose answers are the shortest, and the decode loop takes the remainder. The share falls as the output grows, since prefill is paid once and the loop on every token. Counting prefill, GLANCE remains faster than EAGLE3-VL on the three lower-entropy tasks, by 7.67.6 to 11.1%11.1\%, and every paired bootstrap interval of these margins lies above zero.

Intervals behind the production-engine comparison.

The margin column of Table 1 is EAGLE3-VLms/GLANCEms−1\text{EAGLE3-VL}_{\text{ms}}/\text{GLANCE}_{\text{ms}}-1, the ratio of the two printed speedups. Its interval is a paired bootstrap over the 101101 shared prompts of that statistic with 80008000 resamples. The intervals are [−13.1,−10.6][-13.1,-10.6] on captioning, [−13.1,−9.3][-13.1,-9.3] on TextVQA, [+8.4,+12.8][+8.4,+12.8] on InfographicVQA, [+8.5,+13.6][+8.5,+13.6] on DocVQA, and [+8.7,+13.1][+8.7,+13.1] on ChartQA. GLANCE is the faster arm on 8585, 8181, and 9090 of the 101101 prompts of the three lower-entropy tasks. The same bootstrap, resampling the prompts of every task at once, puts the geometric-mean margin of +1.3%+1.3\% in [+0.3,+2.2][+0.3,+2.2]. Averaging prompt-level ratios instead of taking the ratio of means moves each figure by at most 0.90.9 points and reorders nothing. The autoregressive arm stays within 2%2\% of its mean across the five tasks, so the tasks differ in what they generate rather than in what a token costs to decode.

EAGLE3-VL’s released configuration.

The model card of EAGLE3-VL lists an SGLang launch configuration of three draft steps, top-22 branching, and four draft tokens (AQ-MedAI, 2026). Run under the protocol of Table 1, back to back with the eight-step tree on one RTX A6000 for each task, it accepts 2.212.21 to 2.712.71 tokens a round against 3.513.51 to 4.504.50, and the eight-step tree decodes faster on every task, by 25.3%25.3\% on captioning, 32.9%32.9\% on TextVQA, 18.9%18.9\% on InfographicVQA, 21.0%21.0\% on DocVQA, and 30.0%30.0\% on ChartQA. Every paired bootstrap interval lies above zero, the tree is the faster arm on at least 9898 of the 101101 prompts of each task, and the geometric-mean margin is +25.5%+25.5\%, with a 95%95\% interval from 24.624.6 to 26.5%26.5\%.

Training seeds.

Two further heads were trained with the configuration of Table 3 under seeds 11 and 22, with the target’s fused states computed during training on two GPUs rather than precomputed, at the same global batch. Run under the protocol of Table 1, on one RTX A6000 for each task together with EAGLE3-VL and the primary head, both are faster than EAGLE3-VL on the three lower-entropy tasks, by 8.48.4 and 9.0%9.0\% on InfographicVQA, 9.49.4 and 9.2%9.2\% on DocVQA, and 9.09.0 and 9.5%9.5\% on ChartQA, and every paired bootstrap interval lies above zero. On all five tasks their acceptance lengths lie within 4%4\% of the primary head’s.

Longer contexts, tree against chain, and budget.

Table 8 lists the long-context comparisons, on which GLANCE is the faster arm on 3131 of the 4444 prompts at each context length. Table 9 times the tree against the chain, and Table 10 sweeps the verifier budget.

Comparison Context Result [95%95\% CI]
GLANCE vs. AR 22K 2.30×2.30\times [2.22,2.39][2.22,2.39]
EAGLE3-VL vs. AR 22K 2.15×2.15\times [2.07,2.24][2.07,2.24]
GLANCE vs. EAGLE3-VL 22K +6.9%+6.9\% [+3.4,+10.8][+3.4,+10.8]
GLANCE vs. AR ∼6.7{\sim}6.7K 1.66×1.66\times [1.61,1.73][1.61,1.73]
EAGLE3-VL vs. AR ∼6.7{\sim}6.7K 1.58×1.58\times [1.51,1.64][1.51,1.64]
GLANCE vs. EAGLE3-VL ∼6.7{\sim}6.7K +5.5%+5.5\% [+2.0,+9.1][+2.0,+9.1]
Table 8: Decode-only wall-clock comparisons in SGLang on 4444 full-resolution InfographicVQA prompts, by context length. Percentages are the GLANCE-faster-by margin of Table 1, with paired bootstrap intervals over shared prompts.
ms/tok ↓\downarrow
Task AR tree chain tree vs. chain ↑\uparrow
Captioning 18.3018.30 9.45\mathbf{9.45} 12.8812.88 1.36×1.36\times
TextVQA 18.2218.22 8.26\mathbf{8.26} 11.4211.42 1.38×1.38\times
DocVQA 18.1518.15 8.55\mathbf{8.55} 11.3311.33 1.33×1.33\times
Speedup over AR 1.00×1.00\times 2.09×\mathbf{2.09\times} 1.54×1.54\times 1.36×\mathbf{1.36\times}
Table 9: Decode time of the identical GLANCE head run as a budget-6363 tree and as a width-11 chain in our Hugging Face implementation. The last row gives geometric means. Blue marks the tree.
acceptance length τ↑\tau\ \uparrow speedup over AR ↑\uparrow
Budget NN 1515 3131 4747 6363 1515 3131 4747 6363
Captioning 2.702.70 2.892.89 2.992.99 3.043.04 1.871.87 1.991.99 2.05\mathbf{2.05} 2.05\mathbf{2.05}
TextVQA 2.952.95 3.143.14 3.253.25 3.293.29 1.951.95 2.212.21 2.242.24 2.25\mathbf{2.25}
InfographicVQA 3.213.21 3.433.43 3.643.64 3.633.63 2.222.22 2.392.39 2.46\mathbf{2.46} 2.392.39
DocVQA 3.353.35 3.613.61 3.683.68 3.783.78 2.312.31 2.492.49 2.512.51 2.55\mathbf{2.55}
ChartQA 4.434.43 4.814.81 5.005.00 5.165.16 3.073.07 3.303.30 3.403.40 3.42\mathbf{3.42}
Table 10: Verifier-budget sweep of GLANCE in our Hugging Face implementation. All four budgets of a task run in one process and share its autoregressive baseline.

Round costs.

A GLANCE round is one draft pass and one verify pass, and tree width enters only inside the verify. Regressing the production head’s round time in SGLang on its number of draft passes, over five configurations from 44 to 4848 draft tokens at 33 to 1010 steps on captioning, gives 24.824.8 ms plus 1.821.82 ms for every sequential draft pass. At batch one the verify is therefore nearly flat in tree width, while depth is paid one pass at a time. Widening our tree 4.2×4.2\times costs 33 to 5%5\% of a round, whereas taking that head from two passes to eight costs 38%38\% of one.

Appendix F Classic Two-Model Speculative Decoding

Classic two-model speculation runs on our target through the official Hugging Face assisted-generation path with three drafts, the image-conditioned Qwen3-VL-4B and Qwen3-VL-2B and a text-only Qwen3-1.7B in the spirit of the strong baseline of the first multimodal study (Gagrani et al., 2024). Table 11 reports all three. The 44B draft accepts long blocks, with τ\tau from 3.533.53 to 4.474.47 rising with grounding, yet decodes at 0.610.61 to 0.80×0.80\times, since a half-size draft costs roughly half a target forward for each drafted token and no acceptance amortizes that. The 22B draft accepts less and stays below 1×1\times throughout, so halving the draft again narrows the deficit without closing it. The text-only draft shows what vision access is worth, since its acceptance collapses to 1.421.42 to 2.122.12. Unlike the text-only states of Figure 3, which the block head still reads, this external draft reads nothing of the target at all.

Draft Task nn τ↑\tau\,\uparrow [95%95\% CI] Speedup over AR ↑\uparrow [CI]
Qwen3-VL-4B Captioning 100100 3.533.53 [3.42,3.643.42,3.64] 0.61×0.61\times [0.60,0.630.60,0.63]
TextVQA 100100 3.793.79 [3.65,3.933.65,3.93] 0.67×0.67\times [0.65,0.690.65,0.69]
InfoVQA 100100 3.953.95 [3.74,4.173.74,4.17] 0.70×0.70\times [0.68,0.720.68,0.72]
DocVQA 100100 4.394.39 [4.19,4.604.19,4.60] 0.76×0.76\times [0.74,0.780.74,0.78]
ChartQA 100100 4.474.47 [4.31,4.634.31,4.63] 0.80×0.80\times [0.79,0.820.79,0.82]
Qwen3-VL-2B Captioning 100100 3.113.11 [3.01,3.213.01,3.21] 0.66×0.66\times [0.64,0.680.64,0.68]
TextVQA 100100 3.263.26 [3.15,3.393.15,3.39] 0.72×0.72\times [0.70,0.740.70,0.74]
InfoVQA 100100 3.213.21 [3.07,3.353.07,3.35] 0.74×0.74\times [0.72,0.760.72,0.76]
DocVQA 100100 3.583.58 [3.43,3.743.43,3.74] 0.84×0.84\times [0.82,0.870.82,0.87]
ChartQA 100100 3.973.97 [3.81,4.133.81,4.13] 0.90×0.90\times [0.88,0.920.88,0.92]
Qwen3-1.7B, text-only Captioning 100100 1.491.49 [1.46,1.511.46,1.51] 0.36×0.36\times [0.35,0.380.35,0.38]
TextVQA 100100 1.421.42 [1.40,1.441.40,1.44] 0.32×0.32\times [0.31,0.340.31,0.34]
InfoVQA 100100 1.901.90 [1.84,1.971.84,1.97] 0.47×0.47\times [0.45,0.490.45,0.49]
DocVQA 100100 1.681.68 [1.64,1.721.64,1.72] 0.43×0.43\times [0.41,0.450.41,0.45]
ChartQA 100100 2.122.12 [2.05,2.202.05,2.20] 0.58×0.58\times [0.56,0.610.56,0.61]
Table 11: Classic two-model speculative decoding on Qwen3-VL-8B with three drafts, all on the same prompts and protocol, with speedups from the same run.

Appendix G Additional Law and Transfer Results

Fits on all five tasks.

Table 12 lists the entropy-acceptance fits, and Figure 6 draws them. InfoVQA is logged under the same streaming protocol as captioning, TextVQA, and DocVQA, while ChartQA is logged under its templated evaluation prompts and carries the steepest slope of the five. Both floors clear their own caps, 4.424.42 against 3.223.22 and 6.926.92 against 4.504.50, and the share of rounds accepting eight or more tokens rises with grounding, at 0.00.0, 0.90.9, 2.62.6, 4.44.4, and 11.5%11.5\% from captioning to ChartQA.

Task Rounds b0b_{0} b1b_{1} Rpar2R^{2}_{\mathrm{par}} Riso2R^{2}_{\mathrm{iso}} Rrnd2R^{2}_{\mathrm{rnd}} Ceiling fraction Mean aa
Captioning 40264026 0.8190.819 0.4590.459 0.9490.949 0.990\mathbf{0.990} 0.0970.097 89%89\% 1.8811.881
TextVQA 25072507 0.9360.936 0.5780.578 0.7040.704 0.968\mathbf{0.968} 0.0910.091 63%63\% 2.1202.120
InfoVQA 18581858 1.1691.169 0.8040.804 0.6790.679 0.979\mathbf{0.979} 0.1280.128 62%62\% 2.6722.672
DocVQA 12181218 1.2261.226 0.9540.954 0.4520.452 0.964\mathbf{0.964} 0.0740.074 39%39\% 3.0763.076
ChartQA, templated 21902190 1.5041.504 1.7101.710 0.5430.543 0.991\mathbf{0.991} 0.1380.138 50%50\% 3.7093.709
Pooled, captioning, TextVQA, DocVQA 77517751 0.9640.964 0.6300.630 0.7050.705 0.9170.917 0.1070.107 62%62\% 2.1462.146
Table 12: Entropy-acceptance fits, with the affine-logit coefficients, the parametric, isotonic, and round-level R2R^{2}, the fraction of the entropy-only ceiling realized, and the mean accepted length.
Figure 6: Fitted entropy-acceptance curves for each task.

Brackets.

Figure 7 places each task’s fitted law between two independent bounds, the AdaEDL lower bound (Agrawal et al., 2024) and the verification ceiling 2​log⁡N/H2\log N/H at N=31N{=}31, and every fit stays inside both.

Figure 7: Fitted law of each task, drawn against root entropy HH on a log scale, with the AdaEDL lower bound, the verification ceiling 2​log⁡N/H2\log N/H at N=31N{=}31, and measured entropy deciles as points.

Entropy against image saliency.

Table 13 gives the correlations and decile gaps behind the law. Entropy correlates with acceptance on every task, while an image-saliency signal vtv_{t} does not survive conditioning on entropy, with partial |r|≤0.09|r|\leq 0.09 and gaps near 11 within entropy quintiles.

Quantity Captioning TextVQA DocVQA
Spearman ρ⁡(H,a)\rho(H,a) −0.311-0.311 −0.351-0.351 −0.384-0.384
pp 3.8×10−913.8{\times}10^{-91} 1.5×10−731.5{\times}10^{-73} 4.0×10−444.0{\times}10^{-44}
low/high HH-decile gap 2.30×2.30\times 2.49×2.49\times 3.04×3.04\times
Spearman ρ⁡(vt,a)\rho(v_{t},a) −0.098-0.098 −0.025-0.025 +0.129+0.129
partial r⁡(vt,a∣H)r(v_{t},a\mid H) −0.058-0.058 −0.026-0.026 +0.088+0.088
within-quintile vtv_{t} gap 1.07×1.07\times 1.07×1.07\times 0.94×0.94\times
Table 13: Correlations and gaps for accepted length aa, entropy HH, and image saliency vtv_{t}. The within-quintile gap splits each entropy quintile at its median vtv_{t}.

Mismatched-image probe.

The correct image lengthens accepted blocks by +3.5%+3.5\% on captioning with a 95%95\% interval of [+1.5,+5.7][+1.5,+5.7], by +6.8%+6.8\% on TextVQA with [+4.2,+9.4][+4.2,+9.4], and by +11.7%+11.7\% on DocVQA with [+7.7,+16.3][+7.7,+16.3]. Once entropy is held fixed, the pooled coefficient of the residual image effect is −0.07-0.07.

Transfer across drafters, targets, and modalities.

Table 14 varies the target, and Figure 8 plots the fits. The two further Qwen targets use same-family two-model drafts, which are autoregressive, so the form holds for both kinds of drafter and is a property of the target. On Llama-3.2-11B-Vision, where vision enters by cross-attention rather than as a token prefix, GLANCE remains lossless and the law holds on captioning, and we state the characterization for prefix-fusion targets. Table 15 extends the form beyond vision to speech recognition, code, chat, and a non-Qwen backbone. Grounded vision-language generation and code follow the determinate-copy route of the two-state analysis, with long verbatim runs, while speech behaves like a uniform lift at every position, and both are limits of the same survival family.

Target Qwen2.5-VL Qwen2-VL Llama-3.2
7B 7B 11B-Vision
Drafter two-model two-model GLANCE
b1b_{1} 0.670.67 0.400.40 0.410.41
Curve R2↑R^{2}\uparrow 0.700.70 0.950.95 0.910.91
ρ⁡(H,a)\rho(H,a) −0.42-0.42 −0.30-0.30 −0.38-0.38
Mean a↑a\uparrow 3.673.67 3.083.08 2.862.86
Table 14: Transfer of the law’s form across targets. The two-model drafts are same-family 33B and 22B models, and the Llama column is fitted at horizon L=9L{=}9 on its captioning rounds. Mean aa excludes the bonus token.
Setting Model pair b0b_{0} b1b_{1} τ↑\tau\,\uparrow
speech (ASR) Qwen2.5-Omni 7B++3B 2.072.07 0.490.49 5.415.41
code Qwen2.5-Coder 7B++0.5B 1.941.94 1.761.76 5.775.77
chat (WildChat) two-model 1.661.66 0.640.64 4.334.33
non-Qwen backbone OLMo-2 13B++7B 1.851.85 0.840.84 5.255.25
Table 15: Form of the law across modalities and on a non-Qwen backbone, with two-model pairs.
Figure 8: Entropy-acceptance fits for three targets (a) and the mean accepted length of each task on Qwen2.5-VL-7B (b).

Conditioning ablation.

Table 16 tabulates the values behind panel (b) of Figure 3 together with the zeroed-state control. Pooled over the five tasks and their 300300 prompts, fused visual conditioning lifts acceptance by +14.1%+14.1\%, with a 95%95\% interval of [+12.8,+15.5]%[+12.8,+15.5]\%, and the grounded tasks carry the lift.

τ↑\tau\uparrow by what the head reads
Task fused VL text-only zeroed text/fused
Captioning 2.92\mathbf{2.92} 2.842.84 1.101.10 0.970.97
TextVQA 3.22\mathbf{3.22} 2.952.95 1.121.12 0.920.92
InfoVQA 3.51\mathbf{3.51} 3.123.12 1.101.10 0.890.89
DocVQA 3.62\mathbf{3.62} 2.912.91 1.101.10 0.800.80
ChartQA 4.80\mathbf{4.80} 4.024.02 1.151.15 0.840.84
Mean 3.61\mathbf{3.61} 3.173.17 1.111.11 0.880.88
Table 16: Conditioning ablation at budget 3131 with 6060 prompts in each task. The text-only column uses a text-only Qwen3-8B’s states at text positions and zeros at image positions. Blue marks GLANCE.