Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models
Abstract
Speculative decoding accelerates generation without changing its output, but on vision-language models (VLMs) a self-reinforcing cycle holds it back. Because an autoregressive drafter pays a sequential pass for each drafted token, it must stay small and can ill afford to attend to the image at each pass. Prior work therefore compresses or hides the image, leaving the drafter weakest on the text the image determines. We present GLANCE, a one-pass block drafter that breaks this cycle on an unmodified VLM target. Its block-diffusion head drafts a whole block in one forward pass over the target’s already fused vision-language states, reading the multimodal context once, however deep the draft. The target verifies a wide candidate tree in one pass and commits exactly its greedy output. In one production engine at a fixed round budget, GLANCE decodes up to 3.05 times faster than autoregressive decoding and outpaces the production EAGLE3-VL head on average and by about 11% on grounded tasks. An entropy law explains when drafting pays, predicting the longest accepted blocks on grounded tasks, where the target’s next-token entropy is lowest. Our code is available at https://github.com/js-lee-AI/GLANCE.
1 Introduction
Vision-language models (VLMs) increasingly produce long text about an image. They describe photographs, read scanned documents, and pull values off charts (Bai et al., 2025a; Bai et al., 2025b). Generation remains autoregressive, so every output token needs a full forward pass of the target model, and at small batch sizes this decode loop is bound by memory bandwidth rather than compute. The visual encoder, in contrast, runs once during prefill, and in our workloads the decode loop takes most of the end-to-end time. Faster grounded generation therefore has to come from the decode loop.
Speculative decoding shortens this loop without changing the output. A cheap drafter proposes several future tokens, and the target verifies them in one forward pass, committing the longest prefix it agrees with (Leviathan et al., 2023; Chen et al., 2023). Modern drafters are lightweight heads that read the target’s own hidden features and propose a tree of candidates (Cai et al., 2024; Li et al., 2024b; Li et al., 2024a; Li et al., 2025; Miao et al., 2024), and vision-language serving stacks now ship such heads for VLM targets (AQ-MedAI, 2026).
Adapting this recipe to VLMs, however, has led into a cycle. The drafter is autoregressive, so a candidate tokens deep costs sequential draft passes, and the drafter is kept small so that those passes stay cheap. Each pass also attends over the whole multimodal context, in which image tokens usually outnumber text tokens, and a drafter that looks at the image pays for it again at every step. Heads that keep the full context, such as the production EAGLE3-VL head (AQ-MedAI, 2026), therefore keep their drafts shallow, and methods built for VLMs compress, prune, or hide the image tokens from the drafter (Gagrani et al., 2024; Kang et al., 2025; Huang et al., 2025a; Xie et al., 2026). In both cases drafts stay short, and in the second the drafter also loses the image evidence that grounded text depends on.
On grounded workloads, however, the image makes drafting work. When a VLM answers a question about a document, much of its output already appears on the page, and the next token is often nearly deterministic and frequently copied verbatim, as Figure 1 illustrates. We make this precise with an entropy law, under which the expected accepted length falls with the target’s next-token entropy, and grounded tasks concentrate at the low-entropy end.
To exploit this predictability, we propose long and wide candidate sets in a single draft pass. We present GLANCE (grounded block drafting with one-pass candidate expansion). A block-diffusion head (Chen et al., 2026; Arriola et al., 2025) reads the frozen target’s fused vision-language states and, in a single forward pass, produces a distribution for every position of a block. The highest-scoring prefixes form a wide candidate tree, and the target verifies all of them in one pass. Because the head has already paid for the whole block, width costs no further draft passes (Ringel and Romano, 2026; Zhang et al., 2026c). The committed text is exactly the target’s greedy output, which we prove and also verify in fp32 on every audited prompt. One-pass block drafting has so far been demonstrated only on text-only language models, while lossless drafters built for VLMs have remained autoregressive, and GLANCE is the first one-pass block drafter that is lossless on an unmodified VLM target.
Inside one production engine at a fixed round budget, GLANCE decodes up to faster than autoregressive decoding on chart question answering and outpaces the production EAGLE3-VL head on average and by about on every grounded task, using one draft pass where that head uses eight. Trained on the same corpus and schedule as an EAGLE-3 head and run under the same tree, it decodes to faster on all five tasks, and it keeps its lead when retrained on the corpus of ViSpec, the strongest published VLM drafter. Moreover, the form of the entropy law carries over to other VLM targets and to speech, code, and chat.
This paper makes three contributions.
- •
We diagnose why speculative decoding has underdelivered on VLMs. Autoregressive drafting and reduced vision access reinforce each other, and together they forgo the long verbatim runs of grounded generation, which a one-pass drafter collects in a single pass.
- •
We introduce GLANCE, the first one-pass block drafter that is lossless on an unmodified VLM target. Its head drafts a whole block in one pass from the target’s fused vision-language states, a wide tree drawn from that pass is verified in one target pass, and we build it into SGLang, where it runs against the production head under one engine and one round budget.
- •
We establish an entropy law of draftability and test it with a mismatched-image intervention and a grounded-tail analysis on five tasks, three VLM targets, and four settings beyond vision.
2 Related Work
Draft heads and tree verification.
Speculative decoding accepts drafted tokens with a rejection rule that provably preserves the target distribution (Leviathan et al., 2023; Chen et al., 2023). The practical line drafts from the target’s own hidden features with a lightweight head. Medusa attaches independent heads for each future offset (Cai et al., 2024), Hydra makes them sequentially dependent (Ankner et al., 2024), and EAGLE drafts autoregressively at the feature level over dynamic candidate trees (Li et al., 2024b; Li et al., 2024a), which EAGLE-3 extends with training-time test and multi-layer feature fusion (Li et al., 2025). Many candidates can be verified at once by packing a prefix tree into a single ancestor-masked target pass (Miao et al., 2024), and later work studies tree shape (Chen et al., 2024b; Wang et al., 2025a), theory, and benchmarking (Yin et al., 2024; Huang et al., 2025b; Xia et al., 2024).
Speculative decoding for VLMs.
The first study of speculative decoding for multimodal models found a text-only drafter to be a strong baseline (Gagrani et al., 2024). Subsequent work follows two routes. One shrinks what the drafter reads, by compressing image tokens into a few adaptor embeddings (Kang et al., 2025), pruning visual tokens (Huang et al., 2025a), pruning video tokens under verifier guidance (Ji et al., 2025), or hiding them from the drafter altogether (Xie et al., 2026). The other gives a small drafter its own cheap path to vision, through multimodal distillation (Ganesan et al., 2025), dynamic draft trees (Huo et al., 2025), or refined target features fused by cross-attention (Hu et al., 2025). DREAM uses entropy to weight that fusion inside an autoregressive drafter, whereas we use entropy to predict acceptance and draft in one pass. Semi-autoregressive multimodal drafting relaxes the one-token-a-pass constraint but gives up exactness (Wang et al., 2025b), and a benchmark now covers the setting (Shen et al., 2026). All of these methods keep an autoregressive drafter and engineer its access to vision, whereas GLANCE inherits vision from the target and drafts in one pass.
Block drafting and diffusion drafters.
A parallel line removes the sequential bottleneck inside the drafter. It includes lookahead embeddings (Monea et al., 2023), Jacobi decoding (Fu et al., 2024), block-parallel discrete diffusion language models (Nie et al., 2025; Arriola et al., 2025), diffusion drafters verified by an autoregressive target (Christopher et al., 2025; Li et al., 2026a; Cheng et al., 2025), and DFlash, a block-diffusion head on the frozen target’s hidden states (Chen et al., 2026). DDTree and CaDDTree show that one-pass drafting makes tree width cheap for text-only targets (Ringel and Romano, 2026; Zhang et al., 2026c), and later text-only work builds on such drafters (Zhang et al., 2026a; Wu et al., 2026b; Zhang et al., 2026d; Wang et al., 2026; Li et al., 2026c; Kwon et al., 2026). Block drafting has reached VLMs only by converting the target itself into a self-speculating diffusion model (Wu et al., 2026a; Zhang et al., 2026b), which changes the model’s output. GLANCE instead keeps the target frozen and its output exact.
3 GLANCE
Figure 2 shows one decoding round of GLANCE. A block head drafts a whole block of future tokens in one forward pass, and the target verifies a wide tree of those candidates in one forward pass and commits its own greedy prefix. Only the head is trained, and the target stays frozen.
3.1 Speculative Decoding and Acceptance
Let be the frozen target’s next-token distribution given the multimodal prefix , which holds the text tokens and the encoded image. Greedy decoding emits , one token for each target forward pass. Speculative decoding instead proceeds in rounds (Leviathan et al., 2023; Chen et al., 2023). Given the committed prefix and a pending root token , a drafter proposes candidate continuations, and the target commits the longest prefix that matches its own greedy continuation after , where . When nonroot tokens are accepted, the round commits tokens, namely the accepted prefix and one more token read off the last verified distribution. We report the mean acceptance length
| (1) |
so autoregressive decoding has . For systems that log instead, we add the committed token so that all rows share one scale. A larger means fewer target passes for each generated token, and the wall-clock speedup is discounted by the cost of drafting and verification.
3.2 One-Pass Drafting on Fused States
GLANCE drafts with a block-diffusion head (Chen et al., 2026) of layers, B parameters, and block size , attached to a frozen Qwen3-VL-8B target. The first block position holds the pending root and the other positions are masked, so one forward pass returns a marginal over the vocabulary at every offset with . The head never reads raw visual tokens. Its cross-attention keys and values are the target’s hidden states at five layers spread over the depth of the stack, (Li et al., 2025). By the time these states exist, the target has already merged the image into its text representation, so the drafter inherits visual grounding without an encoder, an adaptor, or compressed image tokens of its own.
Assumption 1 (One-pass marginal interface).
Conditioned on the cached prefix and the pending root , the drafter returns one marginal at every offset in a single forward pass, before any token of the block is committed. A candidate prefix is ranked by its plug-in score .
An autoregressive drafter pays for depth with sequential passes, each attending over the full multimodal context, and the production EAGLE3-VL head takes three passes a round in its released configuration and eight in the production-engine comparison below. The block head pays one pass for the whole block, and the image enters that pass only through states the target has already computed.
3.3 Wide-Tree Verification and Exactness
Definition 1 (Candidate tree and accept rule).
A candidate tree is a prefix-closed set of nonempty strings after the root , of depth at most and budget . Packed into one ancestor-masked target pass, the row of a depth- node yields exactly . With the target’s greedy continuation, the accepted length is . The walk commits and emits the next target-greedy token from the last verified row as the new root.
The tree builder keeps the prefixes with the highest plug-in score , and the kept set is prefix-closed because no prefix scores below its extensions (Ringel and Romano, 2026; Zhang et al., 2026c). The tree is packed into one ancestor-masked target pass (Miao et al., 2024), the decoder walks it along the target’s greedy tokens, and Appendix B gives the complete round as pseudocode. Growing adds verifier work inside that single pass but no draft passes. We use in our Hugging Face implementation and a -node tree in the production engine, where together with the root GLANCE verifies the same draft tokens a round as the production head.
Theorem 1 (Lossless greedy equivalence, informal).
Assume the target’s argmax is unique at every visited prefix and its top- logit gap there exceeds the numerical difference between packed and unpacked evaluation. Then, for any candidate tree, the walk of Definition 1 commits exactly the target’s greedy autoregressive sequence, token for token.
We call a run bitwise identical to greedy decoding when its committed token ids equal those of the target’s own greedy decoding at every position. A weak head can only shorten the accepted prefixes, and no fallback path is needed, because every committed token is the argmax of a target row. We nevertheless measure it instead of assuming it, and the full statement and proof are in Appendix A. In fp32, GLANCE is bitwise identical to greedy decoding on all audited prompts. In bf16, the target does not reproduce its own greedy output across two runs even without any drafter, because kernel choice perturbs near-ties, and GLANCE reproduces that output at least as often as a second run of the target does. Exactness is therefore decided in fp32, and Appendix C reports both audits for every system.
3.4 Training
Only the block head is trained, while the vision tower and the target stay frozen. The training rows are the target’s own greedy generations on COCO-Caption2017 and TextVQA prompts. The objective is top- coverage of the target’s token at every block offset, with offset weighted by , and a reveal scheme exposes a random prefix of the block so that the head learns every reveal length. The head is initialized from a released text-only block head for Qwen3-8B (Chen et al., 2026) and trained for one epoch on one GPU. The recipe contains no document, infographic, or chart data. The full training card is in Appendix C.
4 An Entropy Law of Draftability
Which workloads does a one-pass drafter serve best? We answer with one scalar, the target’s next-token entropy at the round root. Entropy has been used to gate or stop drafting (Agrawal et al., 2024; Tong et al., 2026; Mahmoud, 2026), and we give its relation to acceptance a closed form that the experiments then test.
We model the matches along a block as a survival process. The accepted run continues at offset with probability , and we approximate these hazards by a single match probability that varies slowly with the offset. Taking the match logit to be affine in the entropy gives
| (2) | ||||
with block horizon . Here is the zero-entropy intercept and the entropy slope, both fitted on each task’s round log. The truncation matters only near , since the untruncated mean exceeds the truncated one by the relative amount , which is below whenever . Appendix A derives Equation (2), grounds its monotone part in Fano’s inequality, and bounds how much round-level variance any entropy-only predictor can explain.
Equation (2) describes the top- path of a draft. A wider tree adds siblings at every depth and so raises acceptance at every entropy, and we measure this gain to be a nearly constant factor across tasks.
Two predictions follow. First, strictly decreases in , hence any workload that systematically lowers the target’s entropy yields longer accepted blocks. Grounded generation is such a workload, because it copies determinate strings such as glyphs, numbers, and answer spans from the image. Second, the law has a ceiling that grounded generation breaks.
Theorem 2 (Grounded tail, informal).
Equation (2) caps the near-certain mean at . If near-zero-entropy rounds enter, with probability , a verbatim-copy state whose matches persist, the mean as becomes , where is the mean of the ordinary state, and it exceeds once passes an explicit threshold.
The copy regime is where one-pass drafting gains the most, because a verbatim run of length costs an autoregressive drafter sequential passes and a block drafter one. Appendix A states the two-state model behind Theorem 2 and tests it on prompt-disjoint splits.
| EAGLE3-VL, eight draft passes | GLANCE, one draft pass | ||||||||
| Task | AR ms/tok | ms/tok | speedup | ms/tok | speedup | GLANCE faster by | |||
| Higher-entropy tasks | |||||||||
| Captioning | |||||||||
| TextVQA | |||||||||
| Lower-entropy tasks | |||||||||
| InfographicVQA | |||||||||
| DocVQA | |||||||||
| ChartQA | |||||||||
| Geometric mean | |||||||||
5 Experiments
5.1 Setup
Models and tasks.
The target is Qwen3-VL-8B-Instruct (Bai et al., 2025a), kept frozen. We evaluate five tasks ordered from open-ended to grounded, namely COCO captioning (Lin et al., 2014), TextVQA (Singh et al., 2019), InfographicVQA (InfoVQA) (Mathew et al., 2022), DocVQA (Mathew et al., 2021), and ChartQA (Masry et al., 2022). We call the last three, whose answers are read off a document or chart, the grounded tasks, and they are the three tasks of lowest mean target entropy in Table 1. The rest of the standard VLM suite is multiple-choice or single-word and leaves no decode loop to shorten. Captioning and TextVQA prompts for GLANCE come from the test split of each source, which is disjoint from its training rows, and the other three tasks appear in none of the primary head’s training data.
Baselines.
Following the EAGLE series (Li et al., 2024b; Li et al., 2024a; Li et al., 2025), we compare against released methods that preserve the target’s output. These are n-gram prompt lookup, classic two-model speculation with image-conditioned and text-only drafts, the production EAGLE3-VL head (AQ-MedAI, 2026), and ViSpec (Kang et al., 2025) together with the EAGLE-2 and Medusa baselines of its codebase. No published VLM drafter releases a head for our target, so we train the ViSpec-codebase heads ourselves with the official recipe.
Protocol.
Decoding is batch one and greedy, with up to new tokens. The acceptance length counts tokens rather than time and thus compares systems across engines. Wall-clock speedup is the ratio of the mean decode time for a token, autoregressive over speculative, with both arms measured in one engine on one GPU, and we never compare speedups across engines. Prompt counts, hardware, and results under sampling at temperature are in Appendix C.
| acceptance length | |||||||
| Method (draft passes a round) | Params | Captioning | TextVQA | InfoVQA | DocVQA | ChartQA | Lossless |
| Qwen3-VL-8B, training-free and two-model drafting | |||||||
| n-gram lookup (0) | |||||||
| Classic SD, Qwen3-VL-4B (8) | B | ||||||
| Classic SD, Qwen3-1.7B text-only (8) | B | ||||||
| Qwen3-VL-8B, trained draft heads | |||||||
| EAGLE3-VL, production (5) | B | ||||||
| EAGLE-2, ViSpec codebase (3) | B | ||||||
| ViSpec, official recipe (3) | B | ||||||
| Medusa, ViSpec codebase (1) | B | ||||||
| GLANCE (1) | B | ||||||
| Qwen3-VL-8B, matched training with one corpus, target, batch, schedule, and framework | |||||||
| EAGLE-3 head, depth-3 chain (3) | B | ||||||
| GLANCE, budget-63 tree (1) | B | ||||||
| Qwen3-VL-8B, GLANCE retrained on ViSpec’s corpus | |||||||
| GLANCE, 1 epoch (1) | B | ||||||
| GLANCE, 21 epochs (1) | B | ||||||
| ViSpec’s home target Qwen2.5-VL-7B | |||||||
| ViSpec, released head (3) | B | ||||||
| GLANCE (1) | B | ||||||
5.2 Head-to-Head in a Production Engine
Table 1 shows GLANCE and the production EAGLE3-VL head inside SGLang on one RTX A6000 in bf16, with both CUDA-graph captured and both verifying a tree of draft tokens a round, EAGLE3-VL its own top- dynamic tree and GLANCE its prefix tree. Engine, GPU, and round budget are thus shared, and the one structural difference left is how the draft tokens are produced, in eight sequential passes by EAGLE3-VL and in one by GLANCE. This eight-pass tree decodes to faster than EAGLE3-VL’s released configuration of three passes over four draft tokens on every task. On the three lower-entropy tasks, GLANCE is the faster system by to and reaches the speed of autoregressive decoding on ChartQA. Every paired bootstrap interval excludes zero, GLANCE is faster on at least of the prompts of each of these tasks, and it also leads on the five-task geometric mean, by with a interval from to . None of the three tasks appears in GLANCE’s training data. The lead on these three tasks holds for two further training seeds and at temperature .
On captioning and TextVQA, the two tasks with the highest mean entropy , the eight-pass head leads. The product-of-marginals ranking of Assumption 1 accounts for this split. It is accurate when the tokens of a block are nearly determined by the image, whereas on free-running text an autoregressive drafter stays coherent by construction. Consistently, GLANCE accepts longer blocks on every lower-entropy task than on either higher-entropy task, whereas the production head places TextVQA above both DocVQA and InfographicVQA. The production head’s lead on these two tasks also comes from its training, on K ALLaVA-4V samples (AQ-MedAI, 2026) against our K. With an EAGLE-3 head and a GLANCE head trained on one corpus and run under the same tree, GLANCE is faster on all five tasks, by on captioning, on TextVQA, and to on the three grounded tasks, as Table 4 shows. On full-resolution InfographicVQA prompts, the margin holds at with K tokens of context and at with about K, and both intervals lie above zero. Details, including the prompt-level intervals, are in Appendix E.
5.3 Acceptance Against Released Drafters
Table 2 reports the acceptance length of every drafter released for these targets, each at its own operating point. Among trained heads, GLANCE accepts the longest blocks on all five tasks. Classic two-model speculation with a B draft accepts more on four of the five tasks, but a draft half the size of the target costs about half a target pass for each drafted token, so it slows decoding on every task, as Appendix F shows. The ViSpec codebase isolates what its vision adaptor adds. Trained without the adaptor, the head is exactly the EAGLE-2 drafter, and the two heads differ by at most in acceptance length on any task, so the target’s fused states, which both heads read, already carry what the adaptor’s compressed image tokens add.
Four controls separate the architecture from its training. First, trained from scratch with one K-row corpus, target, global batch, schedule, and framework (Li et al., 2026b), GLANCE decodes faster than an EAGLE-3 head on all five tasks under the shared tree of Section 5.2, where that head’s acceptance length is to . This corpus, unlike our primary recipe, contains document and chart rows. Second, retrained on ViSpec’s own K-row corpus, at one epoch and at ViSpec’s twenty-one, GLANCE stays within one percent of its acceptance under our recipe, pooled over the five tasks, and its lead over the ViSpec family therefore holds on ViSpec’s own data and schedule. Third, on ViSpec’s home target Qwen2.5-VL-7B, a GLANCE head trained on that target accepts longer blocks than the released ViSpec head on four of the five tasks. Fourth, the released text head that GLANCE starts from, run unchanged, trails the production head on all five tasks of Table 1, by on the geometric mean, and our training lengthens its accepted blocks by to , which turns that deficit into GLANCE’s lead. Details of all four controls are in Appendix C.
5.4 Where the Acceptance Comes From
Tree against chain.
GLANCE’s head has the parameters of EAGLE3-VL’s, so a larger head might explain the gains. Panel (a) of Figure 3 measures what the tree adds by running the identical head as a width- chain. The budget- tree accepts between and the chain’s length on every task, a nearly constant factor, as the law anticipates. The tree also decodes faster than the chain in wall-clock time, net of verifying a wider tree. In SGLang, a GLANCE round is to shorter than an EAGLE3-VL round on every task, so the larger head drafts all fifteen offsets in one pass for less than the small head pays for eight.
Vision through fused states.
Panel (b) replaces the target’s fused states with those of a text-only Qwen3-8B, which never sees the image and whose states the head was initialized on. Acceptance falls on every task, and it falls most on grounded ones, since the text-only states retain of GLANCE’s acceptance on captioning but only on DocVQA. Zeroing the conditioning states collapses to , which shows that the head drafts from the target’s context.
Width saturates early.
Panel (c) sweeps the verifier budget. Speedup rises from to on every task and then flattens. Width pays the most on ChartQA, the task with the longest accepted blocks. Appendix E tabulates the sweep and breaks down the cost of a round.
5.5 Testing the Entropy Law
Fit across tasks.
Panel (a) of Figure 4 shows GLANCE’s accepted length falling with entropy on all five tasks, with grounded tasks accepting more at nearly every entropy level. Pooled over the rounds of captioning, TextVQA, and DocVQA, the law has and , and the slope steepens with grounding, from on captioning to on ChartQA. Isotonic fits explain at least of the variance of the ten entropy-decile means on every task. The fits are listed in Appendix G.
Image and entropy.
To test whether the image itself is responsible, we intervene on the image alone. The same prompts are decoded with the correct and with a mismatched image along a shared teacher-forced trajectory. Panel (b) shows that the correct image lengthens accepted blocks on every task, and more so as grounding increases. The correct image also lowers the target’s mean entropy along that trajectory, from to nats on DocVQA, and once entropy is held fixed the remaining image effect is no longer positive. Together with the fused-state ablation, this places the image’s contribution in the entropy channel.
The grounded tail.
Panel (c) compares the cap of Theorem 2 with the measured mean of in each task’s lowest-entropy decile. Every floor lies above its cap, and the excess grows with grounding, reaching against on DocVQA. On ChartQA, of all rounds accept eight or more tokens.
Transfer.
The law keeps its form under a change of drafter, target, and modality, with a positive slope in every case. Two-model drafts on Qwen2.5-VL-7B and Qwen2-VL-7B fit at and , and beyond vision the slope stays positive on speech recognition, chat, code, and a non-Qwen backbone. These targets all feed vision as a token prefix, and Appendix G reports the fits and a cross-attention target.
6 Conclusion
We presented GLANCE, a lossless one-pass block drafter for frozen vision-language models. Its block-diffusion head reads vision through the target’s fused states and fills a whole block in one pass, which suits the long verbatim runs of grounded generation. A wide tree verified in one target pass turns that block into accepted length, and the output stays exactly the target’s greedy decoding. An entropy law explains when this pays and carries over across targets and modalities.
Limitations
Our measurements are at batch one, the regime in which decoding is bound by memory bandwidth and speculative decoding is deployed for latency. We do not measure serving at large batch sizes, which changes the cost of verification. The exactness guarantee concerns greedy decoding. Under sampling, the walk draws each token from the verified row of its node and descends while the draw is a child in the tree. Every committed token is then a draw from the target’s own distribution, and the output distribution is preserved while byte identity is not defined. Our characterization of draftability is stated for targets that feed vision as a token prefix, which covers the Qwen-VL family studied here. Finally, we study single-image tasks whose outputs span tens to hundreds of tokens, since tasks with single-token answers leave no decode loop to shorten.
Ethical Considerations
This work accelerates inference of existing vision-language models without changing their greedy outputs, so it introduces no new model behavior. Risks of the underlying models, such as biased or incorrect descriptions of images, carry over unchanged and are neither amplified nor mitigated. All datasets are public research benchmarks used under their licenses, and no human subjects or personal data are involved. Faster decoding lowers the energy spent on each generated token.
References
- Agrawal et al. (2024) Sudhanshu Agrawal, Wonseok Jeon, and Mingu Lee. 2024. AdaEDL: Early draft stopping for speculative decoding of large language models via an entropy-based lower bound on token acceptance probability. Preprint, arXiv:2410.18351.
- Ankner et al. (2024) Zachary Ankner, Rishab Parthasarathy, Aniruddha Nrusimha, Christopher Rinard, Jonathan Ragan-Kelley, and William Brandon. 2024. Hydra: Sequentially-dependent draft heads for Medusa decoding. In Conference on Language Modeling (COLM). ArXiv:2402.05109.
- AQ-MedAI (2026) AQ-MedAI. 2026. Qwen3-VL-8B-Instruct-eagle3. https://huggingface.co/AQ-MedAI/Qwen3-VL-8B-Instruct-eagle3. Hugging Face model checkpoint.
- Arriola et al. (2025) Marianne Arriola, Aaron Gokaslan, Justin T. Chiu, Zhihan Yang, Zhixuan Qi, Jiaqi Han, Subham Sekhar Sahoo, and Volodymyr Kuleshov. 2025. Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models. In International Conference on Learning Representations.
- Bai et al. (2025a) Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, and 1 others. 2025a. Qwen3-VL Technical Report. arXiv preprint arXiv:2511.21631.
- Bai et al. (2025b) Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, and 1 others. 2025b. Qwen2.5-VL Technical Report. arXiv preprint arXiv:2502.13923.
- Cai et al. (2024) Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D Lee, Deming Chen, and Tri Dao. 2024. Medusa: Simple LLM inference acceleration framework with multiple decoding heads. In International Conference on Machine Learning.
- Chen et al. (2023) Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper. 2023. Accelerating large language model decoding with speculative sampling. arXiv preprint arXiv:2302.01318.
- Chen et al. (2024a) Guiming Hardy Chen, Shunian Chen, Ruifei Zhang, Junying Chen, Xiangbo Wu, Zhiyi Zhang, Zhihong Chen, Jianquan Li, Xiang Wan, and Benyou Wang. 2024a. ALLaVA: Harnessing GPT4V-synthesized data for lite vision-language models. Preprint, arXiv:2402.11684.
- Chen et al. (2026) Jian Chen, Yesheng Liang, and Zhijian Liu. 2026. DFlash: Block Diffusion for Flash Speculative Decoding. In International Conference on Machine Learning.
- Chen et al. (2024b) Zhuoming Chen, Avner May, Ruslan Svirschevski, Yuhsun Huang, Max Ryabinin, Zhihao Jia, and Beidi Chen. 2024b. Sequoia: Scalable and robust speculative decoding. In Advances in Neural Information Processing Systems. ArXiv:2402.12374.
- Cheng et al. (2025) Zicong Cheng, Guo-Wei Yang, Jia Li, Zhijie Deng, Meng-Hao Guo, and Shi-Min Hu. 2025. DEER: Draft with diffusion, verify with autoregressive models. Preprint, arXiv:2512.15176.
- Christopher et al. (2025) Jacob K Christopher, Brian R Bartoldson, Tal Ben-Nun, Michael Cardei, Bhavya Kailkhura, and Ferdinando Fioretto. 2025. Speculative diffusion decoding: Accelerating language generation through diffusion. In Proceedings of the North American Chapter of the Association for Computational Linguistics (NAACL). ArXiv:2408.05636.
- Fu et al. (2024) Yichao Fu, Peter Bailis, Ion Stoica, and Hao Zhang. 2024. Break the sequential dependency of LLM inference using lookahead decoding. In Proceedings of the International Conference on Machine Learning (ICML). ArXiv:2402.02057.
- Gagrani et al. (2024) Mukul Gagrani, Raghavv Goel, Wonseok Jeon, Junyoung Park, Mingu Lee, and Christopher Lott. 2024. On speculative decoding for multimodal large language models. arXiv preprint arXiv:2404.08856. ELVM @ CVPR 2024.
- Ganesan et al. (2025) Mugilan Ganesan, Shane Segal, Ankur Aggarwal, Nish Sinnadurai, Sean Lie, and Vithursan Thangarasa. 2025. MASSV: Multimodal adaptation and self-data distillation for speculative decoding of vision-language models. In Findings of EMNLP.
- Hu et al. (2025) Yunhai Hu, Tianhua Xia, Zining Liu, Rahul Raman, Xingyu Liu, Bo Bao, Eric Sather, Vithursan Thangarasa, and Sai Qian Zhang. 2025. DREAM: Drafting with refined target features and entropy-adaptive cross-attention fusion for multimodal speculative decoding. In Advances in Neural Information Processing Systems. ArXiv:2505.19201.
- Huang et al. (2025a) Haiduo Huang, Fuwei Yang, Zhenhua Liu, Xuanwu Yin, Dong Li, Pengju Ren, and Emad Barsoum. 2025a. SpecVLM: Fast speculative decoding in vision-language models. Preprint, arXiv:2509.11815.
- Huang et al. (2025b) Kaixuan Huang, Xudong Guo, and Mengdi Wang. 2025b. SpecDec++: Boosting speculative decoding via adaptive candidate lengths. In Conference on Language Modeling (COLM).
- Huo et al. (2025) Mingxiao Huo, Jiayi Zhang, Hewei Wang, Jinfeng Xu, Zheyu Chen, Huilin Tai, and Yijun Chen. 2025. Spec-LLaVA: Accelerating vision-language models with dynamic tree-based speculative decoding. arXiv preprint arXiv:2509.11961.
- Ji et al. (2025) Yicheng Ji, Jun Zhang, Heming Xia, Jinpeng Chen, Lidan Shou, Gang Chen, and Huan Li. 2025. SpecVLM: Enhancing speculative decoding of video llms via verifier-guided token pruning. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP). ArXiv:2508.16201.
- Kang et al. (2025) Jialiang Kang, Han Shu, Wenshuo Li, Yingjie Zhai, and Xinghao Chen. 2025. ViSpec: Accelerating vision-language models with vision-aware speculative decoding. In Advances in Neural Information Processing Systems.
- Kwon et al. (2026) Young D. Kwon, Miles Williams, Rui Li, Alexandros Kouris, and Stylianos I. Venieris. 2026. WhiFlash: Accelerating speculative decoding with token-level cross-paradigm routing. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP). ArXiv:2606.07710.
- Leviathan et al. (2023) Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023. Fast inference from transformers via speculative decoding. In International Conference on Machine Learning.
- Li et al. (2026a) Guanghao Li, Zhihui Fu, Min Fang, Qibin Zhao, Ming Tang, Chun Yuan, and Jun Wang. 2026a. DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding. In Findings of ACL.
- Li et al. (2026b) Shenggui Li, Chao Wang, Yikai Zhu, Yubo Wang, Fan Yin, Shuai Shi, Yefei Chen, Xiaomin Dong, Qiaoling Chen, Jin Pan, and 1 others. 2026b. SpecForge: A flexible and efficient open-source training framework for speculative decoding. arXiv preprint arXiv:2603.18567.
- Li et al. (2026c) Tianyi Li, Yaxin Luo, Xinyi Shang, and Zhiqiang Shen. 2026c. DARTree: Speculative diffusion decoding with autoregressive draft trees. Preprint, arXiv:2608.13524.
- Li et al. (2024a) Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. 2024a. EAGLE-2: Faster inference of language models with dynamic draft trees. In Empirical Methods in Natural Language Processing.
- Li et al. (2024b) Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. 2024b. EAGLE: Speculative sampling requires rethinking feature uncertainty. In International Conference on Machine Learning.
- Li et al. (2025) Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. 2025. EAGLE-3: Scaling up inference acceleration of large language models via training-time test. In Advances in Neural Information Processing Systems. ArXiv:2503.01840.
- Lin et al. (2014) Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. In ECCV.
- Mahmoud (2026) Saif Mahmoud. 2026. Acceptance dynamics across cognitive domains in speculative decoding. Preprint, arXiv:2604.14682.
- Masry et al. (2022) Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. 2022. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. In Findings of ACL.
- Mathew et al. (2022) Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawahar. 2022. InfographicVQA. In WACV.
- Mathew et al. (2021) Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. 2021. DocVQA: A dataset for VQA on document images. In WACV.
- Miao et al. (2024) Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, Chunan Shi, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia. 2024. SpecInfer: Accelerating large language model serving with tree-based speculative inference and verification. In ASPLOS.
- Monea et al. (2023) Giovanni Monea, Armand Joulin, and Edouard Grave. 2023. PaSS: Parallel speculative sampling. arXiv preprint arXiv:2311.13581. NeurIPS 2023 Workshop on Efficient Natural Language and Speech Processing.
- Nie et al. (2025) Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, and Chongxuan Li. 2025. Large Language Diffusion Models. In Advances in Neural Information Processing Systems.
- Ringel and Romano (2026) Liran Ringel and Yaniv Romano. 2026. Accelerating Speculative Decoding with Block Diffusion Draft Trees. In Conference on Language Modeling (COLM).
- Shen et al. (2026) Hui Shen, Xin Wang, Ping Zhang, Yunta Hsieh, Qi Han, Zhongwei Wan, Ziheng Zhang, Jingxuan Zhang, Jing Xiong, Ziyuan Liu, Yifan Zhang, Hangrui Cao, Chenyang Zhao, and Mi Zhang. 2026. MMSpec: Benchmarking speculative decoding for vision-language models. Preprint, arXiv:2603.14989.
- Singh et al. (2019) Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. 2019. Towards VQA models that can read. In CVPR.
- Tong et al. (2026) Yujia Tong, Tian Zhang, Yunyang Wan, Kaiwei Lin, Jingling Yuan, and Chuang Hu. 2026. SAGE: Accelerating vision-language models via entropy-guided adaptive speculative decoding. Preprint, arXiv:2602.00523.
- Wang et al. (2025a) Jikai Wang, Yi Su, Juntao Li, Qingrong Xia, Zi Ye, Xinyu Duan, Zhefeng Wang, and Min Zhang. 2025a. OPT-Tree: Speculative decoding with adaptive draft tree structure. Transactions of the Association for Computational Linguistics (TACL). ArXiv:2406.17276.
- Wang et al. (2026) Zheng Wang, Davis Wertheimer, Yu Chin Fabian Lim, Mudhakar Srivatsa, Raghu K. Ganti, Minjia Zhang, and Naigang Wang. 2026. xPress: Parallel refinement for diffusion drafters in speculative decoding. Preprint, arXiv:2608.02438.
- Wang et al. (2025b) Zihua Wang, Ruibo Li, Haozhe Du, Joey Tianyi Zhou, Yu Zhang, and Xu Yang. 2025b. SpecFLASH: A latent-guided semi-autoregressive speculative decoding framework for efficient multimodal generation. Preprint, arXiv:2505.12728.
- Wu et al. (2026a) Chengyue Wu, Shiyi Lan, Yonggan Fu, Sensen Gao, Jin Wang, Jincheng Yu, Jose M. Alvarez, Pavlo Molchanov, Ping Luo, Song Han, Ligeng Zhu, and Enze Xie. 2026a. Fast-dVLM: Efficient block-diffusion vlm via direct conversion from autoregressive vlm. arXiv preprint arXiv:2604.06832.
- Wu et al. (2026b) Tianyu Wu, Yu Yao, Zhenting Qi, Han Zheng, Zhuohan Wang, Haoran Ma, Lawrence Liao, Himabindu Lakkaraju, Ju Li, and Yilun Du. 2026b. D-PACE: Dynamic position-aware cross-entropy for parallel speculative drafting. Preprint, arXiv:2605.18810.
- Xia et al. (2024) Heming Xia, Zhe Yang, Qingxiu Dong, Peiyi Wang, Yongqi Li, Tao Ge, Tianyu Liu, Wenjie Li, and Zhifang Sui. 2024. Unlocking efficiency in large language model inference: A comprehensive survey of speculative decoding. In Findings of the Association for Computational Linguistics (ACL Findings). Spec-Bench.
- Xie et al. (2026) Zhinan Xie, Peisong Wang, Shuang Qiu, and Jian Cheng. 2026. HiViS: Hiding visual tokens from the drafter for speculative decoding in vision-language models. In CVPR Findings.
- Yin et al. (2024) Ming Yin, Minshuo Chen, Kaixuan Huang, and Mengdi Wang. 2024. A theoretical perspective for speculative decoding algorithm. In Advances in Neural Information Processing Systems (NeurIPS). ArXiv:2411.00841.
- Zhang et al. (2026a) Jiebin Zhang, Zhenghan Yu, Song Liu, Eugene J. Yu, Zheng Li, Dawei Zhu, Jiangshan Duo, Weimin Xiong, Yifan Song, Guanghua Yu, and 1 others. 2026a. DFlare: Scaling up draft capacity for block diffusion speculative decoding. Preprint, arXiv:2606.02091.
- Zhang et al. (2026b) Kewei Zhang, Jin Wang, Sensen Gao, Chengyue Wu, Yulong Cao, Songyang Han, Boris Ivanovic, Langechuan Liu, Marco Pavone, Song Han, Daquan Zhou, and Enze Xie. 2026b. Fast-dDrive: Efficient block-diffusion VLM for autonomous driving. Preprint, arXiv:2605.23163.
- Zhang et al. (2026c) Shuai Zhang, Huachuan Qiu, Hongliang He, and Yong Dai. 2026c. Cost-Aware Diffusion Draft Trees for Speculative Decoding. Preprint, arXiv:2606.01813.
- Zhang et al. (2026d) Yaojie Zhang, Linfeng Zhang, Bin Cui, and Xupeng Miao. 2026d. DFlow: Enabling verifier information flow in block diffusion speculative decoding. Preprint, arXiv:2609.06498.
Appendix A Theoretical Statements and Proofs
A.1 Losslessness
Theorem 1 (full).
For the decoder of Definition 1, assume (float-exactness) that at every visited prefix the target is unique and the top logit gap exceeds the packed-versus-unpacked logit perturbation. Then for any candidate tree the committed token sequence equals the target’s autoregressive greedy sequence exactly, and losslessness can fail only at a margin below that perturbation, that is, at a near-tie.
Proof.
Single round. By Definition 1 the verifier row at any prefix reproduces exactly, since ancestor-only masking and token packing change the attention layout, not the conditioning of a row. Let be the target greedy token there. If , the walk descends and commits , and otherwise it stops and emits as the next root. Either way the token produced at that position is , and the tree affects only how many tokens the round commits.
Induction over rounds. The emitted stream concatenates accepted paths and corrections, and each round’s correction is the next round’s root. Every produced token is the target greedy token at its own prefix, so induction on the output position gives the committed sequence bitwise, independent of all budget choices.
Tie sensitivity. The only step that can break is the equality of the packed and the autoregressive . It fails only when falls below the packed-versus-unpacked logit perturbation, that is, at a near-tie. In fp32 the perturbation is below every observed margin, and the fp32 audit shows zero divergences. ∎
A.2 Derivation of the Law
Assumption 2 (Offset survival with marginal hazards).
Conditioned on the round, or equivalently on , let the drafter’s top- token match the target greedy token at offset with probability , and let be the leading run of matches before the first miss, capped at . We assume the survival factorizes into the marginal hazards,
| (3) | ||||
This is an assumption, not a consequence of the chain rule, since it replaces the conditional hazards by the marginals .
Lemma 1 (Slowly varying hazard gives a truncated-geometric mean).
Proof.
Setting in Equation (3) gives the finite geometric series. Its truncated tail is
so the relative truncation error is . At , the error is , , and at , , and , so the closed form is tight except near the entropy floor.
Two errors are folded in, the truncation tail and the slowly varying step itself. The step overstates when decays in . It understates the mean at extremely low entropy when one is fitted across a range that contains a sharp spike at low . The supremum of the DocVQA fit is at and , below the measured DocVQA bottom-decile mean of , which Theorem 2 explains.
Substituting gives Equation (2), which we use as an operational characterization of the conditional mean, not as an exact survival identity. ∎
Proposition 1 (Entropy controls the match, and the affine logit is a surrogate).
Let be the probability that the drafter’s top- token at the root equals the target greedy token, and the target’s root entropy. (i) The target’s own top- error obeys Fano’s inequality
so larger forces larger , and forces . (ii) Modeling the drafter’s logit error as logistic with scale gives for the target’s top- margin , which decreases in by (i). (iii) A first-order expansion of in gives with . Part (i) predicts sign and monotonicity, and parts (ii) and (iii) are modeling steps that select the functional form.
Proof.
- (i)
View the target’s top- token as an estimator of a draw . Fano’s inequality bounds by an increasing function of the error on , so grows with and vanishes as .
- (ii)
The drafter preserves the target’s ordering if and only if the perturbed margin stays positive. A difference of logistics is logistic, so
and on a shell of fixed entropy decreases in by (i).
- (iii)
This is the first-order expansion. Its adequacy over the empirical window is what the curve measures.
A fuller derivation under a temperature family, in which the target’s logits are a fixed shape scaled by an inverse temperature, shows that is affine to to over the resolvable window for Zipf, geometric-gap, and near-binary logit shapes, with self-confidence slopes to .
This reduces the match slope to a single drafter-noise scale through . Matching the measured offset- slopes gives , , and on captioning, TextVQA, and DocVQA. The affine form is exactly a local linearization, and since the binary closed form is provably curved globally, the law is stated over the empirical window only. ∎
Corollary (Grounded is more draftable).
Under the untruncated form of Equation (2),
so is strictly decreasing in , and the truncated form is strictly increasing in and hence decreasing in as well. Any regime with systematically lower next-token entropy therefore has a longer expected accepted block. Grounded and OCR generation copies determinate glyph and answer strings from the image and thereby collapses the target’s entropy. Grounded spans are therefore more draftable than open text, and the mean accepted length of a task is ordered by grounding.
Proof.
The derivative computation is displayed in the statement and uses only and . The ordering claim uses only the measured monotone decrease of in , not the fitted closed form. The measured round logs order the mean accepted length by grounding, at on captioning, on TextVQA, and on DocVQA, and the InfoVQA and ChartQA logs extend the ordering to and , the latter under ChartQA’s templated prompts. The gap between DocVQA and captioning, , has a bootstrap interval of and a Mann-Whitney , and the grounded task carries both the higher and the steeper , as Table 12 lists. ∎
A.3 The Noise Ceiling
Proposition 2 (Entropy-only noise ceiling).
Among all -measurable predictors of , the largest attainable coefficient of determination is , attained by the conditional mean. For a truncated-geometric the within-round variance is of order , so is intrinsically small regardless of the predictor’s quality.
Proof.
For any , orthogonality of the conditional mean gives
which is minimized at , and the law of total variance gives the stated ratio. For the untruncated geometric, with ,
so the within-round variance dominates unless is comparably large. A nonparametric estimate of the between-to-total variance over bins gives ceilings of , , and on captioning, TextVQA, and DocVQA. The law’s round-level values of , , and therefore realize , , and of what any entropy-only predictor could reach, and Table 12 lists the same fractions for InfoVQA and ChartQA. This is why the curve is the appropriate score. ∎
A.4 The Grounded Tail
Assumption 3 (Two-state determinate or uncertain survival).
Conditioned on the round, the matches along a block follow a two-state hidden Markov chain over , with entry distribution , match probability in state , transition matrix upon a match, the run ending at the first miss, and a cap at . State models verbatim copying (, sticky ), and is the ordinary regime with a moderate hazard. The parameters governing are estimated in regime, from near-zero-entropy rounds only, rather than extrapolated from the affine fit. The geometric law of Lemma 1 is the degenerate case .
Theorem 2 (full).
Let . (A) Under the single-hazard law of Equation (2), , a rigorous cap that needs no fit. (B) Under Assumption 3, with , , and ,
| (4) | ||||
whose effective hazard is not constant. Taking gives , which exceeds exactly when . This threshold lies in because , where on DocVQA and . (C) With in-regime maximum-likelihood parameters, Equation (4) recovers the measured DocVQA floor mean of , against the geometric cap of . Proposition 3 gives the falsifiable validation.
Proof.
- (A)
is strictly increasing on , and is maximized as , where it equals . Then gives the bound .
- (B)
Starting from , the joint event of a match at offset followed by state has mass , and each accepted offset multiplies by , so
and the tail-sum identity gives the mean. Survival is a nonnegative combination of the two eigenmodes of , so the hazard moves monotonically toward and is not constant unless . Its first value can approach and exceed , which the constant-hazard family cannot do. As ,
which gives the stated excess over the cap.
- (C)
This is the numerical instantiation, which recovers the measured floor mean against the geometric cap.∎
Proposition 3 (Prompt-disjoint held-out falsification).
Take the bottom-entropy-decile DocVQA rounds, from prompts, and split them by prompt into disjoint sets. (i) On the full floor the geometric run length is rejected, with Pearson , , , while the likelihood ratio decisively favors the two-state form, with , , . (ii) Frozen on one prompt set and scored on the disjoint one, the two-state model attains a smaller held-out than the geometric in of tests, over both split directions and five cross-validation folds, with a mean held-out of against .
Proof.
The proposition records the outcome of the stated estimation procedure, with the construction and statistics as listed. ∎
Appendix B Decoding Round
Algorithm 1 gives one decoding round of GLANCE. The draft pass and the verify pass are the only two forward passes of a round, and the tree budget enters only the size of the verify pass.
Appendix C Implementation Details
Primary drafter.
The primary drafter is a block-diffusion head. It reads the target’s fused hidden states and proposes the whole block in one forward pass, with the vision tower and the target decoder frozen. We use the checkpoint at the end of its one-epoch run, and Table 3 lists its configuration. The remaining constants are an intermediate size of , a block horizon , an image long side of pixels, AdamW betas of and with no weight decay, gradient clipping at , seed , one A6000-class GPU, and bf16 weights of GB.
| Setting | Value |
|---|---|
| Architecture | |
| layers | |
| parameters | B |
| block size | |
| hidden size | |
| attention heads | |
| kept target layers | |
| Objective | |
| loss | coverage CE, offset weight |
| reveal | variable-mask block prefix |
| Data | |
| source | target greedy, new tokens |
| prompts | COCO TextVQA |
| rows | K teacher-forced |
| Optimization | |
| optimizer, lr | AdamW, |
| epochs, accumulation | , |
| max sequence length | |
| warm start | Qwen3-8B DFlash head |
Evaluation prompts.
The drafter trains on COCO-Caption2017 val rows to and on TextVQA validation rows to . The home-target head trains on the first rows of the same two splits. Every GLANCE row of the two main tables therefore takes its captioning and TextVQA prompts from the test split of each source, with no overlap with any training row, namely prompts in each task on Qwen3-VL-8B and captioning and TextVQA prompts on the home target, the counts of ViSpec’s own evaluation. In Table 1, all three systems decode these same test prompts. The other baselines of Table 2 draw their prompts from the validation rows on Qwen3-VL-8B, and the released ViSpec head runs its native loaders on its home target. InfographicVQA, DocVQA, and ChartQA appear in no training corpus of the primary head. The analyses of Figures 3 and 4 compare conditions of one head on shared prompts and draw their captioning and TextVQA prompts from the validation rows. Every system of the two main tables on Qwen3-VL-8B, apart from the ViSpec-codebase heads with their official loaders, uses the same prompt form for each task. Captioning prompts ask the target to describe the image in detail, and TextVQA prompts are the dataset questions as written. The DocVQA prompt reads “Read the document image and answer the question. Quote the exact text from the document, then briefly explain.” before the question, and the InfographicVQA and ChartQA prompts ask in the same way for the exact text, numbers, values, or labels the target reads, followed by an explanation, so that these tasks also leave a decode loop to shorten.
Second and third targets.
The Qwen2.5-VL-7B head is trained from scratch with the OCR-augmented recipe on that target’s own greedy generations, with caption, TextVQA, and OCR rows from DocVQA, ChartQA, and InfoVQA prompts disjoint from the evaluation prompts, rows in total, for seven epochs at global batch . It is evaluated with the same budget- tree and fp32 audit as the primary target, with captioning and TextVQA prompts from the test splits. Prompt counts match the ViSpec row task by task, for captioning and DocVQA and for TextVQA, which is where ViSpec’s staged image set ends, and the fp32 audit finds zero divergences at eight prompts in each task. The Llama-3.2-11B-Vision adapter is warm-started from the released Llama-3.1-8B DFlash head, whose held-out top- accuracy at the first offset climbs from to .
Matched-training comparison.
Both architectures are trained from scratch under one protocol, with the same -row OCR-augmented corpus, the same frozen target, global batch , three epochs, and SpecForge, and with each method’s default learning rate scaled by the square root of the batch. Each drafter then runs at its standard operating point, ours with a budget- tree and the EAGLE-3 head as a depth- chain, greedy, with new tokens, in one Hugging Face implementation on held-out prompts. Pooled over the five tasks the acceptance ratio is , at against , and a depth- chain adds to the EAGLE-3 head. The ALLaVA replication retrains both heads on an ALLaVA-Instruct corpus (Chen et al., 2024a) under the same protocol and gives against pooled over the five tasks. Timed in the same implementation on one A100, with prompts in each task and new tokens, GLANCE decodes at to autoregressive decoding against to for the EAGLE-3 head.
The same two heads also run in SGLang under the protocol of Table 1, with the tree of draft tokens, prompts in each task after a warm-up prompt, and new tokens. Each task’s two arms decode back to back on one RTX A6000. The corpus holds DocVQA and InfographicVQA validation rows and ChartQA test rows, so these three tasks take their prompts from the DocVQA and InfographicVQA test splits and the ChartQA validation split under the corpus’s prompt template, and captioning and TextVQA use the test prompts of Table 1. Table 4 reports the result. GLANCE is faster on to of the prompts of each task.
| EAGLE-3 head | GLANCE | ||||
|---|---|---|---|---|---|
| Task | ms/tok | ms/tok | GLANCE faster by | ||
| Captioning | |||||
| TextVQA | |||||
| InfographicVQA | |||||
| DocVQA | |||||
| ChartQA | |||||
The released text head without our training.
GLANCE starts from the released block head for Qwen3-8B (Chen et al., 2026), which has the same configuration and weight shapes. Run unchanged on Qwen3-VL-8B under the protocol of Table 1, with EAGLE3-VL decoding the same prompts back to back on one RTX A6000, it is slower than EAGLE3-VL on every task, and every paired bootstrap interval lies below zero, as Table 5 shows. EAGLE3-VL is the faster arm on at least of the prompts of each task. Our training lengthens the head’s accepted blocks by to .
| Task | EAGLE3-VL | Released | GLANCE | Released faster by |
|---|---|---|---|---|
| Captioning | ||||
| TextVQA | ||||
| InfographicVQA | ||||
| DocVQA | ||||
| ChartQA | ||||
| Geometric mean | ||||
ViSpec’s corpus and training length.
The training rows are the K LLaVA-Pretrain questions and images regenerated under ViSpec’s own data pipeline, with the same source file, shuffle seed, prompt template, pixel bounds, and sampling temperature, and with responses produced by the frozen target, since self-generation is part of our recipe. The head then trains with the published card unchanged, once for one epoch and once for twenty-one, the latter continuing from the one-epoch checkpoint at the same constant learning rate and global batch of eight rows. Our own corpus receives the same treatment. All arms are evaluated with the budget- tree, branching , prompts in each task, new tokens, and the fp32 audit at eight prompts in each task, back to back on one otherwise idle GPU. Figure 5 evaluates every checkpoint of both runs. Acceptance saturates within about five epochs on either corpus, and twenty-one epochs add on our corpus and on ViSpec’s.
ViSpec-codebase heads on our target.
These heads are trained on Qwen3-VL-8B-Instruct with ViSpec’s official codebase and two-stage schedule, namely text-only drafter pre-training on the authors’ ShareGPT split to convergence followed by multimodal fine-tuning on target-generated data, and Medusa is trained through the same codebase. They decode with the codebase’s tree search and acceptance rule under its official tree of tokens at depth with top- branching. We port the codebase to Qwen3-VL-8B so that the target runs its reference forward pass, with multimodal rotary positions and deepstack visual features. Each head decodes the prompts of each task that the GLANCE row uses, with the same -pixel images, greedy, batch one, and up to new tokens. The EAGLE-2 row is this recipe without the vision adaptor, and the ViSpec row is the same schedule with the adaptor, both trained from the same stage-one checkpoint, data, and seed, so the gap between the rows is the adaptor’s contribution alone.
ViSpec on its home target.
The released ViSpec head for Qwen2.5-VL-7B-Instruct runs with its native loaders and its default tree of tokens at depth , with top- branching and two query heads, greedy, batch one, and up to new tokens. ViSpec logs the bonus-excluded length, which we raise by the committed token.
Other protocol details.
Tables 1 and 2 run the same GLANCE head, the first in SGLang with images at native resolution and a tree of draft tokens, the second in our Hugging Face implementation at pixels with the budget- tree, so their acceptance lengths differ. In Table 1 the three arms of each task decode back to back on one RTX A6000, and GLANCE’s draft pass replays from a CUDA graph over a static key-value cache, with one graph for each number of tokens committed in the previous round. The mean entropy of Table 1 is the target’s mean root entropy over the rounds of the law fits in Section 5.5. EAGLE3-VL, the production AQ-MedAI checkpoint for this target, and the n-gram baseline run in vLLM , with prompts in each task and the warm-up sample excluded, and EAGLE3-VL runs as a chain of , which Table 7 sweeps from to . Prompt lookup proposes a draft only when an n-gram match exists, so its acceptance length is the length of the blocks it does propose. The classic two-model rows are the endpoints of the three-draft sweep of Table 11. The mismatched-image probe uses prompts in each task, with the three image conditions scored on a shared teacher-forced trajectory and prompt-clustered bootstrap intervals. Speedup is the ratio of mean decode time for a token over prompts.
Output equivalence.
Our implementation decodes prompts in each of three tasks, and each system’s generated tokens are compared position by position with the target’s own greedy decoding of the same prompts, in the same arithmetic. Table 6 reports the result. Decoding the target twice with no drafter shows that bf16 identity measures kernel arithmetic rather than any drafter, so the guarantee is decided in fp32, where both the drafter-free control and GLANCE reach of . The ViSpec-codebase heads and the released ViSpec head, compared the same way on prompts in each task with up to new tokens, reproduce all in fp32, as exact acceptance requires.
| System | Arithmetic | Identical |
| Our implementation | ||
| autoregressive, no drafter | bf16 | |
| autoregressive, no drafter | fp32 | |
| GLANCE, budget- tree | bf16 | |
| GLANCE, budget- tree | fp32 | |
| GLANCE, matched-training head | bf16 | |
| EAGLE-3 head, matched training | bf16 | |
| ViSpec codebase | ||
| EAGLE-2, ViSpec codebase | fp32 | |
| ViSpec, official recipe | fp32 | |
| Medusa, ViSpec codebase | fp32 | |
| ViSpec, released head, home target | fp32 | |
Sampling.
At temperature , with a top- of , no top- cut, and the protocol of Table 1 otherwise unchanged, GLANCE is faster than EAGLE3-VL on the three lower-entropy tasks, by on InfographicVQA, on DocVQA, and on ChartQA. Every paired bootstrap interval lies above zero, GLANCE is the faster arm on at least of the prompts of each of these tasks, and the geometric-mean margin over the five tasks is , with a interval from to . Both arms verify with SGLang’s tree sampling, which commits draws from the target’s own distribution.
| acceptance length | ms/token | |||||
|---|---|---|---|---|---|---|
| Chain length | Cap. | TVQA | Doc | Cap. | TVQA | Doc |
Appendix D Reproducibility
Section 3 and Appendix C specify the drafter’s architecture, training data, objective, and hyperparameters, and Appendix B gives one decoding round as pseudocode. Appendix A contains the complete proofs of all theoretical statements. Every measurement states its prompt count, and Appendix E lists each wall-clock comparison with its engine and sample size. Our decoding code, exactness audit, and law-fitting code are available at https://github.com/js-lee-AI/GLANCE, and we will release the drafter heads, the training code, and the round-level logs behind the entropy-law fits.
Appendix E Additional Wall-Clock Results
Where the cost sits.
Across the five tasks of Table 1, prefill including the visual encoder takes to of the end-to-end time of autoregressive decoding, the most on DocVQA, whose answers are the shortest, and the decode loop takes the remainder. The share falls as the output grows, since prefill is paid once and the loop on every token. Counting prefill, GLANCE remains faster than EAGLE3-VL on the three lower-entropy tasks, by to , and every paired bootstrap interval of these margins lies above zero.
Intervals behind the production-engine comparison.
The margin column of Table 1 is , the ratio of the two printed speedups. Its interval is a paired bootstrap over the shared prompts of that statistic with resamples. The intervals are on captioning, on TextVQA, on InfographicVQA, on DocVQA, and on ChartQA. GLANCE is the faster arm on , , and of the prompts of the three lower-entropy tasks. The same bootstrap, resampling the prompts of every task at once, puts the geometric-mean margin of in . Averaging prompt-level ratios instead of taking the ratio of means moves each figure by at most points and reorders nothing. The autoregressive arm stays within of its mean across the five tasks, so the tasks differ in what they generate rather than in what a token costs to decode.
EAGLE3-VL’s released configuration.
The model card of EAGLE3-VL lists an SGLang launch configuration of three draft steps, top- branching, and four draft tokens (AQ-MedAI, 2026). Run under the protocol of Table 1, back to back with the eight-step tree on one RTX A6000 for each task, it accepts to tokens a round against to , and the eight-step tree decodes faster on every task, by on captioning, on TextVQA, on InfographicVQA, on DocVQA, and on ChartQA. Every paired bootstrap interval lies above zero, the tree is the faster arm on at least of the prompts of each task, and the geometric-mean margin is , with a interval from to .
Training seeds.
Two further heads were trained with the configuration of Table 3 under seeds and , with the target’s fused states computed during training on two GPUs rather than precomputed, at the same global batch. Run under the protocol of Table 1, on one RTX A6000 for each task together with EAGLE3-VL and the primary head, both are faster than EAGLE3-VL on the three lower-entropy tasks, by and on InfographicVQA, and on DocVQA, and and on ChartQA, and every paired bootstrap interval lies above zero. On all five tasks their acceptance lengths lie within of the primary head’s.
Longer contexts, tree against chain, and budget.
Table 8 lists the long-context comparisons, on which GLANCE is the faster arm on of the prompts at each context length. Table 9 times the tree against the chain, and Table 10 sweeps the verifier budget.
| Comparison | Context | Result [ CI] |
|---|---|---|
| GLANCE vs. AR | K | |
| EAGLE3-VL vs. AR | K | |
| GLANCE vs. EAGLE3-VL | K | |
| GLANCE vs. AR | K | |
| EAGLE3-VL vs. AR | K | |
| GLANCE vs. EAGLE3-VL | K |
| ms/tok | ||||
|---|---|---|---|---|
| Task | AR | tree | chain | tree vs. chain |
| Captioning | ||||
| TextVQA | ||||
| DocVQA | ||||
| Speedup over AR | ||||
| acceptance length | speedup over AR | |||||||
|---|---|---|---|---|---|---|---|---|
| Budget | ||||||||
| Captioning | ||||||||
| TextVQA | ||||||||
| InfographicVQA | ||||||||
| DocVQA | ||||||||
| ChartQA | ||||||||
Round costs.
A GLANCE round is one draft pass and one verify pass, and tree width enters only inside the verify. Regressing the production head’s round time in SGLang on its number of draft passes, over five configurations from to draft tokens at to steps on captioning, gives ms plus ms for every sequential draft pass. At batch one the verify is therefore nearly flat in tree width, while depth is paid one pass at a time. Widening our tree costs to of a round, whereas taking that head from two passes to eight costs of one.
Appendix F Classic Two-Model Speculative Decoding
Classic two-model speculation runs on our target through the official Hugging Face assisted-generation path with three drafts, the image-conditioned Qwen3-VL-4B and Qwen3-VL-2B and a text-only Qwen3-1.7B in the spirit of the strong baseline of the first multimodal study (Gagrani et al., 2024). Table 11 reports all three. The B draft accepts long blocks, with from to rising with grounding, yet decodes at to , since a half-size draft costs roughly half a target forward for each drafted token and no acceptance amortizes that. The B draft accepts less and stays below throughout, so halving the draft again narrows the deficit without closing it. The text-only draft shows what vision access is worth, since its acceptance collapses to to . Unlike the text-only states of Figure 3, which the block head still reads, this external draft reads nothing of the target at all.
| Draft | Task | [ CI] | Speedup over AR [CI] | |
|---|---|---|---|---|
| Qwen3-VL-4B | Captioning | [] | [] | |
| TextVQA | [] | [] | ||
| InfoVQA | [] | [] | ||
| DocVQA | [] | [] | ||
| ChartQA | [] | [] | ||
| Qwen3-VL-2B | Captioning | [] | [] | |
| TextVQA | [] | [] | ||
| InfoVQA | [] | [] | ||
| DocVQA | [] | [] | ||
| ChartQA | [] | [] | ||
| Qwen3-1.7B, text-only | Captioning | [] | [] | |
| TextVQA | [] | [] | ||
| InfoVQA | [] | [] | ||
| DocVQA | [] | [] | ||
| ChartQA | [] | [] |
Appendix G Additional Law and Transfer Results
Fits on all five tasks.
Table 12 lists the entropy-acceptance fits, and Figure 6 draws them. InfoVQA is logged under the same streaming protocol as captioning, TextVQA, and DocVQA, while ChartQA is logged under its templated evaluation prompts and carries the steepest slope of the five. Both floors clear their own caps, against and against , and the share of rounds accepting eight or more tokens rises with grounding, at , , , , and from captioning to ChartQA.
| Task | Rounds | Ceiling fraction | Mean | |||||
|---|---|---|---|---|---|---|---|---|
| Captioning | ||||||||
| TextVQA | ||||||||
| InfoVQA | ||||||||
| DocVQA | ||||||||
| ChartQA, templated | ||||||||
| Pooled, captioning, TextVQA, DocVQA |
Brackets.
Figure 7 places each task’s fitted law between two independent bounds, the AdaEDL lower bound (Agrawal et al., 2024) and the verification ceiling at , and every fit stays inside both.
Entropy against image saliency.
Table 13 gives the correlations and decile gaps behind the law. Entropy correlates with acceptance on every task, while an image-saliency signal does not survive conditioning on entropy, with partial and gaps near within entropy quintiles.
| Quantity | Captioning | TextVQA | DocVQA |
|---|---|---|---|
| Spearman | |||
| low/high -decile gap | |||
| Spearman | |||
| partial | |||
| within-quintile gap |
Mismatched-image probe.
The correct image lengthens accepted blocks by on captioning with a interval of , by on TextVQA with , and by on DocVQA with . Once entropy is held fixed, the pooled coefficient of the residual image effect is .
Transfer across drafters, targets, and modalities.
Table 14 varies the target, and Figure 8 plots the fits. The two further Qwen targets use same-family two-model drafts, which are autoregressive, so the form holds for both kinds of drafter and is a property of the target. On Llama-3.2-11B-Vision, where vision enters by cross-attention rather than as a token prefix, GLANCE remains lossless and the law holds on captioning, and we state the characterization for prefix-fusion targets. Table 15 extends the form beyond vision to speech recognition, code, chat, and a non-Qwen backbone. Grounded vision-language generation and code follow the determinate-copy route of the two-state analysis, with long verbatim runs, while speech behaves like a uniform lift at every position, and both are limits of the same survival family.
| Target | Qwen2.5-VL | Qwen2-VL | Llama-3.2 |
|---|---|---|---|
| 7B | 7B | 11B-Vision | |
| Drafter | two-model | two-model | GLANCE |
| Curve | |||
| Mean |
| Setting | Model pair | |||
|---|---|---|---|---|
| speech (ASR) | Qwen2.5-Omni 7B3B | |||
| code | Qwen2.5-Coder 7B0.5B | |||
| chat (WildChat) | two-model | |||
| non-Qwen backbone | OLMo-2 13B7B |
Conditioning ablation.
Table 16 tabulates the values behind panel (b) of Figure 3 together with the zeroed-state control. Pooled over the five tasks and their prompts, fused visual conditioning lifts acceptance by , with a interval of , and the grounded tasks carry the lift.
| by what the head reads | ||||
|---|---|---|---|---|
| Task | fused VL | text-only | zeroed | text/fused |
| Captioning | ||||
| TextVQA | ||||
| InfoVQA | ||||
| DocVQA | ||||
| ChartQA | ||||
| Mean | ||||