For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Speculative decoding
Speculative decoding accelerates LLM token generation without changing the model's outputs. A smaller, faster draft step proposes several candidate tokens, and the larger target model verifies them in one forward pass. MAX accepts the prefix the target agrees with and resamples at the first disagreement, so quality matches running the target alone.
The speedup comes from batching verification across K candidate positions.
When the target accepts all K drafts, you get K+1 tokens per step instead
of one, converting a memory-bandwidth-bound workload into one that better
uses available compute.
Supported methods
MAX supports common speculative-decoding methods like EAGLE, EAGLE3, MTP, and DFlash.
| Method | Draft source | Supported targets and hardware |
|---|---|---|
eagle | A trained EAGLE or EAGLE3 draft that shares the target's embedding and lm_head | Llama 3 (1 GPU), Kimi K2.5 and K2.6 (8× B200). |
mtp | A multi-token prediction head, either inside the target checkpoint or as an assistant draft | DeepSeek V3 and derivatives (8× B200), GLM-5.2 (8× B200), Gemma 4 (1 GPU), Inkling (8× B200). |
dflash | A trained block draft (DFlash or DSpark) paired with the target | Llama 3 (1 GPU), Gemma 4 (1 GPU), Kimi K2.5 (8× B200). |
Serve with speculative decoding
Pick a tab for the method you want to run. Each example starts max serve with the right target and draft, then sends a chat-completion
request to the local endpoint from the OpenAI Python client.
To call the endpoint, install the OpenAI Python client:
- pixi
- uv
- pip
- conda
pixi add openaiuv add openaipip install openaiconda install openai- EAGLE
- MTP
- DFlash
Serve Llama 3.1 8B Instruct with a pretrained EAGLE checkpoint as the
draft:
max serve \
--model meta-llama/Llama-3.1-8B-Instruct \
--speculative-method eagle \
--draft-model-path atomicapple0/EAGLE-LLaMA3.1-Instruct-8B \
--num-speculative-tokens 2 \
--devices gpuWe use atomicapple0/EAGLE-LLaMA3.1-Instruct-8B because it ships
safetensors weights in bfloat16. Some EAGLE repos ship float16 PyTorch
.bin checkpoints instead, which MAX doesn't support. For the full
list of weight formats MAX supports, see
WeightsFormat.
The endpoint is ready when you see this message printed in your terminal:
Server ready on http://0.0.0.0:8000 (Press CTRL+C to quit)Send a chat-completion request:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="EMPTY",
)
response = client.chat.completions.create(
model="meta-llama/Llama-3.1-8B-Instruct",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What are the benefits of speculative decoding?"},
],
max_completion_tokens=500,
)
print(response.choices[0].message.content)DeepSeek V3 and GLM-5.2 bake the MTP head into the target checkpoint, so you don't pass a separate draft model:
max serve \
--model deepseek-ai/DeepSeek-V3 \
--speculative-method mtp \
--num-speculative-tokens 2 \
--devices gpu:0,1,2,3,4,5,6,7The endpoint is ready when you see this message printed in your terminal:
Server ready on http://0.0.0.0:8000 (Press CTRL+C to quit)Send a chat-completion request:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="EMPTY",
)
response = client.chat.completions.create(
model="deepseek-ai/DeepSeek-V3",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What are the benefits of speculative decoding?"},
],
max_completion_tokens=500,
)
print(response.choices[0].message.content)Serve Llama 3.1 8B Instruct with a pretrained DFlash checkpoint as
the draft:
max serve \
--model meta-llama/Llama-3.1-8B-Instruct \
--speculative-method dflash \
--draft-model-path z-lab/LLaMA3.1-8B-Instruct-DFlash-UltraChat \
--devices gpuDFlash drafts are trained at a fixed block size, and MAX sets
--num-speculative-tokens to the trained width, one less than the
block size. If you pass a different value, MAX overrides it and logs
a warning such as DFlash draft was trained at block_size=10, so num_speculative_tokens is being overridden from 2 to 9.
DSpark drafts use the same serving path.
The endpoint is ready when you see this message printed in your terminal:
Server ready on http://0.0.0.0:8000 (Press CTRL+C to quit)Send a chat-completion request:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="EMPTY",
)
response = client.chat.completions.create(
model="meta-llama/Llama-3.1-8B-Instruct",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What are the benefits of speculative decoding?"},
],
max_completion_tokens=500,
)
print(response.choices[0].message.content)Supported model architectures
When you serve a model with a draft, MAX loads a dedicated speculative
decoding architecture for the pair and prints its name in the startup
banner. For example, serving deepseek-ai/DeepSeek-V3 with
--speculative-method mtp resolves to UnifiedMTPDeepseekV3ForCausalLM.
The name comes from the target's architecture and the draft checkpoint's
declared architecture:
| Target | Draft | Resolved architecture |
|---|---|---|
| Llama 3 | EAGLE draft | UnifiedEagleLlama3ForCausalLM |
| Llama 3 | DFlash draft | UnifiedDflashLlama3ForCausalLM |
| Gemma 4 | Assistant draft (Gemma4AssistantForCausalLM) | UnifiedMTPGemma4ForCausalLM |
| Gemma 4 31B | DFlash draft | UnifiedDflashGemma4_31BForCausalLM |
| Gemma 4 31B | DSpark draft | UnifiedDSparkGemma4_31BForCausalLM |
| Gemma 4 12B | DSpark draft | UnifiedDSparkGemma4_12BForCausalLM |
| DeepSeek V3 | None (built-in MTP head) | UnifiedMTPDeepseekV3ForCausalLM |
| GLM-5.2 | None (built-in MTP head) | UnifiedMTPGlmMoeDsaForCausalLM |
| Kimi K2.5 | DFlash draft | UnifiedDflashKimiK25ForCausalLM |
Monitor acceptance rates
Once you send traffic, the scheduler logs per-batch acceptance stats in this format:
Draft Tokens: 145/160 (90.62%) accepted, Acceptance Len: 1.45 / 2 toks (accepted drafts/step; 2.45 toks/step incl bonus), Per-Pos: [p0=95%, p1=86%] |Each field means:
- Accepted / generated: tokens the target confirmed, over tokens the draft proposed.
- Acceptance length: average number of drafted tokens accepted per
verification pass. A value of
1.45 / 2means on average 1.45 of the 2 drafted tokens survive verification. The parenthetical counts the bonus token the target produces after them: 2.45 tokens per step including it. - Per-position: acceptance rate at each draft position, conditional on all earlier positions accepting. Later positions are always rarer.
Low acceptance rates (below roughly 50%) usually mean the draft doesn't
match the target well. Try a smaller --num-speculative-tokens or a
better-matched draft checkpoint.
Tune speculative decoding
The following flags control how MAX drafts tokens and how verification decides to accept them. The Config column shows the config class each flag maps to when you configure a pipeline programmatically.
| Flag | Description | Config |
|---|---|---|
--num-speculative-tokens | Number of tokens the draft proposes per step. Defaults to 2 for EAGLE and MTP. For DFlash-style block drafters, MAX reads the value from the draft checkpoint's trained width. If the checkpoint doesn't declare a width, you must set this flag explicitly. Larger values raise peak speedup but hurt acceptance at later positions. | SpeculativeConfig |
--num-speculative-tokens-per-batch-size | How many of the drafted tokens the target verifies, set by decode batch-size range. Takes a JSON list of inclusive ranges, such as [{"batch_start": 1, "batch_end": 16, "num_tokens": 3}, {"batch_start": 17, "batch_end": 64, "num_tokens": 1}]. | SpeculativeConfig |
--synthetic-acceptance-rate | Benchmarking-only knob that accepts each drafted token with a calibrated probability, ignoring real logits. Use this to model hypothetical speedups without changing the draft. | SpeculativeConfig |
--draft-proposal | How the draft model proposes tokens. argmax (default) selects tokens deterministically. sampled draws from the draft model's distribution and preserves that distribution for verification. sampled requires a GPU, a static vocabulary size, and a serving architecture that supports it. | SpeculativeConfig |
--use-relaxed-acceptance-for-thinking | Whether to accept a draft token inside a <think>...</think> block when it matches any of the target's top candidates within a probability threshold, instead of requiring an exact match. Outside the thinking span, the strict acceptance rule still applies. Requires --draft-proposal argmax and a serving architecture that supports relaxed acceptance. Defaults to false. | SpeculativeConfig |
--relaxed-topk | The number of top candidates from the target distribution to consider when relaxed acceptance is active. Ignored when relaxed acceptance is off. Defaults to 10. | SpeculativeConfig |
--relaxed-delta | The probability gap below the top candidate within which candidates remain eligible for relaxed acceptance. Ignored when relaxed acceptance is off. Defaults to 0.6. | SpeculativeConfig |
--use-greedy-acceptance | Whether to use greedy (argmax) draft acceptance instead of the stochastic sampler, which lets MAX capture the fused speculative graph with CUDA graphs. Valid only for greedy serving (temperature 0, top-k 1), and incompatible with relaxed and synthetic acceptance. Requires a serving architecture that supports it. Defaults to false. | SpeculativeConfig |
--draft-devices | Device list for the draft model. Useful when you want the draft and target on different GPUs. | Draft model config |
--device-memory-utilization | Fraction of device memory MAX may use. Speculative decoding allocates KV cache for both the target and the draft, so leave more headroom than you would for single-model serving. | KVCacheConfig |
MAX auto-enables the overlap scheduler for every speculative architecture. It also auto-enables device graph capture on the architectures that support it. Both reduce per-step latency and need no additional flags.
Compatibility and limits
When using speculative decoding:
- The
--enable-echoflag isn't supported. - The
--max-lengthand--draft-max-lengthare both capped at the draft model's maximum supported length. MAX clamps both to that limit if either exceeds it. - Structured output is supported. MAX applies the constraint bitmask to each speculated position during verification. This feature requires a GPU deployment.
- Repetition, frequency, and presence penalties aren't supported when
using a separate draft model (
--draft-model-path), including EAGLE. Repetition, frequency, and presence penalties are supported when using MTP.
Next steps
You can combine speculative decoding with prefix caching and with disaggregated inference. The following topics go deeper on performance and deployment:
- Prefix caching: Use prefix caching when serving a model with the MAX CLI.
- Benchmark MAX on NVIDIA or AMD GPUs: Learn how to use our benchmarking script to measure the performance of MAX.
- Deploy MAX on GPU with self-hosted endpoints: Learn how to deploy MAX pipelines to cloud.
- Speculative decoding: Learn more about it in the LLM Inference Handbook.