Removable and Irreducible:
A Token-Cost Ledger for the Multilingual Tokenization Tax
Abstract
Large language models pay a well-documented tax on non-English text: the same content costs several times more tokens, and because attention is quadratic in sequence length, far more compute. We ask how much of this tax is removable. Framing the token layer as source coding – transformer compute is monotone in sequence length, whose per-atom floor is the Shannon rate , an object already applied to tokenizers in prior work – we assemble a token-cost ledger that splits each language’s cost, at fixed parallel content, into a removable coding redundancy, a residual coding slack, an intrinsic-content term, and an orthogonal, irreducible grapheme-to-phoneme term that governs the multimodal rather than the text cost. On FLORES-200 across eight languages, a production tokenizer costs up to more tokens for Indic scripts than for English; a script-matched code trained on sentences removes a median of that excess (bootstrap 95% CI ), and a script-fair information floor shows the intrinsic content differs by under – the tax is representational, not informational. A constructed code removes of a controlled source’s redundancy, and the token tax implies up to attention cost. We are explicit about scope and failure: this is compute-and-memory accounting, not a model-quality claim; we neither measure nor claim the cross-lingual direction of the orthographic term; and our matched code is a conservative small-data demonstration. We contribute the unifying ledger, the removable-versus-intrinsic attribution, and an open one-command harness.
1 Introduction
It costs more to say the same thing to a language model in Telugu than in English. On identical, professionally-translated content, a widely-used production tokenizer emits up to as many tokens for an Indic script as for English (Section 4); an older English-centric tokenizer reaches – per word. Because self-attention is in sequence length (Vaswani et al., 2017), a token inflation is a attention-compute inflation and a shrink of the effective context window. This multilingual “token tax” is well documented as a phenomenon and a fairness problem (Ahia et al., 2023; Petrov et al., 2023; Lundin and others, 2025).
The question this paper asks is narrower and, we think, more useful: how much of the tax is removable by choosing a better code, and how much is intrinsic to the language? We answer it by assembling a token-cost ledger. Holding semantic content fixed with a parallel corpus, we decompose each language’s normalized sequence length into (i) a removable coding redundancy – the penalty for using a code fit to the wrong (English-dominated) distribution; (ii) a residual coding slack from a finite, finitely-trained vocabulary; (iii) an intrinsic-content term measured by a script-fair compressor; and, on an orthogonal axis, (iv) an irreducible grapheme-to-phoneme term that does not affect text sequence length at all but governs the multimodal (speech) ledger. Terms (i)–(iii) live on the text-compute axis; (iv) is the organic cost of a deep orthography.
None of the ledger’s individual pieces is new. The fertility floor is Shannon source coding (Shannon, 1948; Cover and Thomas, 2006), applied to tokenizers as an “efficiency” or “capacity-utilization” metric by Zouhar et al. (2023) and Erdogan et al. (2026), the latter of whom also give the additive floor-plus-redundancy split we use. Orthographic depth as an information-theoretic quantity is established (Torres and Futrell, 2025). A compute-optimal vocabulary size is derived by Tao et al. (2024). Our contribution is (a) to unify the removable text term and the irreducible orthographic term into one cost ledger; (b) to attribute the observed multilingual tax empirically to removable versus intrinsic components on real parallel data, with a constructed code that drives the removable term to zero; and (c) to keep the whole thing scoped strictly to compute and memory, which is what makes the accounting well-posed.
Scope fence (stated once, up front).
This is a cost accounting, not a quality claim. We do not assert that fewer tokens make a better model; Schmidt et al. (2024) and others show it need not, and we agree. Everything here concerns FLOPs, KV-cache bytes, and context-window occupancy, quantities that are monotone in sequence length regardless of downstream accuracy.
Contributions.
-
•
A unified token-cost ledger (Section 2) that puts the removable coding redundancy and the irreducible orthographic term on one accounting identity, with a single estimand – the removable fraction – for “how synthetic is this language’s tax?”
-
•
An empirical attribution (Section 4) on FLORES-200: a script-matched code trained on sentences removes a median of the production-tokenizer tax (CI ), and the script-fair information floor shows intrinsic content varies by across Indic languages – the tax is representational.
-
•
A constructed floor-approaching code and the quadratic-cost consequence (Section 5): a matched code removes of a controlled source’s redundancy, sitting bits above the entropy floor; the token tax implies up to attention cost.
-
•
An open, one-command harness and an honest limitations section (Section 7), including a pre-registered control that we report as not established.
2 The token-cost ledger
Cost is monotone in sequence length.
A decoder-only transformer of width and layers processing tokens pays attention FLOPs, feed-forward FLOPs, KV-cache memory, and an vocabulary term (Vaswani et al., 2017).
Proposition 1 (cost monotonicity).
For fixed , transformer compute and KV memory for a fixed piece of content are non-decreasing in , and strictly increasing once (the attention term dominates).
Every term is non-decreasing in ; the term’s derivative overtakes the linear terms once . The consequence is the objective: minimize expected for fixed content. In the quadratic regime a sequence-length ratio is an attention-cost ratio .
The floor (prior art).
Model content as atoms drawn from a source with entropy bits/atom, encoded to tokens over a -ary alphabet.
Lemma 1 (fertility floor; Cover and Thomas, 2006; Zouhar et al., 2023; Erdogan et al., 2026).
Any uniquely-decodable -ary code has expected tokens per atom , with equality approached within one token by a matched Huffman code and exactly in the block limit. If the code is optimal for a wrong distribution , then with .
We state Lemma 1 for self-containedness and attribute it; is Erdogan et al.’s capacity utilization, and the split is their Appendix-C decomposition. is the removable part: refit the code to and .
The ledger (our synthesis).
Hold content fixed with a parallel corpus and normalize every quantity to a reference language (English ). Write for the total tokens of language under encoder divided by English’s under the same encoder. Then the excess of a production encoder decomposes additively into individually measurable terms that are non-negative for every language whose tax we attribute:111The identity telescopes exactly for all languages; the residual-slack term can nonetheless turn marginally negative for a shallow-orthography Latin control whose intrinsic content already sits near the English baseline. German is the one such case here (): the small matched BPE and the LZMA content proxy are distinct instruments, each normalized to English under itself, and agree to within noise this close to . For all five Indic languages—our focus—every term is strictly positive (Table 1).
| (1) |
and, on an orthogonal axis, an irreducible term that leaves text untouched but governs the speech ledger. The unification of the removable text term with the irreducible orthographic term is the contribution; neither half is ours (the removable term is Erdogan et al., 2026; orthographic depth as algorithmic mutual compressibility is Torres and Futrell, 2025, whose Kolmogorov instrument we replace with a Shannon one in the LLM-cost setting). The headline estimand is the removable fraction
| (2) |
the share of the production tax a language-matched code removes: means the tax is a code artifact (synthetic), means it is intrinsic (organic).
A one-line remark on vocabulary.
3 Experimental setup
Real parallel data. FLORES-200 (NLLB Team et al., 2022) (CC-BY-SA), professionally translated and sentence-aligned, so line of every language file is the same sentence. We report on eight languages: a Latin control (English, Spanish, German), Devanagari (Hindi), Bengali, and three Dravidian abugidas (Telugu, Tamil, Kannada). We measure on the -sentence devtest split. Encoders. UTF-8 bytes; Unicode extended grapheme clusters (\X, the akshara-preserving atom); three production byte-level BPE tokenizers (GPT-2, cl100k_base, o200k_base); and a per-language matched BPE trained held out on the FLORES dev split and evaluated on devtest. Information floor. A script-fair content estimate: LZMA over the grapheme-cluster-ID stream (not UTF-8 bytes), so a script is not charged for its 3-byte-per-codepoint UTF-8 assignment. Constructed source. A Zipf() source over concepts with a known entropy. Apparatus. Single-thread Python; tiktoken and the standalone tokenizers library (no GPU, no model weights); deterministic seeds; one command regenerates every number and figure.
Pre-registration.
Hypotheses, decision rules, and four negative controls were frozen before the confirmatory run. The calibration control (NC2) requires the stack to recover a known entropy: on a uniform-32 source it returns bits and a Huffman length in ; it passed before any language number was read.
4 Results: the removable tax
Table 1 and Figure 1 report the ledger. Under cl100k_base the production tax reaches (Telugu), (Kannada), and (Hindi); under the older GPT-2 tokenizer, fertility per word reaches – for Dravidian scripts. A script-matched code trained on only sentences (learned vocabulary –) brings these to , , and ; and the script-fair information floor sits at – – the intrinsic content of the same sentences is within across all Indic languages (this is our pre-registered content-invariance control, NC3: no language exceeded the threshold, so no part of the tax is re-attributed to intrinsic content). Bootstrapping sentences, the median removable fraction across the five Indic languages is with a CI of : a matched code removes about two-thirds of the production tax, and the information floor shows nearly all of the rest is coding slack rather than content. That production tokenizers themselves disagree by on the same Telugu content ( under cl100k vs. under the more multilingual o200k) is independent evidence that the tax is a property of the code, not the language.
| Language | bytes/char | ||||
|---|---|---|---|---|---|
| English | — | ||||
| Spanish | |||||
| German | |||||
| Hindi | |||||
| Bengali | |||||
| Telugu | |||||
| Tamil | |||||
| Kannada | |||||
| Indic median | 0.64 CI[0.638,0.647] |
Bytes-per-char predicts the tax (H2, a replication).
5 Decomposition, a constructed code, and quadratic cost
Figure 2 decomposes the Indic excess of Eq. 1. The removable term (vocabulary mismatch) dominates; the intrinsic-content sliver is in every case. The residual coding slack – the gap between our small matched code and the floor – is itself removable in principle (a better-trained code closes it), so it is a lower bound on removability, not a second intrinsic term; that a k-vocabulary production tokenizer (o200k) already reaches on Telugu confirms the slack is training, not content.
A constructed code hits the floor.
On the controlled Zipf source (entropy bits/concept), a mismatched fixed-width code pays bits/concept; the matched Huffman code (Huffman, 1952) (“Silicon Vernacular”) pays – bits above the floor, within Lemma 1’s one-bit guarantee – removing of the redundancy (Figure 3, left). Our pre-registered no-sub-floor control (NC1) requires that no code we report encodes below , and the Kraft sum of every code we build satisfies . We therefore deliberately do not report an empirical block code beating the floor: at block sizes the -gram alphabet is undersampled and an in-sample Huffman would appear to beat from finite-sample bias – a dishonest number. The residual sub-bit gap is closed by block/arithmetic coding as a theorem, not a measurement. We position this construction as the discrete-code limit of byte-entropy patching (Pagnoni et al., 2024), not as a proposed human language.
Quadratic amplification.
Because attention is , the token tax is amplified in compute: the same content costs up to (Kannada) and (Telugu) the attention work of English under cl100k, collapsing to – under the matched code (Figure 3, right).
6 Related work
Information theory of tokenizers. Zouhar et al. (2023) frame the tokenizer as a channel and define an entropy-to-max-entropy efficiency (the parent of ); Erdogan et al. (2026) define capacity utilization and the additive floor-plus-redundancy decomposition we build on; Rajaraman et al. (2024) and Gastaldi et al. (2025) give complementary theories of tokenization. We contribute neither the floor nor the split – we contribute their use as a multilingual accounting. The multilingual tax. Ahia et al. (2023), Petrov et al. (2023), and Lundin and others (2025) document the cross-language cost disparity; Indic-specific tokenizers (Rana et al., 2025) reduce fertility. We explain the disparity as a removable KL redundancy above a script-fair floor – and note that Petrov et al.’s residual byte-level disparity is exactly the intrinsic + orthographic remainder our ledger predicts. Byte-level models. MEGABYTE (Yu et al., 2023) and the Byte Latent Transformer (Pagnoni et al., 2024) remove the discrete code and let entropy set patch boundaries; our ledger explains what they can remove (the redundancy ) and cannot (intrinsic content, orthographic ). Compression is not quality. Schmidt et al. (2024) and Bostrom and Durrett (2020) show fewer tokens need not improve models; our scope fence makes this orthogonal – we account for cost, not quality. Orthographic depth. Torres and Futrell (2025) formalize transparency as Kolmogorov mutual compressibility; we borrow the quantity (as Shannon ) for the irreducible axis. Optimal vocabulary. Tao et al. (2024) and Limisiewicz et al. (2026) derive the compute-optimal vocabulary/granularity we merely cite.
7 Limitations and honest negatives
(1) We did not establish the orthographic direction. We pre-registered a control (NC4) that English’s grapheme-to-phoneme ambiguity exceeds shallow Indic scripts’. We measure only the English side (CMUdict homograph entropy bits/type); we have no Indic pronunciation lexicon, so the cross-lingual direction is literature-consistent but not established by our instrument. Per the pre-registration we report this as a negative and future work, not a result. (2) The matched code is a small-data demonstration, trained on 1k sentences (vocabulary –k); it understates removability, making a lower bound. (3) The information floor is an LZMA estimate, an upper bound on intrinsic content; a tighter estimator would shrink the intrinsic term further, again in the direction of “more removable.” (4) This is not a quality claim. We measure compute and memory; we make no statement about accuracy or loss, and explicitly do not contradict Schmidt et al. (2024). (5) Scope: eight languages, one parallel benchmark, text only; speech/multimodal cost is argued, not measured. (6) We prove no new theorem: the floor and its redundancy split are cited, not claimed.
8 Conclusion
The multilingual tokenization tax is, in the compute ledger, mostly removable: on real parallel text a script-matched code trained on a thousand sentences erases about two-thirds of it, and a script-fair floor shows the intrinsic content of the same sentences differs by under six percent. What remains – the organic grapheme-to-phoneme cost of a deep orthography – is real but lives on a different (multimodal) axis, and flips the ranking. We contribute the unifying ledger, the removable-versus-intrinsic attribution, a constructed code that reaches the entropy floor, and an open harness; we claim neither the floor, nor the optimal vocabulary, nor a model-quality benefit. The larger program these results open – design principles of a near-optimal language for foundational models, spanning grammar and attention routing, acoustic isomorphism, and human learnability – we name as future work and a thesis, not a claim of this paper.
Reproducibility.
Code, data pointers, pre-registration, and one-command reproduction: https://github.com/samyama-ai/token-cost-ledger.
References
- Do all languages cost the same? tokenization in the era of commercial language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), Cited by: §1, §4, §6.
- Byte pair encoding is suboptimal for language model pretraining. In Findings of the Association for Computational Linguistics: EMNLP 2020, Note: arXiv:2004.03720 Cited by: §6.
- Elements of information theory. 2nd edition, Wiley-Interscience. Cited by: §1, Lemma 1.
- An information-theoretic perspective on llm tokenizers. arXiv preprint arXiv:2601.09039. Cited by: §1, §2, §2, §6, Lemma 1.
- The foundations of tokenization: statistical and computational concerns. In International Conference on Learning Representations (ICLR), Note: arXiv:2407.11606 Cited by: §6.
- A method for the construction of minimum-redundancy codes. Proceedings of the IRE 40 (9), pp. 1098–1101. Cited by: §5.
- Compute optimal tokenization. arXiv preprint arXiv:2605.01188. Cited by: §2, §6.
- The token tax: systematic bias in multilingual tokenization. arXiv preprint arXiv:2509.05486. Cited by: §1, §6.
- No language left behind: scaling human-centered machine translation. arXiv preprint arXiv:2207.04672. Note: FLORES-200 evaluation benchmark Cited by: §3.
- Byte latent transformer: patches scale better than tokens. arXiv preprint arXiv:2412.09871. Cited by: §5, §6.
- Language model tokenizers introduce unfairness between languages. In Advances in Neural Information Processing Systems (NeurIPS), Note: arXiv:2305.15425 Cited by: §1, §4, §6.
- Toward a theory of tokenization in LLMs. arXiv preprint arXiv:2404.08335. Cited by: §6.
- IndicSuperTokenizer: an optimized tokenizer for indic multilingual LLMs. arXiv preprint arXiv:2511.03237. Cited by: §6.
- Tokenization is more than compression. arXiv preprint arXiv:2402.18376. Cited by: §1, §6, §7.
- A mathematical theory of communication. Bell System Technical Journal 27, pp. 379–423. Cited by: §1.
- Scaling laws with vocabulary: larger models deserve larger vocabularies. In Advances in Neural Information Processing Systems (NeurIPS), Note: arXiv:2407.13623 Cited by: §1, §2, §6.
- Clarifying orthography: orthographic transparency as compressibility. arXiv preprint arXiv:2505.13657. Cited by: §1, §2, §6.
- Attention is all you need. Advances in Neural Information Processing Systems (NeurIPS). Cited by: §1, §2.
- MEGABYTE: predicting million-byte sequences with multiscale transformers. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: §6.
- Tokenization and the noiseless channel. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL), Note: arXiv:2306.16842 Cited by: §1, §6, Lemma 1.