The LLM Token Tax: the same sentence, 2–10× the tokens.
Every model splits text into tokens differently. For the same meaning, a German sentence, an Arabic sentence, a Hindi sentence cost far more tokens on some models than others — and tokens are the unit you pay for: API price, GPU time, and the real size of your context window. We measured 9 open tokenizers across 204 languages on 997 parallel sentences (FLORES-200) — including the just-released Kimi K2. This is an efficiency measurement, not a quality one.
Leaderboard — average token tax across 204 languages
Mean tokens-per-sentence relative to English, averaged over every language. Lower is better. Big-vocabulary, multilingual-first tokenizers win; 32k-vocab models are the most expensive off English.
The full table — token tax by language × model
Each cell = tokens for that language ÷ tokens for the same English sentences. 1.00× = English baseline; 2.00× = twice the cost for the same meaning. Hover a cell for detail.
| Qwen2.5 | Llama-3.1 | Llama-2 | Mistral-v0.3 | Gemma-2 | DeepSeek-V3 | BLOOM | Phi-3.5 | |
|---|---|---|---|---|---|---|---|---|
| English | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| Chinese | 1.02 | 1.30 | 2.02 | 1.64 | 1.09 | 0.96 | 0.94 | 2.02 |
| Spanish | 1.52 | 1.55 | 1.46 | 1.58 | 1.26 | 1.50 | 1.21 | 1.46 |
| Arabic | 1.63 | 1.66 | 3.43 | 3.48 | 1.50 | 1.62 | 1.16 | 3.43 |
| Hindi | 4.45 | 2.54 | 4.62 | 4.58 | 1.86 | 2.82 | 1.30 | 4.62 |
| Portuguese | 1.46 | 1.48 | 1.43 | 1.55 | 1.23 | 1.43 | 1.12 | 1.43 |
| Russian | 1.77 | 1.64 | 1.66 | 1.89 | 1.40 | 1.61 | 2.50 | 1.66 |
| Japanese | 1.41 | 1.52 | 2.24 | 2.15 | 1.17 | 1.43 | 1.79 | 2.24 |
| German | 1.55 | 1.57 | 1.41 | 1.57 | 1.25 | 1.51 | 1.68 | 1.41 |
| French | 1.56 | 1.58 | 1.46 | 1.59 | 1.34 | 1.53 | 1.19 | 1.46 |
| Korean | 1.64 | 1.49 | 3.20 | 2.47 | 1.64 | 1.66 | 2.78 | 3.20 |
| Italian | 1.62 | 1.63 | 1.46 | 1.60 | 1.33 | 1.54 | 1.61 | 1.46 |
| Turkish | 1.61 | 1.40 | 2.08 | 2.20 | 1.39 | 2.02 | 1.96 | 2.08 |
| Vietnamese | 1.43 | 1.38 | 2.93 | 2.92 | 1.36 | 2.11 | 1.27 | 2.93 |
| Polish | 1.77 | 1.89 | 1.71 | 1.90 | 1.46 | 1.69 | 2.14 | 1.71 |
| Ukrainian | 2.52 | 1.66 | 1.73 | 1.95 | 1.67 | 2.18 | 2.75 | 1.73 |
| Thai | 2.56 | 2.21 | 4.37 | 4.22 | 1.83 | 1.78 | 4.59 | 4.37 |
| Greek | 4.90 | 2.25 | 5.01 | 5.25 | 2.29 | 2.70 | 3.79 | 5.01 |
| Bengali | 5.05 | 5.79 | 5.40 | 4.93 | 2.69 | 2.02 | 1.18 | 5.40 |
| Swahili | 1.93 | 1.93 | 1.86 | 1.94 | 1.60 | 1.94 | 1.24 | 1.86 |
The just-released Kimi K2 is in the leaderboard above (4th of 9) and gets its own head-to-head write-up — what a 1-trillion-parameter model's tokenizer actually costs → The heatmap keeps the original eight for readability.
What we are — and are not — claiming
This is efficiency, not quality. A high token tax does not mean a model is “bad” at that language — it means its tokenizer is inefficient there, so that language costs more and fits less context. A model can be excellent at Hindi and still be token-expensive at it; the two are separate measurements and we do not conflate them.
This is a tokenizer measurement, not our internal X-Ray scan. It needs no weights and no GPU — just each model’s public tokenizer — which is exactly why we can run it for every open model, in the open, and why you can reproduce it yourself. Our layer-level X-Ray is a different, deeper instrument; this table is its lightweight companion.
Method (reproducible)
Corpus: FLORES-200 dev — 997 sentences, professionally translated into every language, so each column measures the same meaning. Metric: total tokens per language ÷ total tokens for the identical English sentences, per tokenizer (add_special_tokens=False). Tokenizers: Qwen2.5, Llama-3.1, Llama-2, Mistral-v0.3, Gemma-2, DeepSeek-V3, BLOOM, Phi-3.5 (public, unmodified). One caveat on the corpus: FLORES is news/encyclopedic prose; code, chat and domain text will differ. Token count is a property of the tokenizer, so it is identical across model sizes within a family — a 0.5B and a 32B of the same family tokenize a sentence the same way.
Scan your model → building a multilingual model? See your own language economy before and after you extend the vocab.
← X-Ray for LLMsAll research notes
— Tetracta AI Teams · for humans, like humans.