LanguagebeginnerGPU optional~25 minColab

Tokenizers and the Arabic tax

المُجزِّئات والضريبة العربية

Fertility is a price, paid per token

The same sentence costs more in Arabic than in English. Not metaphorically — measurably, in tokens, and tokens are what you pay for and what fills your . So an Arabic reader pays twice: a higher bill for identical content, and a shorter effective context from the same model.

What this workshop measures is not just that the gap exists but how much of it is anyone's fault. Run it and you will find the same Arabic sentences costing one production about four times what they cost another — a factor of 4.505 on the worst of the three. That spread is the whole point: the tax is not a property of Arabic, it is a procurement decision, and somebody made it.

The goal

Measure fertility — tokens per word — for the same parallel text in Arabic and English across three production tokenizers, convert the ratio into what it actually costs in context window and inference price, and find the single change to the text that moves the number most.

Colab opens a read-only copy. Save a copy to Drive to keep your edits.

The notebook needs a keyboard — best opened on a desktop.

The papers behind this

Everything a language model charges you for is counted in tokens. Context windows are quoted in tokens, prices are quoted per million tokens, and rate limits are enforced in tokens. So the question "how many tokens does this sentence cost?" is not a technical curiosity — it is the bill. A tokenizer decides that number, and it decides it by having been trained on some particular mixture of text. Whatever was common in that mixture became whole pieces. Whatever was rare became fragments.

Preflight, then the — verified before anything is measured.

import azimuth_nb as azimuth

env = azimuth.setup(SLUG, lang=LANG, profile=PROFILE)
Workshop code
Tokenizers and the Arabic tax
NVIDIA GeForce GTX 1070 with Max-Q Design · 8.0 GB · 15.9 GB RAM · PyTorch 2.6.0+cu124
profile: free
assets:
  · parallel.tsv — already present
ready · sampleSentences=1000, pricePerMillionTokens=3.0, contextWindow=8192, seed=17

Word counts first. is a ratio, and the denominator has to be defined before the numerator means anything.

import random

ARABIC_RANGE = ("\u0600", "\u06ff")


def arabic_share(text):
    """Fraction of letters that are Arabic script."""
    letters = [ch for ch in text if ch.isalpha()]
    if not letters:
        return 0.0
    return sum(ARABIC_RANGE[0] <= ch <= ARABIC_RANGE[1] for ch in letters) / len(letters)


rows = []
with open(env.assets["parallel.tsv"], encoding="utf-8") as fh:
    for line in fh:
        parts = line.rstrip("\n").split("\t")
        if len(parts) >= 2 and parts[0].strip() and parts[1].strip():
            rows.append((parts[0].strip(), parts[1].strip()))

# WHICH COLUMN IS ARABIC IS DETECTED, NOT ASSUMED.
#
# An earlier run of this workshop measured a corpus whose columns were the
# other way round and reported that Arabic was CHEAPER than English — a
# confident, precise, exactly-backwards result. Nothing failed: both columns
# are text, both tokenize, and the arithmetic is identical. Only the sign of
# the conclusion changed.
#
# Reading the script itself costs one pass and removes the assumption. A
# column order is a property of whoever exported the file; the alphabet is a
# property of the language.
sample = rows[: min(200, len(rows))]
share_0 = sum(arabic_share(a) for a, _ in sample) / len(sample)
share_1 = sum(arabic_share(b) for _, b in sample) / len(sample)

if max(share_0, share_1) < 0.5:
    raise SystemExit(
        "Neither column looks like Arabic script — check the encoding of "
        "parallel.tsv before trusting anything below."
    )

if share_0 >= share_1:
    pairs_raw = rows
else:
    pairs_raw = [(b, a) for a, b in rows]
    print("note: columns were (english, arabic) — swapped to match")

# A fixed seed so the sample — and therefore every number below — is the same
# on your machine, on CI, and on the page.
random.seed(env.cfg["seed"])
limit = env.cfg["sampleSentences"]
pairs = random.sample(pairs_raw, min(limit, len(pairs_raw))) if limit else pairs_raw

# WHITESPACE, deliberately. "Word" has to mean the same operation in both
# languages or the ratio compares nothing. Whitespace undercounts Arabic
# morphology — one written word often carries article, stem and pronoun — so
# this makes the result a LOWER BOUND on the real gap rather than a flattering
# one. Overstating the case would be the easiest way to lose the argument.
ar_words = sum(len(ar.split()) for ar, _ in pairs)
en_words = sum(len(en.split()) for _, en in pairs)

if env.lang == "ar":
    print(f"أزواج: {len(pairs):,} من أصل {len(pairs_raw):,}")
    print(f"كلمات: {ar_words:,} عربية · {en_words:,} إنجليزية")
else:
    print(f"pairs: {len(pairs):,} of {len(pairs_raw):,}")
    print(f"words: {ar_words:,} Arabic · {en_words:,} English")
Workshop code
note: columns were (english, arabic) — swapped to match
pairs: 1,000 of 24,999
words: 8,676 Arabic · 9,995 English

Three vocabularies, three different bets about which languages matter.

from tokenizers import Tokenizer

# Three different bets about which languages deserve vocabulary budget.
# Loaded from the Hub by name — small JSON files, no model weights.
SPECS = [
    ("gpt2", "openai-community/gpt2", "50k pieces, almost all English"),
    ("bloom", "bigscience/bloom", "250k pieces shared across 46 languages"),
    ("xlm-roberta", "FacebookAI/xlm-roberta-base", "250k pieces, 100 languages"),
]

tokenizers = {}
for name, repo, note in SPECS:
    try:
        tokenizers[name] = Tokenizer.from_pretrained(repo)
        print(f"  · {name:12} {note}")
    except Exception as exc:
        # A tokenizer that will not load is reported and skipped, not fatal:
        # the comparison still means something with two, and a dead Hub should
        # not cost you the whole workshop.
        print(f"  ! {name:12} unavailable ({type(exc).__name__}) — skipped")

n_tokenizers = len(tokenizers)
Workshop code
  · gpt2         50k pieces, almost all English
  · bloom        250k pieces shared across 46 languages
  · xlm-roberta  250k pieces, 100 languages

Read DOWN the ratio column, not across one row. The spread between the tokenizers is the finding — the same Arabic sentences cost one of them four times what they cost another.

def count_tokens(tok, texts):
    """Total pieces across a list of strings, no special tokens."""
    return sum(len(tok.encode(t, add_special_tokens=False).ids) for t in texts)


ar_texts = [ar for ar, _ in pairs]
en_texts = [en for _, en in pairs]

table = []
for name, tok in tokenizers.items():
    ar_tokens = count_tokens(tok, ar_texts)
    en_tokens = count_tokens(tok, en_texts)
    ar_f = ar_tokens / ar_words
    en_f = en_tokens / en_words
    table.append(
        {
            "tokenizer": name,
            "ar_fertility": round(ar_f, 3),
            "en_fertility": round(en_f, 3),
            "ratio": round(ar_f / en_f, 3),
        }
    )

header = f"{'tokenizer':14}{'ar/word':>10}{'en/word':>10}{'ratio':>9}"
print("\n" + header)
print("-" * len(header))
for row in table:
    print(
        f"{row['tokenizer']:14}{row['ar_fertility']:>10.2f}"
        f"{row['en_fertility']:>10.2f}{row['ratio']:>9.2f}"
    )

# The headline numbers come from the WORST tokenizer for Arabic, because that
# is the one an Arabic reader is most likely to be paying for.
worst = max(table, key=lambda r: r["ratio"])
ar_fertility = worst["ar_fertility"]
en_fertility = worst["en_fertility"]
ratio = worst["ratio"]
Workshop code
tokenizer        ar/word   en/word    ratio
-------------------------------------------
gpt2                5.93      1.32     4.50
bloom               1.70      1.34     1.26
xlm-roberta         1.91      1.51     1.26

One sentence, cut three ways. The � marks a token smaller than a single letter — the tokenizer is not splitting words, it is splitting characters. It is not a broken file or a broken font: the third block decodes the SAME bytes into readable Arabic, which is the proof that the fragmentation is a choice somebody made rather than a property of the script.

env.explain("subword")


# One sentence, cut both ways. The average is an argument; this is the
# evidence for it.
#
# DECODE EACH PIECE, never print `.tokens` directly. GPT-2 is a BYTE-level
# BPE: its raw token strings are the bytes re-encoded as printable Latin-1, so
# an Arabic word shows up as `ĠاÙĦ | ت | ع` — mojibake that looks like a
# bug in the notebook rather than the finding. Decoding each id individually
# puts the actual fragments back on screen, and where a token is half a UTF-8
# character it shows as `�` — which IS the finding: the tokenizer is cutting
# below the level of a letter.
# CHOOSE a demonstrative pair; do not take pairs[0].
#
# This corpus contains rows whose Arabic column is untranslated — UN document
# titles, mostly — and the seeded sample happened to land on one. The result
# was the showcase cell printing the SAME English sentence under both labels,
# 15 tokens against 15, on the one screen the entire argument rests on. The
# averages above were right; the evidence for them was gibberish.
#
# So the pair is picked, not indexed: genuinely Arabic on one side, genuinely
# not on the other, and long enough to show fragmentation. Deterministic,
# because it scans in order and takes the first that qualifies.
def demonstrative(candidates):
    for ar, en in candidates:
        if arabic_share(ar) > 0.8 and arabic_share(en) < 0.2 and 8 <= len(ar.split()) <= 20:
            return ar, en
    return candidates[0]


sample_ar, sample_en = demonstrative(pairs)
worst_name = worst["tokenizer"]
tok = tokenizers[worst_name]


def pieces_of(tokenizer, text):
    """Decoded pieces, plus how many were not even whole characters.

    A byte-level BPE splits UTF-8, and an Arabic letter is two bytes. So a
    single token id can be HALF A LETTER, and decoding it alone yields no
    character at all — U+FFFD, the replacement character.

    That is not a corpus problem and not a rendering problem. It is the
    measurement: the same file decodes perfectly through BLOOM two blocks
    below. Counting the fragments turns the confusing symbol into the number
    it was always standing for.
    """
    ids = tokenizer.encode(text, add_special_tokens=False).ids
    parts, fragments = [], 0
    for i in ids:
        piece = tokenizer.decode([i])
        if not piece or "\ufffd" in piece:
            fragments += 1
            parts.append("\ufffd")
        else:
            parts.append(piece)
    return parts, fragments


def show(tokenizer, label, text, name):
    parts, fragments = pieces_of(tokenizer, text)
    print(f"\n{label} — {len(parts)} tokens, {len(text.split())} words  [{name}]")
    print("  " + " | ".join(parts))
    if fragments:
        share = fragments / len(parts) * 100
        if env.lang == "ar":
            print(
                f"  ← {fragments} من {len(parts)} رمزاً ({share:.0f}%) ليست حروفاً كاملة."
                " رمز \ufffd يعني قطعة أصغر من الحرف الواحد — وهذا هو القياس لا خطأ عرض."
            )
        else:
            print(
                f"  ← {fragments} of {len(parts)} tokens ({share:.0f}%) are not whole"
                " characters. A \ufffd is a piece SMALLER than one letter — that is the"
                " measurement, not a rendering fault."
            )


worst_name = worst["tokenizer"]
tok = tokenizers[worst_name]

show(tok, "العربية" if env.lang == "ar" else "Arabic", sample_ar, worst_name)
show(tok, "الإنجليزية" if env.lang == "ar" else "English", sample_en, worst_name)

# The same Arabic sentence through a tokenizer that bought vocabulary for it.
# This block is what proves the file is fine and the fragmentation is a choice.
best_name = min(table, key=lambda r: r["ratio"])["tokenizer"]
if best_name != worst_name:
    show(
        tokenizers[best_name],
        "العربية" if env.lang == "ar" else "Arabic",
        sample_ar,
        best_name,
    )
Workshop code
subword — A piece smaller than a word. It lets a fixed vocabulary cover words it has never seen, at the cost of splitting them.

Arabic — 71 tokens, 10 words  [gpt2]
  ال | � | � | � | � | � | � | ر | � | � | ه | م | ي | ة | � | � | � | � | ن | � | � | � | � | ي | ن | ة |  ال | � | � | ي | � | � | ة | � | � | ال | � | � | � | � | ر | ة | � | � | ت | م | ت | ل | � | � | � | � | � | � | د | � | � | ن | ة |  ال | ع | و | � | � | م |  ال | س | ا | م | ة | .
  ← 38 of 71 tokens (54%) are not whole characters. A � is a piece SMALLER than one letter — that is the measurement, not a rendering fault.

English — 19 tokens, 14 words  [gpt2]
  More |  importantly | , |  the |  driving |  c | abs |  of |  the |  locom | ot | ives |  would |  fill |  with |  poisonous |  exhaust |  fumes | .

Arabic — 22 tokens, 10 words  [bloom]
  الأ | كثر |  أهمية | ، |  أن |  كاب | ينة |  القيادة |  بالق | اطرة |  ستم | تل | ئ |  بأ | د | خ | نة |  العو | ادم |  الس | امة | .

Notice which tokenizer is expensive. GPT-2's was spent almost entirely on English, so Arabic falls through to its byte-level fallback and is billed one per byte or so. BLOOM spent 250,000 pieces across forty-six languages, and Arabic got some of them — which is why the same sentences cost it barely more than English.

Neither is a bug. They are different purchases, and the difference lands on different readers.

The ratio, converted into the two things it actually costs: money and room to think.

env.explain("context window")

price = env.cfg["pricePerMillionTokens"]
window = env.cfg["contextWindow"]

# What the ratio means once it leaves the spreadsheet.
ar_cost_multiple = round(ratio, 3)
effective_context_ar = int(window / ar_fertility)
effective_context_en = int(window / en_fertility)

ar_million_words_cost = (ar_fertility * 1_000_000 / 1_000_000) * price
en_million_words_cost = (en_fertility * 1_000_000 / 1_000_000) * price

if env.lang == "ar":
    print(
        f"لكل مليون كلمة: {ar_million_words_cost:.2f}$ بالعربية · {en_million_words_cost:.2f}$ بالإنجليزية"
    )
    print(f"القارئ العربي يدفع {ar_cost_multiple:.2f}× للمحتوى ذاته")
    print(
        f"نافذة {window:,} رمز تسع {effective_context_ar:,} كلمة عربية · {effective_context_en:,} كلمة إنجليزية"
    )
else:
    print(
        f"per million words: ${ar_million_words_cost:.2f} Arabic · ${en_million_words_cost:.2f} English"
    )
    print(f"an Arabic reader pays {ar_cost_multiple:.2f}× for the same content")
    print(
        f"a {window:,}-token window holds {effective_context_ar:,} Arabic words · {effective_context_en:,} English"
    )
Workshop code
context window — How many tokens a model can hold at once. Measured in tokens, not words — so a language with higher fertility gets less of it.
per million words: $17.80 Arabic · $3.95 English
an Arabic reader pays 4.50× for the same content
a 8,192-token window holds 1,380 Arabic words · 6,220 English

Exercise

Find the single cheapest change that moves the number. Try stripping diacritics, normalizing alef and yaa variants, or removing tatweel — each is one line, and each changes how the vocabulary matches.

Then ask the harder question: which of these changes is safe? Stripping diacritics is nearly free on modern prose and destructive on Qur'anic or poetic text. A tokenizer optimization that quietly corrupts your corpus is not an optimization. Report both the tokens saved and what you gave up.

# YOUR TURN.
#
# Each of these is one line, and each changes how the vocabulary matches.
# Turn them on, re-measure, and report BOTH what you saved and what you lost.
import re

STRIP_DIACRITICS = False  # harmless on modern prose, destructive on Qur'anic text
NORMALIZE_ALEF = False  # أ إ آ -> ا ; ى -> ي
STRIP_TATWEEL = False  # the ـ elongation character


def normalize(text):
    if STRIP_DIACRITICS:
        text = re.sub(r"[\u064B-\u0652\u0670]", "", text)
    if NORMALIZE_ALEF:
        text = re.sub(r"[أإآ]", "ا", text).replace("ى", "ي")
    if STRIP_TATWEEL:
        text = text.replace("\u0640", "")
    return text


enabled = [
    name
    for name, on in (
        ("diacritics", STRIP_DIACRITICS),
        ("alef", NORMALIZE_ALEF),
        ("tatweel", STRIP_TATWEEL),
    )
    if on
]

if not enabled:
    # Saying so beats printing "6.018 -> 6.018 (-0.0%)", which reads like the
    # normalization does not work rather than like it was never switched on.
    if env.lang == "ar":
        print("لم تُفعَّل أي معالجة بعد — أعد أحد الثوابت أعلاه إلى True ثم شغّل الخلية.")
    else:
        print("No normalization enabled yet — set one of the flags above to True and re-run.")
else:
    # Measured on the WORST tokenizer, the one the headline number came from.
    # Normalizing against a tokenizer that already handles Arabic well would
    # show almost no gain and teach the opposite of the truth.
    normalized = [normalize(t) for t in ar_texts]
    after = count_tokens(tok, normalized) / ar_words
    saved = (ar_fertility - after) / ar_fertility * 100 if ar_fertility else 0.0
    print(f"{worst_name}: {ar_fertility:.3f} -> {after:.3f} tokens/word  ({-saved:+.1f}%)")
    print(f"  enabled: {', '.join(enabled)}")
Workshop code

A hint is available in the notebook — env.hint(1)

Both checks: the ratio you measured, and that you measured it on more than one tokenizer.

fertility_ratio_final = ratio
ratio_ok = env.check("fertility-measured", fertility_ratio_final)
count_ok = env.check("tokenizers-compared", n_tokenizers)
Workshop code
✓ Arabic fertility, as a multiple of English: 4.505 (needs ≥ 2.0)
✓ Tokenizers measured: 3 (needs ≥ 3)
receipt = env.receipt()
Workshop code
Workshop complete.

Completion code: AZ-██████████
Paste it on the workshop's page on Azimuth to record it.
Last verified: 2026-08-27 · NVIDIA GeForce GTX 1070 with Max-Q Design · PyTorch 2.6.0+cu124 · Python 3.13.5 · cd18d2f

Terms in this workshop