اللغةbeginnerمعالج رسوميات اختياري~25 دقيقةColab
المُجزِّئات والضريبة العربية
Tokenizers and the Arabic tax
الخصوبة ثمنٌ تدفعه رمزاً رمزاً
خذ جملة واحدة، واكتبها بالعربية ثم بالإنجليزية، وأرسل النسختين إلى نموذج لغوي. ستجد أن النسخة العربية تستهلك أكثر. ولا أقصد هذا مجازاً: الفرق يُقاس بعدد الرموز. والرموز هي وحدة الحساب كلّها، فبها يُسعَّر الاستخدام، وبها تُقاس . لذلك يدفع القارئ العربي مرّتين: فاتورة أعلى على المحتوى نفسه، ومساحة سياق فعلية أضيق وهو يستخدم النموذج نفسه.
لكن هذه الورشة لا تكتفي بإثبات أن الفجوة موجودة؛ إنها تقيس كم منها يعود إلى قرار اتّخذه أحدٌ ما. حين تشغّل الكود سترى أن الجمل العربية نفسها تكلّف أحد المستخدمة فعلاً في نماذج حقيقية نحو أربعة أضعاف ما تكلّفه في مُجزِّئ آخر. وعلى أسوأ المُجزِّئات الثلاثة، تبلغ نسبة رموز الكلمة العربية إلى رموز الكلمة الإنجليزية 4.505.
هذا التفاوت بين المُجزِّئات هو بيت القصيد. فالضريبة ليست صفةً في اللغة العربية نفسها. إنها قرار، واتّخذه أحدٌ ما.
الهدف
قِس <Term en="Fertility">الخصوبة</Term>، أي عدد الرموز لكل كلمة، على <Term en="Parallel Corpus">مدوّنة متوازية</Term>: الجمل نفسها بالعربية والإنجليزية، عبر ثلاثة مُجزِّئات مستخدمة في نماذج حقيقية. ثم حوِّل النسبة بين اللغتين إلى ما تكلّفه فعلاً: من نافذة السياق، ومن <Term en="Inference Cost">تكلفة الاستدلال</Term>. وأخيراً ابحث عن التعديل الواحد على النص الذي يحرّك هذا الرقم أكثر من غيره.
يفتح Colab نسخة للقراءة فقط. احفظ نسخة في Drive للاحتفاظ بتعديلاتك.
يحتاج الدفتر إلى لوحة مفاتيح — يُفضَّل فتحه على حاسوب مكتبي.
الأوراق وراء هذه الورشة
كل ما يحاسبك عليه النموذج اللغوي يُعدّ بالرموز: نافذة السياق تُقاس بالرموز، والسعر يُحدَّد لكل مليون رمز، وحتى حدود الاستخدام تُطبَّق بالرموز. لذلك حين تسأل «كم رمزاً تكلّفني هذه الجملة؟» فأنت لا تسأل سؤالاً أكاديمياً. أنت تقرأ الفاتورة. الذي يحدّد هذا الرقم هو المُجزِّئ، ويحدّده بناءً على النصوص التي تدرَّب عليها. ما تكرّر كثيراً في بيانات أصبح قطعة واحدة كاملة، وما كان نادراً تفتّت إلى شظايا.
تهيئة أوّلية ثم تحميل المدوّنة المتوازية — مع التحقّق منها قبل بدء أي قياس.
import azimuth_nb as azimuth
env = azimuth.setup(SLUG, lang=LANG, profile=PROFILE)المُجزِّئات والضريبة العربية
NVIDIA GeForce GTX 1070 with Max-Q Design · 8.0 غ.ب · ذاكرة 15.9 غ.ب · PyTorch 2.6.0+cu124
الملف: free
البيانات:
· parallel.tsv — already present
جاهز · sampleSentences=1000, pricePerMillionTokens=3.0, contextWindow=8192, seed=17نبدأ بعدّ الكلمات. نسبة، ولا قيمة للبسط ما لم نحدّد المقام أولاً.
import random
ARABIC_RANGE = ("\u0600", "\u06ff")
def arabic_share(text):
"""Fraction of letters that are Arabic script."""
letters = [ch for ch in text if ch.isalpha()]
if not letters:
return 0.0
return sum(ARABIC_RANGE[0] <= ch <= ARABIC_RANGE[1] for ch in letters) / len(letters)
rows = []
with open(env.assets["parallel.tsv"], encoding="utf-8") as fh:
for line in fh:
parts = line.rstrip("\n").split("\t")
if len(parts) >= 2 and parts[0].strip() and parts[1].strip():
rows.append((parts[0].strip(), parts[1].strip()))
# WHICH COLUMN IS ARABIC IS DETECTED, NOT ASSUMED.
#
# An earlier run of this workshop measured a corpus whose columns were the
# other way round and reported that Arabic was CHEAPER than English — a
# confident, precise, exactly-backwards result. Nothing failed: both columns
# are text, both tokenize, and the arithmetic is identical. Only the sign of
# the conclusion changed.
#
# Reading the script itself costs one pass and removes the assumption. A
# column order is a property of whoever exported the file; the alphabet is a
# property of the language.
sample = rows[: min(200, len(rows))]
share_0 = sum(arabic_share(a) for a, _ in sample) / len(sample)
share_1 = sum(arabic_share(b) for _, b in sample) / len(sample)
if max(share_0, share_1) < 0.5:
raise SystemExit(
"Neither column looks like Arabic script — check the encoding of "
"parallel.tsv before trusting anything below."
)
if share_0 >= share_1:
pairs_raw = rows
else:
pairs_raw = [(b, a) for a, b in rows]
print("note: columns were (english, arabic) — swapped to match")
# A fixed seed so the sample — and therefore every number below — is the same
# on your machine, on CI, and on the page.
random.seed(env.cfg["seed"])
limit = env.cfg["sampleSentences"]
pairs = random.sample(pairs_raw, min(limit, len(pairs_raw))) if limit else pairs_raw
# WHITESPACE, deliberately. "Word" has to mean the same operation in both
# languages or the ratio compares nothing. Whitespace undercounts Arabic
# morphology — one written word often carries article, stem and pronoun — so
# this makes the result a LOWER BOUND on the real gap rather than a flattering
# one. Overstating the case would be the easiest way to lose the argument.
ar_words = sum(len(ar.split()) for ar, _ in pairs)
en_words = sum(len(en.split()) for _, en in pairs)
if env.lang == "ar":
print(f"أزواج: {len(pairs):,} من أصل {len(pairs_raw):,}")
print(f"كلمات: {ar_words:,} عربية · {en_words:,} إنجليزية")
else:
print(f"pairs: {len(pairs):,} of {len(pairs_raw):,}")
print(f"words: {ar_words:,} Arabic · {en_words:,} English")note: columns were (english, arabic) — swapped to match
أزواج: 1,000 من أصل 24,999
كلمات: 8,676 عربية · 9,995 إنجليزيةثلاثة مُجزِّئات، ولكلٍّ منها مفرداته ورهانه الخاص على اللغات التي تستحق نصيباً منها.
from tokenizers import Tokenizer
# Three different bets about which languages deserve vocabulary budget.
# Loaded from the Hub by name — small JSON files, no model weights.
SPECS = [
("gpt2", "openai-community/gpt2", "50k pieces, almost all English"),
("bloom", "bigscience/bloom", "250k pieces shared across 46 languages"),
("xlm-roberta", "FacebookAI/xlm-roberta-base", "250k pieces, 100 languages"),
]
tokenizers = {}
for name, repo, note in SPECS:
try:
tokenizers[name] = Tokenizer.from_pretrained(repo)
print(f" · {name:12} {note}")
except Exception as exc:
# A tokenizer that will not load is reported and skipped, not fatal:
# the comparison still means something with two, and a dead Hub should
# not cost you the whole workshop.
print(f" ! {name:12} unavailable ({type(exc).__name__}) — skipped")
n_tokenizers = len(tokenizers) · gpt2 50k pieces, almost all English
· bloom 250k pieces shared across 46 languages
· xlm-roberta 250k pieces, 100 languagesلا تقرأ صفّاً واحداً؛ اقرأ عمود النسبة من أعلى إلى أسفل. الاكتشاف هنا هو الفرق بين المُجزِّئات: الجمل العربية نفسها تكلّف أحدها أربعة أضعاف ما تكلّفه في آخر.
def count_tokens(tok, texts):
"""Total pieces across a list of strings, no special tokens."""
return sum(len(tok.encode(t, add_special_tokens=False).ids) for t in texts)
ar_texts = [ar for ar, _ in pairs]
en_texts = [en for _, en in pairs]
table = []
for name, tok in tokenizers.items():
ar_tokens = count_tokens(tok, ar_texts)
en_tokens = count_tokens(tok, en_texts)
ar_f = ar_tokens / ar_words
en_f = en_tokens / en_words
table.append(
{
"tokenizer": name,
"ar_fertility": round(ar_f, 3),
"en_fertility": round(en_f, 3),
"ratio": round(ar_f / en_f, 3),
}
)
header = f"{'tokenizer':14}{'ar/word':>10}{'en/word':>10}{'ratio':>9}"
print("\n" + header)
print("-" * len(header))
for row in table:
print(
f"{row['tokenizer']:14}{row['ar_fertility']:>10.2f}"
f"{row['en_fertility']:>10.2f}{row['ratio']:>9.2f}"
)
# The headline numbers come from the WORST tokenizer for Arabic, because that
# is the one an Arabic reader is most likely to be paying for.
worst = max(table, key=lambda r: r["ratio"])
ar_fertility = worst["ar_fertility"]
en_fertility = worst["en_fertility"]
ratio = worst["ratio"]tokenizer ar/word en/word ratio
-------------------------------------------
gpt2 5.93 1.32 4.50
bloom 1.70 1.34 1.26
xlm-roberta 1.91 1.51 1.26الجملة نفسها مقطّعة بثلاث طرق. حين ترى � فمعناه أن الرمز أصغر من حرف واحد: المُجزِّئ هنا لا يقطّع الكلمات، بل يشطر الحرف نفسه. وهذا ليس عطلاً في الملف ولا في الخط. الكتلة الثالثة تفكّ البايتات نفسها فتظهر عربيةً مقروءة. وهذا هو الدليل: التفتيت قرارٌ اتّخذه أحدٌ ما، لا صفةٌ في الكتابة العربية.
env.explain("subword")
# One sentence, cut both ways. The average is an argument; this is the
# evidence for it.
#
# DECODE EACH PIECE, never print `.tokens` directly. GPT-2 is a BYTE-level
# BPE: its raw token strings are the bytes re-encoded as printable Latin-1, so
# an Arabic word shows up as `ĠاÙĦ | ت | ع` — mojibake that looks like a
# bug in the notebook rather than the finding. Decoding each id individually
# puts the actual fragments back on screen, and where a token is half a UTF-8
# character it shows as `�` — which IS the finding: the tokenizer is cutting
# below the level of a letter.
# CHOOSE a demonstrative pair; do not take pairs[0].
#
# This corpus contains rows whose Arabic column is untranslated — UN document
# titles, mostly — and the seeded sample happened to land on one. The result
# was the showcase cell printing the SAME English sentence under both labels,
# 15 tokens against 15, on the one screen the entire argument rests on. The
# averages above were right; the evidence for them was gibberish.
#
# So the pair is picked, not indexed: genuinely Arabic on one side, genuinely
# not on the other, and long enough to show fragmentation. Deterministic,
# because it scans in order and takes the first that qualifies.
def demonstrative(candidates):
for ar, en in candidates:
if arabic_share(ar) > 0.8 and arabic_share(en) < 0.2 and 8 <= len(ar.split()) <= 20:
return ar, en
return candidates[0]
sample_ar, sample_en = demonstrative(pairs)
worst_name = worst["tokenizer"]
tok = tokenizers[worst_name]
def pieces_of(tokenizer, text):
"""Decoded pieces, plus how many were not even whole characters.
A byte-level BPE splits UTF-8, and an Arabic letter is two bytes. So a
single token id can be HALF A LETTER, and decoding it alone yields no
character at all — U+FFFD, the replacement character.
That is not a corpus problem and not a rendering problem. It is the
measurement: the same file decodes perfectly through BLOOM two blocks
below. Counting the fragments turns the confusing symbol into the number
it was always standing for.
"""
ids = tokenizer.encode(text, add_special_tokens=False).ids
parts, fragments = [], 0
for i in ids:
piece = tokenizer.decode([i])
if not piece or "\ufffd" in piece:
fragments += 1
parts.append("\ufffd")
else:
parts.append(piece)
return parts, fragments
def show(tokenizer, label, text, name):
parts, fragments = pieces_of(tokenizer, text)
print(f"\n{label} — {len(parts)} tokens, {len(text.split())} words [{name}]")
print(" " + " | ".join(parts))
if fragments:
share = fragments / len(parts) * 100
if env.lang == "ar":
print(
f" ← {fragments} من {len(parts)} رمزاً ({share:.0f}%) ليست حروفاً كاملة."
" رمز \ufffd يعني قطعة أصغر من الحرف الواحد — وهذا هو القياس لا خطأ عرض."
)
else:
print(
f" ← {fragments} of {len(parts)} tokens ({share:.0f}%) are not whole"
" characters. A \ufffd is a piece SMALLER than one letter — that is the"
" measurement, not a rendering fault."
)
worst_name = worst["tokenizer"]
tok = tokenizers[worst_name]
show(tok, "العربية" if env.lang == "ar" else "Arabic", sample_ar, worst_name)
show(tok, "الإنجليزية" if env.lang == "ar" else "English", sample_en, worst_name)
# The same Arabic sentence through a tokenizer that bought vocabulary for it.
# This block is what proves the file is fine and the fragmentation is a choice.
best_name = min(table, key=lambda r: r["ratio"])["tokenizer"]
if best_name != worst_name:
show(
tokenizers[best_name],
"العربية" if env.lang == "ar" else "Arabic",
sample_ar,
best_name,
)subword — قطعة أصغر من الكلمة. تتيح لمفردات ثابتة أن تغطي كلمات لم ترها قط، بثمن تفتيتها.
العربية — 71 tokens, 10 words [gpt2]
ال | � | � | � | � | � | � | ر | � | � | ه | م | ي | ة | � | � | � | � | ن | � | � | � | � | ي | ن | ة | ال | � | � | ي | � | � | ة | � | � | ال | � | � | � | � | ر | ة | � | � | ت | م | ت | ل | � | � | � | � | � | � | د | � | � | ن | ة | ال | ع | و | � | � | م | ال | س | ا | م | ة | .
← 38 من 71 رمزاً (54%) ليست حروفاً كاملة. رمز � يعني قطعة أصغر من الحرف الواحد — وهذا هو القياس لا خطأ عرض.
الإنجليزية — 19 tokens, 14 words [gpt2]
More | importantly | , | the | driving | c | abs | of | the | locom | ot | ives | would | fill | with | poisonous | exhaust | fumes | .
العربية — 22 tokens, 10 words [bloom]
الأ | كثر | أهمية | ، | أن | كاب | ينة | القيادة | بالق | اطرة | ستم | تل | ئ | بأ | د | خ | نة | العو | ادم | الس | امة | .انظر أيّ المُجزِّئات هو المُكلِف. أنفق GPT-2 كلّها تقريباً على الإنجليزية، فلم يبقَ فيها للعربية شيء يُذكر. لذلك حين يصله نصّ عربي يرجع إلى البايتات الخام، ويحاسبك بنحو رمز لكل بايت. أما BLOOM فوزّع مئتين وخمسين ألف قطعة على ستّ وأربعين لغة، ونالت العربية منها حصّة حقيقية. ولهذا لا تكلّفه الجمل العربية إلا أكثر قليلاً من نظيرتها الإنجليزية.
لا خطأ في أيّ منهما. إنهما صفقتان مختلفتان، والثمن يدفعه قرّاء مختلفون.
هنا تتحوّل النسبة إلى ما تكلّفه فعلاً: مالٌ تدفعه، ومساحةٌ من نافذة السياق تخسرها.
env.explain("context window")
price = env.cfg["pricePerMillionTokens"]
window = env.cfg["contextWindow"]
# What the ratio means once it leaves the spreadsheet.
ar_cost_multiple = round(ratio, 3)
effective_context_ar = int(window / ar_fertility)
effective_context_en = int(window / en_fertility)
ar_million_words_cost = (ar_fertility * 1_000_000 / 1_000_000) * price
en_million_words_cost = (en_fertility * 1_000_000 / 1_000_000) * price
if env.lang == "ar":
print(
f"لكل مليون كلمة: {ar_million_words_cost:.2f}$ بالعربية · {en_million_words_cost:.2f}$ بالإنجليزية"
)
print(f"القارئ العربي يدفع {ar_cost_multiple:.2f}× للمحتوى ذاته")
print(
f"نافذة {window:,} رمز تسع {effective_context_ar:,} كلمة عربية · {effective_context_en:,} كلمة إنجليزية"
)
else:
print(
f"per million words: ${ar_million_words_cost:.2f} Arabic · ${en_million_words_cost:.2f} English"
)
print(f"an Arabic reader pays {ar_cost_multiple:.2f}× for the same content")
print(
f"a {window:,}-token window holds {effective_context_ar:,} Arabic words · {effective_context_en:,} English"
)context window — عدد الرموز التي يسع النموذج حملها دفعة واحدة. تُقاس بالرموز لا بالكلمات — فاللغة الأعلى خصوبة تنال منها أقل.
لكل مليون كلمة: 17.80$ بالعربية · 3.95$ بالإنجليزية
القارئ العربي يدفع 4.50× للمحتوى ذاته
نافذة 8,192 رمز تسع 1,380 كلمة عربية · 6,220 كلمة إنجليزيةتمرين
ابحث عن أرخص تعديل يحرّك الرقم. جرّب حذف التشكيل، أو توحيد أشكال الألف والياء، أو إزالة التطويل (الكشيدة). كلٌّ منها سطر واحد، وكلٌّ منها يغيّر القطع التي يجدها المُجزِّئ مطابِقةً في مفرداته.
ثم اسأل السؤال الأصعب: أيّ هذه التعديلات آمن؟ حذف التشكيل لا يكاد يغيّر شيئاً في النثر الحديث، لكنه يُتلف النص القرآني والشعر. وأيّ تحسين في يُفسد مدوّنتك بصمت ليس تحسيناً أصلاً. لذلك اذكر في تقريرك الأمرين معاً: كم رمزاً وفّرت، وماذا خسرت في المقابل.
# YOUR TURN.
#
# Each of these is one line, and each changes how the vocabulary matches.
# Turn them on, re-measure, and report BOTH what you saved and what you lost.
import re
STRIP_DIACRITICS = False # harmless on modern prose, destructive on Qur'anic text
NORMALIZE_ALEF = False # أ إ آ -> ا ; ى -> ي
STRIP_TATWEEL = False # the ـ elongation character
def normalize(text):
if STRIP_DIACRITICS:
text = re.sub(r"[\u064B-\u0652\u0670]", "", text)
if NORMALIZE_ALEF:
text = re.sub(r"[أإآ]", "ا", text).replace("ى", "ي")
if STRIP_TATWEEL:
text = text.replace("\u0640", "")
return text
enabled = [
name
for name, on in (
("diacritics", STRIP_DIACRITICS),
("alef", NORMALIZE_ALEF),
("tatweel", STRIP_TATWEEL),
)
if on
]
if not enabled:
# Saying so beats printing "6.018 -> 6.018 (-0.0%)", which reads like the
# normalization does not work rather than like it was never switched on.
if env.lang == "ar":
print("لم تُفعَّل أي معالجة بعد — أعد أحد الثوابت أعلاه إلى True ثم شغّل الخلية.")
else:
print("No normalization enabled yet — set one of the flags above to True and re-run.")
else:
# Measured on the WORST tokenizer, the one the headline number came from.
# Normalizing against a tokenizer that already handles Arabic well would
# show almost no gain and teach the opposite of the truth.
normalized = [normalize(t) for t in ar_texts]
after = count_tokens(tok, normalized) / ar_words
saved = (ar_fertility - after) / ar_fertility * 100 if ar_fertility else 0.0
print(f"{worst_name}: {ar_fertility:.3f} -> {after:.3f} tokens/word ({-saved:+.1f}%)")
print(f" enabled: {', '.join(enabled)}")يتوفّر تلميح في الدفتر — env.hint(1)
فحصان: الأول يتأكّد من النسبة التي قِستَها، والثاني من أنك قِستَها على أكثر من مُجزِّئ واحد.
fertility_ratio_final = ratio
ratio_ok = env.check("fertility-measured", fertility_ratio_final)
count_ok = env.check("tokenizers-compared", n_tokenizers)✓ خصوبة العربية منسوبة إلى الإنجليزية: 4.505 (المطلوب ≥ 2.0)
✓ عدد المُجزِّئات المقيسة: 3 (المطلوب ≥ 3)receipt = env.receipt()اكتملت الورشة.
رمز الإتمام: AZ-██████████
الصقه في صفحة الورشة على أزيموث لتسجيل إتمامها.مصطلحات هذه الورشة
- تجزئة النصوصTokenization
- جزء الكلمةSubword
- قاموس المفردات للنموذجVocabulary
- نافذة السياقContext Window