Language Models2022intermediate12 min read

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

BLOOM: نموذج لغوي متعدد اللغات مفتوح الوصول بـ 176 مليار معامل

BigScience Workshop · Le Scao, T. · Fan, A. · Akiki, C. · Pavlick, E. · Ilić, S. · Hesslow, D. · Castagné, R. · Luccioni, A. S. — arXiv

The problem

By 2022, large language models like GPT-3, PaLM, and Chinchilla had demonstrated remarkable capabilities — learning, code generation, reasoning — but they were all developed behind closed doors by well-funded organizations and never released to the public. The research community could study their outputs but not their weights, data, or engineering decisions. This created a two-tier system: a handful of companies controlled the most powerful models, while everyone else could only guess how they worked. Moreover, most LLMs were overwhelmingly English-centric, leaving billions of speakers of other languages underserved.

The contribution

BLOOM: a 176B-parameter decoder-only , trained openly by a collaboration of over 1,000 researchers from 70+ countries. It was trained on the ROOTS corpus — 1.6 TB of curated text across 46 natural languages and 13 programming languages — using the Jean Zay supercomputer in Paris for 117 days on 384 A100 GPUs. The architecture introduces positional embeddings and an LayerNorm for training stability. After (BLOOMZ), the model achieves competitive performance on benchmarks. Everything — weights, code, data documentation, and training logs — was released under the Responsible AI License (RAIL).

The impact

BLOOM proved that the open-science model can produce frontier-scale LLMs. It was the first 100B+ model with fully open weights, data documentation, and training code. It inspired a wave of open models — Falcon, LLaMA, Mistral — and established the template for large-scale collaborative AI research. Its Responsible AI License became a reference for balancing openness with safety. The ROOTS corpus set a standard for multilingual data curation and documentation.

Imagine the most powerful telescope in the world — but only one company owns it, it only points at the English-speaking sky, and nobody is allowed to look through the lens. That was the state of large language models before BLOOM.

The BigScience project said: what if we built our own telescope, together? Over a thousand researchers from 70+ countries pooled their expertise, gathered starlight from 46 languages, and built a 176-billion-parameter lens that anyone can point wherever they want. The blueprints, the glass, the mirror — everything is open.

The problem: powerful models behind locked doors

When GPT-3 appeared in 2020, it stunned the world with its ability to write essays, translate languages, and solve problems from just a few examples. But OpenAI never released its weights. Google followed with PaLM and Chinchilla — equally impressive, equally closed. The pattern was clear: only organizations with billions of dollars and thousands of GPUs could build frontier models, and they had no obligation to share.

This created a deep asymmetry. Researchers at universities could read the papers, but not reproduce the results. Developers in the Global South had to rely on API access controlled by companies in Silicon Valley. Languages other than English were afterthoughts — GPT-3 was trained almost entirely on English text. The most powerful technology of the decade was being shaped by a narrow group, for a narrow audience.

BigScience set out to challenge this. Organized by Hugging Face with support from the French government, it gathered over 1,000 researchers from more than 70 countries into working groups spanning engineering, data governance, ethics, and evaluation. The goal: build a large language model in the open, document every decision, and release everything.

Open in Lab
Explore the scale of the BigScience collaboration: over 1,000 researchers from 70+ countries organized into specialized working groups.
The demo wakes as you arrive…

The data: ROOTS — a multilingual garden

A language model is only as good as the text it trains on. Most LLMs use crawled web data with minimal curation — Common Crawl, the Pile — and the result is models that reflect the biases and language distribution of the English-language internet.

BLOOM was trained on the ROOTS corpus: 1.6 terabytes of text spanning 498 sources across 46 natural languages and 13 programming languages. But ROOTS was not just big — it was deliberately curated. A dedicated data governance working group sourced data from books, scientific papers, government documents, Wikipedia, and community-gathered datasets. Each source was documented with a data card describing its provenance, license, language, and potential biases.

Crucially, the team made conscious choices about language representation. Rather than letting data availability dictate proportions (which would make the model 90%+ English), they actively sought high-quality sources for underrepresented languages including several African and Southeast Asian languages. The corpus was deduplicated, personally identifiable information was redacted, and a quality filtering pipeline removed low-quality text.

Open in Lab
Explore the language distribution of the ROOTS corpus. Toggle between language families to see how data was sourced across 46 languages.
The demo wakes as you arrive…

The tokenizer: one vocabulary for all languages

The — the component that splits text into tokens the model processes — is often treated as a detail. But for a multilingual model, it is a make-or-break design decision. A tokenizer trained mostly on English will split Chinese or Arabic words into many tiny pieces, wasting the model's context window and degrading performance.

BLOOM uses Byte-Pair Encoding at the byte level with a vocabulary of 250,680 tokens. Three design choices were critical. First, no Unicode normalization was applied — this preserves scripts like Arabic diacritics and CJK characters faithfully. Second, the vocabulary size was set large enough to keep fertility (the average number of tokens per word) close to monolingual tokenizers for most languages. Third, the vocabulary size was constrained to be divisible by 128 (for GPU efficiency) and by 4 (for ).

The result: a single tokenizer that handles 46 natural languages and 13 programming languages without destroying any of them. Text in Arabic, French, Vietnamese, and Python all tokenize efficiently through the same pipeline.

Open in Lab
Type or select text in different languages and see how BLOOM's tokenizer splits it compared to a typical English-centric tokenizer.
The demo wakes as you arrive…

Architecture: a decoder-only Transformer with two twists

BLOOM follows the decoder-only Transformer architecture — the same family as GPT-3. The model has 70 layers, 112 heads, a hidden dimension of 14,336, and processes sequences of 2,048 tokens. The total parameter count is 176 billion, of which 3.6 billion are in the embedding layer alone. It uses activations and tied input-output embeddings.

But the BigScience team did not simply copy GPT-3. Through a systematic at the 1-billion-parameter scale, they tested architectural variations and adopted two key deviations from the standard Transformer.

ALiBi Positional Embeddings. Standard Transformers add positional information to embeddings — either through learned vectors or sinusoidal functions. ALiBi (Attention with Linear Biases) takes a completely different approach: instead of modifying the embeddings, it adds a linear penalty directly to the attention scores. The further apart two tokens are, the larger the penalty — like a rubber band that makes distant tokens harder to attend to. Each attention head gets a different slope for this penalty, so some heads focus locally while others reach further.

The team found that ALiBi produced smoother training dynamics and better downstream performance than both learned embeddings and rotary embeddings (RoPE), even without its original motivation of sequence length extrapolation.

softmax ⁣(qi⋅kj⊤d−m⋅∣i−j∣)\text{softmax}\!\Bigl(\frac{q_i \cdot k_j^\top}{\sqrt{d}} - m \cdot |i - j|\Bigr)
ALiBi attention — a distance-dependent penalty replaces positional embeddings — Instead of adding position information to the input embeddings, ALiBi subtracts m⋅∣i−j∣m \cdot |i - j| from each attention score, where mm is a head-specific slope and ∣i−j∣|i - j| is the distance between positions. Heads with small slopes see the full context; heads with large slopes focus on nearby tokens.
Open in Lab
Visualize how ALiBi penalizes attention scores based on distance. Compare different head slopes and see how each head specializes.
The demo wakes as you arrive…

Embedding LayerNorm. The second deviation is an additional immediately after the embedding layer — before the first Transformer block. In preliminary experiments at 104B parameters (using float16), this extra normalization dramatically improved training stability, preventing the loss spikes and divergence that plague large-scale training runs.

There is a trade-off: embedding LayerNorm slightly penalizes generalization. The team accepted this cost because training stability at 176B scale was non-negotiable — a single divergence could waste weeks of supercomputer time. Later research suggested that switching from float16 to bfloat16 (which BLOOM's final training used) may reduce the need for this extra normalization.

Open in Lab
Explore BLOOM's architecture layer by layer. Click on each component to see its purpose, dimensions, and design rationale.
The demo wakes as you arrive…

Engineering: training 176B parameters across 384 GPUs

Training a 176B-parameter model does not fit on a single GPU — or even a single machine. BLOOM was trained on the Jean Zay supercomputer in Paris using 384 NVIDIA A100 80GB GPUs organized into 48 nodes of 8 GPUs each. The training ran for 117 days, from March 11 to July 6, 2022, processing approximately 366 billion tokens.

To distribute the work across GPUs, the team combined three parallelism strategies, often called "," building on NVIDIA's Megatron-LM framework with DeepSpeed integration.

Open in Lab
See how Data, Tensor, and Pipeline Parallelism split BLOOM's training across 384 GPUs. Toggle each strategy to understand what it distributes.
The demo wakes as you arrive…

replicates the entire model on multiple GPU groups. Each group processes a different mini-batch, then gradients are averaged via all-reduce. Think of it as photocopying the model and giving each copy a different homework assignment, then combining the answers.

Tensor Parallelism splits individual layers across GPUs within a single node. A 14,336-dimension matrix multiplication can be distributed across 4 GPUs, each handling a slice. This requires fast interconnects (NVLink) because GPUs must exchange intermediate results.

assigns different layers to different GPU groups. GPU group 1 handles layers 1–18, group 2 handles layers 19–35, and so on. The data flows through GPUs sequentially, like an assembly line. The challenge is keeping all GPUs busy — "pipeline bubbles" occur when some stages wait for others.

BLOOM used Data Parallelism with 8 copies, Tensor Parallelism of degree 4 within each node, and Pipeline Parallelism of degree 12 across nodes. These fused CUDA kernels from Megatron-LM — fused LayerNorm, fused GeLU with bias — squeezed maximum throughput from the hardware.

Evaluation: from zero-shot to multitask finetuning

The pretrained BLOOM model performs competitively with GPT-3 on English benchmarks and shows strong multilingual capabilities. On SuperGLUE in zero-shot and few-shot settings, BLOOM's performance is similar to comparably-sized GPT models trained on the Pile.

But the real story is what happens after multitask prompted finetuning. The team created BLOOMZ by finetuning BLOOM on xP3 — a multilingual collection of prompted tasks spanning 46 languages. The result was dramatic: BLOOMZ showed substantially stronger zero-shot task generalization across languages, closing the gap with models that had been finetuned on much more English-centric data.

This demonstrated a crucial insight: a multilingual model, when properly finetuned with multilingual prompts, can generalize to new tasks in languages it saw during finetuning — and sometimes even in languages it did not see at all, through cross-lingual transfer.

Open in Lab
Compare BLOOM and BLOOMZ performance on multilingual benchmarks. Toggle between zero-shot and few-shot settings.
The demo wakes as you arrive…

Licensing: open but responsible

BLOOM was not released under a standard open-source license. Instead, the team created the Responsible AI License (RAIL) — a novel approach that grants broad access and modification rights while restricting specific harmful uses. Users can download, modify, and deploy BLOOM freely, but the license prohibits using it to generate disinformation, conduct surveillance, or cause deliberate harm.

This was a deliberate middle path between two extremes: fully closed models (GPT-3, PaLM) that restrict access entirely, and fully permissive licenses (MIT, Apache) that impose no use restrictions. The RAIL approach acknowledges that a 176B-parameter model capable of generating fluent text in 46 languages is powerful enough to warrant some guardrails, even when the goal is maximum openness.

Legacy and what came next

  1. 2020

    GPT-3 (175B) — closed

    OpenAI trains a 175B model with impressive few-shot abilities but never releases the weights. API access only.

  2. 2022

    OPT (175B) — partially open

    Meta releases OPT with weights available for research. A step toward openness but with restricted commercial use.

  3. 2022

    BLOOM (176B) — fully open

    BigScience releases BLOOM with open weights, data documentation, training code, and the RAIL license. The first truly open 100B+ model.

  4. 2023

    LLaMA (65B) — research release

    Meta releases LLaMA, smaller but more compute-efficient, showing that open models can match much larger closed ones.

  5. 2023

    Falcon (180B) — open compute-optimal

    TII builds on the open-model movement BLOOM started, releasing Falcon with state-of-the-art performance among open models.

BLOOM's most important contribution was not its scores — it was the proof of concept. Before BLOOM, many assumed that building a frontier LLM required the resources of a big tech company. BLOOM showed that an international collaboration, organized openly and transparently, could match the scale of corporate efforts. It opened the door for every open model that followed.

The model also established norms that the field now takes for granted: publishing training details, documenting data sources, measuring carbon footprints, and thinking carefully about licensing. These practices were pioneering when BLOOM did them. They are now expected.

Loading and using BLOOM with Hugging Facepython

Simplified to show the idea — not the real implementation.

from transformers import AutoModelForCausalLM, AutoTokenizer
# Load the BLOOM tokenizer (250,680 vocab, byte-level BPE) tokenizer = AutoTokenizer.from_pretrained("bigscience/bloom")
# Load model — use device_map="auto" for multi-GPU inference model = AutoModelForCausalLM.from_pretrained(
    "bigscience/bloom",
    device_map="auto",       # Distributes layers across GPUs
    torch_dtype="auto"       # Uses bfloat16 if available
)
# Generate text — works in 46 languages prompt = "Translate English to French: The open science movement" inputs = tokenizer(prompt, return_tensors="pt").to("cuda") outputs = model.generate(**inputs, max_new_tokens=50) print(tokenizer.decode(outputs[0], skip_special_tokens=True))

CitationBigScience Workshop et al.. BLOOM: A 176B-Parameter Open-Access Multilingual Language Model. arXiv, 2022.

Terms in this paper