Language Models2023intermediate12 min read
The Falcon Series of Open Language Models
سلسلة Falcon من النماذج اللغوية المفتوحة
Almazrouei, E. · Alobeidli, H. · Alshamsi, A. · Cappelli, A. · Cojocaru, R. · Debbah, M. · Goffinet, E. · Hesslow, D. · Launay, J. · Malartic, Q. · Mazzotta, D. · Noune, B. · Pannier, B. · Penedo, G. — arXiv
The problem
By 2023, building state-of-the-art language models required massive curated datasets combining web crawls with hand-picked sources like books, academic papers, and code repositories. This curation was expensive, legally fraught, and hard to scale beyond 1–2 trillion tokens. Meanwhile, models were growing to hundreds of billions of parameters, demanding distributed across thousands of GPUs — often on expensive supercomputer clusters with high-bandwidth interconnects. No openly documented model had yet followed the Chinchilla at the 180B scale, and the gap between open and closed models was widening.
The contribution
The Falcon series — three causal -only models with 7B, 40B, and 180B parameters — trained predominantly on RefinedWeb, a 5-trillion- filtered and deduplicated web dataset. Falcon-180B was trained on 3.5 trillion tokens across 4,096 A100 GPUs on AWS cloud with only 50 Gbps interconnect, making it the largest openly documented training run following Chinchilla-optimal scaling. The team demonstrated that stringently filtered web data alone can match or outperform curated corpora, introduced for efficiency, and released the models and a 600B-token data extract under permissive licenses.
The impact
Falcon challenged the prevailing belief that curated data is essential for top performance, proving that scale and quality filtering of web data is a viable alternative. Falcon-180B became one of the strongest open-weight models at its release, nearing PaLM-2 Large performance while being fully documented and reproducible. The RefinedWeb dataset and pipeline influenced subsequent data-centric efforts, and the series accelerated the open ecosystem of large language models alongside LLaMA.
Imagine you're building a library to teach a student everything about the world. The traditional approach is like hiring curators to hand-pick the best books, academic journals, and encyclopedias — expensive, slow, and limited by what curators can find.
Falcon takes a different approach: it sends trucks to collect every publication in every bookstore, then runs each page through a meticulous quality filter — removing duplicates, stripping spam, checking language quality. The result? A library so vast and so clean that a student trained on it alone outperforms one who only read the curated collection.
The Falcon family: three scales, one philosophy
The Falcon series comprises three causal decoder-only models developed by the Technology Innovation Institute (TII) in Abu Dhabi. All three share the same fundamental design philosophy: scale high-quality web data, not hand-picked corpora.
Falcon-7B was trained on 1,500 billion tokens using 384 A100 GPUs. Falcon-40B consumed 1,000 billion tokens on the same GPU count. Falcon-180B — the flagship — was trained on 3,500 billion tokens across 4,096 A100 40GB GPUs, making it the largest openly documented run at the time. Falcon-7B can run on consumer hardware like an Apple M2 chip, while Falcon-180B typically requires 8×A100 80GB for inference.
The design philosophy rests on three pillars of scalability: performance scalability (following Chinchilla-optimal scaling laws so that more compute translates to predictable improvements), data scalability (building a pipeline that produces trillions of unique, high-quality tokens so data never needs to be repeated), and hardware scalability (engineering a distributed training system that runs efficiently on cloud infrastructure with limited interconnect bandwidth).
Data: the web is all you need (if you clean it well)
The conventional wisdom in 2023 was clear: to build a great language model, you need curated data. GPT-3 mixed web crawls with books and Wikipedia. The Pile combined 22 individual sources including academic papers, GitHub code, and Project Gutenberg books. PaLM devoted 50% of its training to conversational data from social media.
The Falcon team challenged this assumption head-on. They built a data processing pipeline called Macrodata Refinement that transforms raw CommonCrawl web data into a high-quality training corpus through three stages: text extraction and language identification, heuristic filtering (removing boilerplate, short documents, and low-quality text), and stringent deduplication (both exact and fuzzy matching to eliminate redundant content).
The result is RefinedWeb — a dataset of 5 trillion English tokens, the largest publicly documented web dataset at the time. In controlled experiments, small models trained on RefinedWeb alone outperformed models trained on The Pile (a carefully curated dataset) on natural language benchmarks. Even more surprisingly, when curated data was added to a strong web baseline, performance did not consistently improve — and in some cases it even degraded, likely due to loss of diversity.
The final Falcon training dataset is roughly 84% web data from RefinedWeb, with small additions of technical curated data (2%), books (6%), code (3%), and conversational data (5%). Critically, no data was repeated — every token was seen exactly once during training. This stands in contrast to models like LLaMA which trained for 1.7 epochs over their data. The team's ablations on code and multilinguality showed that adding 5% code and 10% multilingual European data caused negligible degradation in English -shot performance, validating a safe recipe for broadening the model's capabilities.
Architecture: proven recipes with inference-friendly tweaks
The Falcon architecture is based on PaLM, but each design decision was independently validated through small-scale ablations. The team prioritized hardware scalability and inference efficiency over marginal accuracy gains, reasoning that architectural tweaks rarely improve task performance as much as better data does.
The key architectural choices include:
Multiquery and Grouped Query . Standard multihead attention stores separate key-value pairs for every head — expensive during inference because the entire KV Cache must be stored. (from Shazeer, 2019) shares a single key-value head across all query heads, slashing the KV Cache by a factor equal to the number of heads. Falcon-7B and Falcon-40B use pure multiquery attention. For Falcon-180B, the team extended this to Grouped Query Attention with 8 KV heads (one per tensor parallel shard), balancing memory savings with quality. This was a pragmatic choice: with across 8 GPUs, each shard needs at least one KV head to avoid expensive communication.
Positional encoding. Falcon-7B and Falcon-40B use Rotary Positional Embeddings (RoPE), while Falcon-180B uses (Attention with Linear Biases). In ablations, RoPE showed only a marginal edge over ALiBi, and the team chose ALiBi for the largest model because it was already validated in their pipeline when Falcon-180B training began.
Parallel attention and feed-forward layers. Instead of computing attention and the feed-forward network sequentially, Falcon computes them in parallel within each Transformer block and sums the results. This reduces communication overhead in tensor- parallel setups and slightly improves .
No biases in linear layers. Removing terms from all linear layers reduces all-reduce communication in tensor parallelism (biases must be synced across shards) with no measurable quality loss.
Training at scale: 4,096 GPUs on the cloud
Training Falcon-180B required solving engineering challenges that most open projects had not faced. The team built Gigatron, a custom distributed training codebase combining three complementary parallelism strategies:
weaves together (each GPU processes a different mini-batch), tensor parallelism (a single layer is split across GPUs — Falcon-180B uses 8-way tensor parallelism), and (different layers live on different GPUs, with micro-batches flowing through them in sequence). This gives fine-grained control over memory and communication.
ZeRO-1 optimizer sharding partitions the optimizer states (momentum and in ) across data-parallel ranks, reducing per-GPU memory by roughly 3×. This was critical to fitting Falcon-180B on 40GB A100s instead of requiring 80GB cards.
Custom Triton kernels fused multiple operations (like attention + + mask) into single GPU kernels, reducing memory round-trips and achieving state-of-the-art training throughput.
All training was done in (bfloat16) precision — including the optimizer states. The team found that Mixed Precision Training with FP32 master weights was unnecessary for their setup, and pure BF16 simplified the pipeline while saving memory.
The entire system ran on AWS cloud infrastructure with only 50 Gbps of interconnect per GPU — dramatically lower than the 400-800 Gbps typical of dedicated supercomputers like NVIDIA's DGX clusters. Achieving high utilization under this constraint required careful overlap of computation and communication, and motivated architectural choices like parallel layers and bias removal that minimize cross-GPU synchronization.
Simplified to show the idea — not the real implementation.
# Falcon's parallel formulation inside each Transformer block
# Instead of sequential: y = FFN(Attn(x))
# Falcon computes: y = x + Attn(x) + FFN(x) (in parallel)
def falcon_block(x, layer_norm, attn, ffn):
# Single layer norm applied once
h = layer_norm(x)
# Attention and FFN computed in PARALLEL (not sequential)
attn_out = attn(h) # multiquery or grouped-query
ffn_out = ffn(h) # standard GeLU FFN, no GLU
# Sum both outputs with residual
return x + attn_out + ffn_outResults: competing with the best
The Falcon team evaluated their models on a broad set of natural language tasks using the EleutherAI Evaluation Harness, as well as on selected tasks from the PaLM and GPT-3 evaluation suites. The results tell a clear story of consistent scaling:
Falcon-7B achieved an aggregate zero-shot score of 60.8, outperforming models of similar scale. Falcon-40B reached 67.1, competitive with the original Chinchilla (70B parameters). Falcon-180B scored 70.3 on the same aggregate, outperforming PaLM (540B) and LLaMA 2 (70B), while nearing the performance of PaLM-2 Large — placing it among the top three language models at the time alongside GPT-4 and PaLM-2 Large.
On individual benchmarks, Falcon-180B demonstrated strong performance across common sense reasoning (HellaSwag, PIQA, Winogrande), question answering (ARC, OpenBookQA), and reading comprehension (LAMBADA, RACE). The model showed consistent gains across all scales, validating the data-centric approach.
The team was careful to note evaluation challenges: different papers use different prompts, different task selections, and different evaluation harnesses, making direct comparisons difficult. They built custom aggregates for their ablations and reported results with variance estimates to ensure transparency.
Training recipe: the details that matter
Falcon uses the optimizer with a cosine schedule and linear . The team validated several hyperparameter choices through ablations:
z-loss regularization adds a small penalty to the squared log of the softmax partition function, stabilizing training at large scale. This technique, borrowed from PaLM, prevents logit drift without affecting final performance.
of 0.1 proved optimal in ablations. Higher values degraded performance, while lower values offered no benefit.
Learning rate was carefully searched for each model scale. The team found that common recipes (like scaling LR inversely with model size) worked well but required for their specific setup.
The vocabulary uses a with 65,024 entries, trained on a representative subset of the pretraining data. For Falcon-180B, the input and output matrices are tied (shared weights), saving parameters without measurable quality loss.
Limitations and open questions
The Falcon team was transparent about limitations. Their ablations were conducted at small scale (1-3B parameters, 30-60B tokens), which may not capture behaviors that only emerge at larger scales — like outlier features that affect quantization at the 6B scale, or memorization patterns that disproportionately affect larger models.
The evaluation focused exclusively on English natural language tasks. Models trained primarily on web data may underperform on specialized domains like code or scientific reasoning compared to models with targeted curated data for those domains.
The team also acknowledged that Falcon is a base model without instruction-tuning or alignment. Without RLHF or similar alignment techniques, the model may generate harmful, biased, or factually incorrect content. The paper focuses on pretraining and leaves downstream alignment as future work.
Finally, the debate over data quality is not fully settled. While RefinedWeb's web-only approach works well on aggregate benchmarks, it is possible that task-specific performance (like code generation or mathematical reasoning) benefits from targeted curation at larger scales — a hypothesis the small-scale ablations cannot fully test.
The Falcon timeline and legacy
2022
RefinedWeb data pipeline
Development of the Macrodata Refinement pipeline begins. The goal: extract trillions of high-quality tokens from CommonCrawl through aggressive filtering and deduplication.
2022
Training begins
Custom Gigatron codebase kicks off pretraining in December 2022 on AWS cloud infrastructure with cost-efficient A100 40GB GPUs.
2023
Falcon-7B & Falcon-40B released
Both models released under Apache 2.0. Falcon-40B briefly topped the HuggingFace Open LLM Leaderboard, demonstrating the web-data-first approach works.
2023
Falcon-180B released
The flagship model is released with a responsible use license. At 3.5 trillion training tokens, it is the largest openly documented pretraining run and the best open model at the time.
2024
Falcon-2 multimodal
The second generation introduces Falcon-2 11B with vision-language capabilities, building a multimodal ecosystem on the foundations laid by the first generation.
Falcon's deepest contribution is methodological: it proved that data engineering — filtering, deduplication, and quality control at industrial scale — can substitute for expensive manual curation. This insight influenced subsequent models and datasets, and helped democratize access to high-quality training data.
The series also demonstrated that large-scale training does not require billion-dollar supercomputers. By engineering their software stack to work with commodity cloud interconnects, the Falcon team showed that a well-designed distributed system can close the gap between academic clusters and purpose-built AI supercomputers.
CitationAlmazrouei, Alobeidli, Alshamsi, Cappelli, Cojocaru, Debbah, Goffinet, Hesslow, Launay, Malartic, Mazzotta, Noune, Pannier, Penedo. The Falcon Series of Open Language Models. arXiv, 2023.
Terms in this paper
- Causal Language Modelالنموذج اللغوي السببي
- Decoderمفكّ الترميز
- Pretrainingالتدريب المسبق
- Scaling Lawsقوانين التوسعة
- Multiquery Attentionانتباه الاستعلام الأحادي
- Grouped Query Attentionانتباه الاستعلام المُجمَّع
- Tensor Parallelismتوازي مصفوفات الموتّرات
- Data Parallelismتوازي البيانات
- Pipeline Parallelismالتوازي المتسلسل للطبقات
- BF16الدقة العائمة بنظام برين 16 بت
- Tokenizerمـُجزئ النصوص
- Byte Pair Encoding (BPE)ترميز زوج البايت
- Fine-Tuningالضبط الدقيق
- Zero-Shotالنمط الصفري
- Benchmarkالمعيار المرجعي