Language Models2023intermediate12 min read
Toolformer: Language Models Can Teach Themselves to Use Tools
Toolformer: نماذج اللغة تستطيع تعليم نفسها استخدام الأدوات
Schick, T. · Dwivedi-Yu, J. · Dessì, R. · Raileanu, R. · Lomeli, M. · Zettlemoyer, L. · Cancedda, N. · Scialom, T. — NeurIPS
The problem
Large language models can solve new tasks from just a few examples, yet they paradoxically fail at basic skills that simple specialized systems handle easily — arithmetic, factual lookup, calendar reasoning, even translating low-resource languages. Scaling alone does not fix these weaknesses. Previous approaches either required massive human annotation to teach , or restricted the model to narrow task-specific settings where tools were always invoked, stripping away the model's general-purpose flexibility.
The contribution
A self-supervised method that teaches a to use external tools — calculator, Q&A system, search engines, translation system, and calendar — via embedded directly in the text. The model itself decides when to call which tool, what arguments to pass, and how to incorporate results. The pipeline requires only a handful of demonstrations per API: the model generates candidate API calls, executes them, and keeps only those that reduce on subsequent tokens. Fine-tuned on GPT-J (6.7B), Toolformer achieves performance competitive with GPT-3 (175B) on several downstream tasks.
The impact
Toolformer established that tool use could be learned without heavy human supervision, inspiring the wave of tool-augmented LLMs that followed. It directly influenced HuggingGPT, Gorilla, and the function-calling capabilities now standard in commercial APIs like GPT-4 and Claude. The key insight — that the model's own loss signal can supervise tool-use learning — became a foundational idea in the agentic AI movement.
Imagine a brilliant author locked in a room with nothing but a typewriter. She can write beautifully about the history of Rome, but ask her to calculate compound interest and she freezes. Now imagine sliding tools under the door — a calculator, an encyclopedia, a calendar — but with no instruction manual. She picks each one up, tries it on different sentences, and keeps using it only when the answer genuinely makes her writing better.
That is Toolformer. The language model doesn't need a teacher to explain when to use the calculator. It discovers this itself by testing whether the tool's response actually helps predict the next word. No human labels, no rigid rules — just a model learning to reach for the right tool at the right moment.
The paradox: brilliant at language, blind at basics
Large language models can write poetry, summarize legal documents, and pass exams — yet a pocket calculator outperforms them at arithmetic, and a simple database lookup beats them at retrieving facts. This is the core paradox: the same scaling that produces emergent reasoning abilities does not fix these fundamental blind spots.
The problem runs deeper than simple ignorance. Language models hallucinate facts because they have no mechanism to verify claims against an authoritative source. They fail at arithmetic because their training objective — predicting the next — never required learning precise computation. And they cannot access real-time information because their knowledge is frozen at the moment training data was collected.
Humans solve this effortlessly. We reach for a calculator when numbers get complex, search the web when memory fails, and check a calendar when dates matter. The question is: can a language model learn to do the same?
The key idea: API calls are just text
The first clever insight of Toolformer is representational: an API call can be written as a text sequence using special tokens. The call has a name, an input, and a response, all enclosed in markers. For example, a calculator call looks like: [Calculator(171 + 312) → 483].
By encoding tool calls this way, the model does not need a separate module or architecture change. It simply learns to generate these special tokens at appropriate positions in the text. To the , calling a tool is no different from generating any other sequence of tokens — the magic is in learning when to generate them.
This design makes Toolformer API-agnostic: any tool whose input and output can be expressed as text can be plugged in. The paper demonstrates five tools — a calculator, two search engines (Wikipedia and a web search), a question-answering system, a machine translation system, and a calendar — but the framework is extensible to any text-based API.
The pipeline: sample, execute, filter
The heart of Toolformer is a three-step pipeline that converts a plain text dataset into an augmented dataset containing useful API calls. Think of it as a quality-control assembly line: the model proposes many possible tool insertions, then reality checks them, and finally only the genuinely helpful ones survive.
Step 1 — Sample: For each sentence in the training data, the model proposes candidate API calls at promising positions. It does this using : a few demonstrations showing example API calls are prepended to the text, and the model generates candidate calls by sampling. Positions are ranked by how likely the model thinks a <API> token should appear there.
Step 2 — Execute: Each candidate API call is actually executed — the calculator computes the arithmetic, the search engine retrieves results, the Q&A model answers the question. This is the reality check: the tool's actual output replaces the model's speculation.
Step 3 — Filter: This is the critical step. A call is kept only if inserting the tool's response reduces the model's loss on predicting the tokens that follow. If the tool's answer does not help — or actively hurts — the call is discarded. The surviving calls form the augmented training dataset.
The filter: did the tool actually help?
The filtering criterion is what makes the entire approach self-supervised. The idea is simple: compare how well the model predicts the tokens after the API call in two scenarios — with the tool's response included, and without it (or with an empty response).
The model computes a weighted cross-entropy loss over the tokens following the API call position. If the loss with the tool's result is substantially lower than the loss without it, the API call is considered helpful and retained. The threshold controls how much improvement is required.
Intuitively, this asks: "Does knowing the tool's answer make the model less surprised by what comes next?" If the calculator returns 483 and the text continues with "483," then yes — the tool dramatically reduced the model's uncertainty. If the tool's answer is irrelevant, the loss barely changes and the call is discarded.
The toolbox: five APIs, one model
Toolformer demonstrates its approach with five diverse tools, each addressing a different weakness of language models:
Calculator — handles arithmetic that language models routinely get wrong. The model learns to emit expressions like Calculator(171 × 3) and insert the result.
Wikipedia Search — retrieves relevant passages when factual grounding is needed. The model generates queries like WikiSearch("Eiffel Tower height").
Web Search — a broader search engine for general knowledge retrieval when Wikipedia does not cover the topic.
Question Answering — calls an external QA model that can answer factual questions based on retrieved evidence, complementing the search tools.
Machine Translation — translates text to or from other languages, particularly useful for low-resource language content.
Calendar — returns the current date, allowing the model to answer time-dependent questions that would otherwise be impossible given a fixed training cutoff.
Each tool requires only 5–10 hand-written demonstration examples in the prompt. From these minimal seeds, the model bootstraps thousands of training examples through the sample-execute-filter pipeline.
Training: fine-tuning GPT-J on augmented data
The base model is GPT-J with 6.7 billion parameters. The pipeline annotates a subset of the CCNet dataset with API calls — up to 25,000 examples per API tool. Crucially, the augmented dataset also contains the original unannotated examples, so the model does not forget how to generate regular text without tools.
Training uses a standard autoregressive language modeling loss. The model learns to predict the API call tokens (including the special <API> and </API> markers) as part of the regular token stream. At time, a small modification to the decoding strategy allows API calls: whenever the <API> token appears among the top- most likely next tokens (not just the top-1), the model is allowed to begin a tool call. This increases tool use without requiring the model to be fully confident that a tool is needed.
The training is lightweight: up to 2,000 steps on 8 NVIDIA A100 GPUs using DeepSpeed ZeRO-3, with a of 128 and of . The key finding is that this does not degrade the model's core language modeling ability — perplexity on standard benchmarks remains unchanged.
Simplified to show the idea — not the real implementation.
def should_keep_api_call(model, text, position, api_call, response, tau):
"""Keep an API call only if the tool's response reduces loss."""
# Loss WITH the tool response included
loss_with_tool = compute_loss(model, text, position,
insert=api_call + " → " + response)
# Baseline 1: no API call at all
loss_no_call = compute_loss(model, text, position, insert=None)
# Baseline 2: API call without the response
loss_no_response = compute_loss(model, text, position,
insert=api_call + " → ")
# Use the better (lower) baseline
loss_baseline = min(loss_no_call, loss_no_response)
# Keep only if tool response helps enough
return (loss_baseline - loss_with_tool) >= tauResults: a 6.7B model challenges a 175B giant
Toolformer is evaluated in a zero-shot setting — no task-specific examples are provided at inference time. The results are striking:
On math benchmarks (ASDiv, SVAMP, MAWPS), Toolformer with the calculator tool dramatically outperforms both GPT-J and the much larger OPT (66B) and GPT-3 (175B). The calculator gives Toolformer near-perfect arithmetic, which no amount of scaling can match.
On factual QA (LAMA's SQuAD, GoogleRE, T-REx subsets), Toolformer with search tools outperforms GPT-J by a wide margin. The search and QA tools ground the model's answers in retrieved facts rather than memorized (and potentially hallucinated) knowledge.
On temporal reasoning using the TempLAMA , the calendar tool gives Toolformer a substantial edge — it can answer questions about current dates that any frozen model would get wrong.
Importantly, on tasks where tools are not helpful (creative generation, general NLU), Toolformer performs identically to the base GPT-J model. It does not lose generality.
Model size: tool use is an emergent ability
An important finding is that tool-use ability does not emerge at all model sizes. The authors tested GPT-2 variants from 124M to 1.6B parameters. Smaller models could not learn to use tools effectively — their performance with Toolformer was barely better than the baseline, and sometimes worse.
This suggests that learning when to use a tool requires a certain level of language understanding that only emerges at larger scales. The model needs enough capacity to represent both the concept of tool calling (a meta-skill) and the actual language understanding needed to decide which tool is appropriate for a given context.
At the GPT-J scale (6.7B), the capability clearly emerges: the model consistently and correctly inserts API calls where they help and refrains from calling tools where they would not. This aligns with the broader observation that many interesting LLM capabilities are emergent — appearing only beyond a threshold of scale.
The road to tool-augmented LLMs
2021
WebGPT — tool use via human feedback
OpenAI trained GPT-3 to use a web browser via human demonstration and reinforcement learning. Effective, but required thousands of human-labeled episodes.
2022
Self-Instruct — bootstrapping instruction data
Wang et al. showed that LLMs can generate their own instruction-following data, reducing the need for human annotation. Toolformer extends this idea to tool use.
2023
Toolformer
Self-supervised tool use via API calls as text tokens. A 6.7B model with tools competes with a 175B model without them. No human annotation of tool calls needed.
2023
ReAct — reasoning and acting together
Yao et al. interleaved chain-of-thought reasoning with tool actions, allowing the model to reason about what tool to use next. Addressed Toolformer's inability to chain.
2023
HuggingGPT — orchestrating AI models as tools
Extended the tool-use paradigm to treat entire AI models as tools. A controller LLM plans, selects, and orchestrates specialist models from Hugging Face.
2024
Function calling becomes standard
GPT-4, Claude, and Gemini all ship with built-in function-calling capabilities. Tool use moves from research novelty to production standard, building on the foundations laid by Toolformer.
Toolformer's lasting contribution is not any specific tool or benchmark result — it is the demonstration that a language model's own loss signal can supervise tool-use learning without human annotation. This insight unlocked the tool-augmented LLM era. Today, every major language model can call functions, browse the web, execute code, and interact with databases. The line from Toolformer to modern agentic AI is direct and unbroken.
CitationSchick, Dwivedi-Yu, Dessì, Raileanu, Lomeli, Zettlemoyer, Cancedda, Scialom. Toolformer: Language Models Can Teach Themselves to Use Tools. NeurIPS, 2023.
Terms in this paper
- Tool Useاستخدام الأدوات
- Self-Supervised Learningالتعلم ذاتي الإشراف
- API Callsاستدعاءات الواجهات البرمجية
- Perplexityمعيار الحيرة الاحتمالية
- Fine-Tuningالضبط الدقيق
- In-Context Learningالتعلم في السياق
- Zero-Shotالنمط الصفري
- Language Modelالنموذج اللغوي
- Cross Entropyالعشوائية المتقاطعة
- Autoregressive Modelالنموذج التوليدي التراجعي
- Tokenوحدة لغوية (رمز)
- Tokenizationتجزئة النصوص
- Few-Shotالنمط القليل العيّنات
- Pretrainingالتدريب المسبق
- Benchmarkالمعيار المرجعي