Reinforcement Learning2023intermediate11 min read
HuggingGPT: Solving AI Tasks with ChatGPT and Its Friends in Hugging Face
HuggingGPT: كيف يقود ChatGPT فريقاً من النماذج المتخصّصة على Hugging Face
Shen, Y. · Song, K. · Tan, X. · Li, D. · Lu, W. · Zhuang, Y. — NeurIPS
The problem
By 2023, thousands of specialized AI models existed on platforms like Hugging Face — models for , image generation, speech synthesis, text classification, and more. But each model only handled one narrow task. If a user wanted something complex, like "detect the pose of a person in this photo, generate a new image based on that pose, then describe it aloud," they had to manually find the right models, chain them together, and handle data flow between them. Meanwhile, large language models like ChatGPT were excellent at understanding requests and reasoning, but could not process images, audio, or video directly. The gap was clear: LLMs had the brain but lacked the hands, and expert models had the hands but lacked the brain.
The contribution
HuggingGPT: a system that uses an LLM (ChatGPT) as a controller to orchestrate specialized AI models from Hugging Face. The workflow has four stages — (decomposing user requests into subtasks with dependencies), (choosing the best expert model for each subtask based on model descriptions), task execution (running models with proper resource dependency handling), and response generation (summarizing all results into a coherent answer). By treating language as a universal interface between the LLM controller and expert models, HuggingGPT can tackle complex tasks across vision, speech, language, and video domains without any architectural changes.
The impact
HuggingGPT was among the first systems to demonstrate the "LLM as controller" paradigm for orchestrating external tools, directly inspiring the wave of systems that followed — from AutoGPT and BabyAGI to ChatGPT Plugins and LangChain frameworks. It showed that language itself could serve as the middleware between a reasoning engine and a universe of specialist models, pioneering a design pattern now central to AI agent development.
Imagine a hospital emergency room. A patient arrives with a complex case: they need an X-ray, blood work, an MRI, and finally a specialist consultation. No single doctor does all of this. Instead, a triage coordinator reads the situation, decides which specialists to call and in what order, routes the patient through each one, and finally assembles a complete report.
HuggingGPT works the same way. ChatGPT is the triage coordinator — it reads your request, decides which AI specialist models to call from Hugging Face, routes your data through each one in the right order, and assembles a final answer. The coordinator never performs surgery, but without it the specialists cannot collaborate.
The core idea: language as a universal interface
The key insight behind HuggingGPT is deceptively simple: every AI model on Hugging Face comes with a text description of what it does. An object detection model says "I detect objects in images," a model says "I convert text to audio." These descriptions are already in natural language — exactly what LLMs understand best.
So the authors asked: what if we feed these model descriptions into ChatGPT's prompt and let it act as a dispatcher? The user describes what they want in plain language, ChatGPT reads the model descriptions and selects the right ones, calls them in the right order, and stitches the outputs together. Language becomes the middleware — the glue between the reasoning LLM and hundreds of specialist models.
This is a radical shift from the traditional approach of building a single monolithic model that tries to do everything. Instead, HuggingGPT treats the LLM as a brain that coordinates many small experts, each excellent at one thing. The result is a system that inherits the reasoning power of ChatGPT and the specialist capabilities of the entire Hugging Face ecosystem.
Stage 1: Task planning — the brain decomposes the problem
When a user sends a request like "detect the pose of the person in this photo, generate a new image of a girl reading based on that pose, then describe the new image aloud," HuggingGPT's first job is to parse this into structured subtasks. ChatGPT analyzes the request and produces a JSON task list, where each task has a type (like "pose-detection" or "text-to-speech"), an ID, a dependency list linking it to previous tasks, and arguments specifying inputs.
The dependencies are critical. In the example above, the pose-to-image task depends on the pose-detection task — it cannot run until pose detection produces the pose image. HuggingGPT uses a special <resource> tag to represent these dynamic dependencies. During planning, if task 2 needs the output of task 1, its arguments reference <resource>-1 instead of a fixed file. At execution time, this placeholder gets replaced with the actual output.
To help ChatGPT understand what good planning looks like, the prompt includes few-shot demonstrations: example user requests paired with their expected task decompositions. These demonstrations teach the LLM the JSON format, the available task types, and how to handle dependencies between tasks.
Stage 2: Model selection — choosing the right expert
Once the task plan is ready, HuggingGPT needs to assign a specific model to each task. Hugging Face hosts thousands of models, so putting all of them into the prompt would exceed the . HuggingGPT solves this with a two-step filtering process.
First, it filters models by task type: if the task is "object-detection," only object detection models are considered. Second, it ranks those models by their download count on Hugging Face — a proxy for popularity and quality — and selects the top-K candidates. These K model descriptions (including the model name, metadata, and function description) are placed in the prompt, and ChatGPT picks the single best one, outputting its choice as a JSON object with the model ID and its reasoning.
This design is deliberately open-ended: to add a new specialist model, you just upload it to Hugging Face with a good description. No retraining, no prompt rewriting, no code changes — the system discovers and uses new models automatically through their text descriptions.
Stage 3: Task execution — running the experts
With models assigned, HuggingGPT executes each task. The critical challenge here is handling resource dependencies — task outputs that serve as inputs to downstream tasks. HuggingGPT dynamically replaces the <resource>-task_id placeholders with actual generated outputs before launching each task.
Tasks without dependencies run in parallel to maximize efficiency. For example, if the plan includes both and object detection on the same image, these run simultaneously since neither depends on the other. Only tasks that need outputs from previous ones wait.
HuggingGPT uses a hybrid endpoint strategy: commonly used or slow models run on local servers for speed, while less common models run through Hugging Face's cloud API. Local endpoints take priority, and the system falls back to cloud endpoints only when a model is not deployed locally.
Stage 4: Response generation — assembling the answer
After all expert models finish, HuggingGPT collects their structured outputs — bounding boxes with confidence scores, generated images, transcribed text, answer distributions — and feeds them back to ChatGPT along with the full execution log. ChatGPT then synthesizes everything into a natural language response that directly addresses the user's original request.
This is not mere concatenation. The LLM actively interprets the results: it reads detection confidence scores, compares model outputs, and provides a coherent narrative. For example, after running object detection, image classification, and on the same image, ChatGPT might say: "The image shows a herd of giraffes and zebras. I detected five objects with scores above 97%." The response includes file paths for any generated images or audio.
How well does it plan? Evaluating the brain
Since task planning is the most critical stage — a bad plan means the entire pipeline fails — the authors focused their quantitative evaluation on measuring how well different LLMs plan. They tested on three categories of increasing complexity.
Single tasks are requests requiring just one model call, like "classify this image." Sequential tasks require multiple tasks in a chain, like "detect pose → generate image from pose." Graph tasks are the most complex — they decompose into directed acyclic graphs where some tasks run in parallel and others depend on multiple predecessors.
The results reveal a clear capability hierarchy. GPT-3.5 achieved 52.6% accuracy on single tasks, 51.9 F1 on sequential tasks, and 50.5 GPT-4 Score on graph tasks — dramatically outperforming open-source models like Alpaca-7b (6.5% single-task accuracy) and Vicuna-7b (23.9%). On a human-annotated dataset of especially complex requests, even GPT-4 achieved only 41.4% accuracy on sequential tasks, underscoring how hard reliable task planning remains.
The ablation studies revealed two important findings about demonstrations in the prompt. First, increasing the variety of task types covered by the demonstrations consistently improved planning quality — showing that the LLM benefits from seeing diverse examples, not just more of them. Second, performance gains from adding more demonstrations plateau after about 4 shots; beyond that, more examples don't help much.
Human evaluation on 130 diverse requests confirmed the objective findings: GPT-3.5 achieved a 91.2% passing rate on task planning (meaning the plan could actually be executed) and a 63.1% overall success rate (the final response actually solved the user's request). Open-source models lagged far behind, with Alpaca-13b at 6.9% success rate.
Limitations: where the system breaks
The authors are candid about four key limitations. First, planning quality is entirely bottlenecked by the LLM — if ChatGPT misunderstands the request or creates a wrong dependency, the entire pipeline fails with no self-correction mechanism. Second, efficiency is a challenge: each request requires multiple LLM calls (one for planning, one per model selection, one for response generation), plus model inference time, making the end-to-end significant. Third, the context window limits how many model descriptions can fit in the prompt, constraining the selection pool. Fourth, LLM outputs are inherently stochastic — ChatGPT may occasionally produce malformed JSON or select inappropriate models, causing runtime exceptions.
Legacy: from concept to paradigm
2023
Toolformer
Showed LLMs can learn to insert API calls into text sequences — the first step toward teaching language models to use external tools.
2023
HuggingGPT (JARVIS)
Connected ChatGPT to the entire Hugging Face ecosystem with a four-stage pipeline. Demonstrated that language can serve as the universal interface between a reasoning controller and specialist models.
2023
AutoGPT & BabyAGI
Autonomous agents using iterative planning. Broader task scope but less structured execution than HuggingGPT.
2023
ChatGPT Plugins & Function Calling
OpenAI productionized the tool-use pattern. LLMs select and call APIs via structured function definitions — essentially HuggingGPT's model selection stage at industrial scale.
2024
Agentic frameworks mature
LangChain agents, CrewAI, and other frameworks formalized the decompose → route → execute → synthesize pattern that HuggingGPT pioneered.
HuggingGPT's deepest contribution is not the system itself but the architectural pattern it demonstrated: separate the reasoning engine (LLM) from the execution engines (specialist models), and use natural language as the protocol between them. This pattern now appears everywhere — from ChatGPT's function calling to Claude's to open-source agent frameworks. The idea that an LLM does not need to be a vision model or a speech model, it just needs to know when to call one, has become the foundational principle of agentic AI.
CitationShen, Song, Tan, Li, Lu, Zhuang. HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face. NeurIPS, 2023.
Terms in this paper
- Task Planningتخطيط المهام
- Model Selectionاختيار النموذج
- Tool Useاستخدام الأدوات
- Agentوكيل
- In-Context Learningالتعلم في السياق
- Prompt Engineeringهندسة التحفيز
- Large Language Modelالنموذج اللغوي الكبير
- Zero-Shotالنمط الصفري
- Fine-Tuningالضبط الدقيق
- Transfer Learningنقل التعلم
- Multimodalمتعدد الوسائط
- Inferenceالاستدلال
- Benchmarkالمعيار المرجعي