Robotics2022intermediate14 min read
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
نفِّذ ما أستطيع، لا ما أقول: ربط اللغة بقدرات الروبوت الفعلية
Ahn, M. · Brohan, A. · Brown, N. · Chebotar, Y. · Cortes, O. · David, B. · Finn, C. · Fu, C. · Gopalakrishnan, K. · Hausman, K. · Herzog, A. · Ho, D. · Hsu, J. · Ibarz, J. · Ichter, B. · Irpan, A. · Jang, E. · Jauregui Ruano, R. · Jeffrey, K. · Jesmonth, S. · Joshi, N. J. · Julian, R. · Kalashnikov, D. · Kuang, Y. · Lee, K.-H. · Levine, S. · Lu, Y. · Luu, L. · Parada, C. · Pastor, P. · Quiambao, J. · Rao, K. · Rettinghouse, J. · Reyes, D. · Sermanet, P. · Sievers, N. · Tan, C. · Toshev, A. · Vanhoucke, V. · Xia, F. · Xiao, T. · Xu, P. · Xu, S. · Yan, M. · Zeng, A. — CoRL
The problem
By 2022, large language models like PaLM and GPT-3 had absorbed enormous knowledge about everyday tasks from web text, making them capable of describing multi-step procedures in natural language. However, these models had never interacted with the physical world: they could suggest "use a vacuum cleaner" even if no vacuum existed in the scene, or propose actions the robot physically could not perform. Meanwhile, learned robotic skills could execute low-level actions reliably, but had no capacity for high-level reasoning or understanding abstract instructions like "I spilled my drink, can you help?". The gap between what language models say and what robots can do was the core problem.
The contribution
SayCan: a framework that combines a (the "Say" — what is useful) with learned functions (the "Can" — what is physically possible) to enable a robot to execute long-horizon, abstract natural-language instructions. The LLM scores each available skill by how likely it advances the instruction, the scores each skill by how likely it can succeed in the current state, and their product selects the best action. Evaluated on 101 real-world kitchen tasks with a mobile manipulator, SayCan achieved 84% success and 74% execution success — nearly doubling ungrounded baselines. It also demonstrated that improving the underlying LLM directly improves robot performance.
The impact
SayCan was one of the earliest demonstrations that large language models can serve as effective planners for real-world robots when properly grounded. It established the "LLM as planner + affordance " paradigm that influenced a generation of robotic systems including RT-1, RT-2, Inner Monologue, and Code as Policies. The insight that robot performance scales with LLM quality opened a bridge between NLP progress and robotic capability, inspiring the modern wave of vision-language-action models.
Imagine a brilliant chef who has memorized every recipe ever written but has never set foot in your kitchen. You ask them: "I spilled juice on the counter, can you help?" They might answer: "Use the dishwasher's steam-clean cycle!" — a fine idea, except your kitchen has no dishwasher, and the chef cannot see what's on your counter.
Now give the chef a helper who lives in the kitchen. The helper can't plan a meal, but they know exactly which tools are within arm's reach and which drawers actually open. Together, the chef proposes steps and the helper vetoes anything impossible.
SayCan works exactly this way: the language model is the chef (rich knowledge, no physical sense), the robot's value functions are the helper (grounded awareness, no planning ability), and their collaboration produces plans that are both smart and feasible.
The problem: language models dream, robots live
Large language models like PaLM and GPT-3 encode an impressive wealth of procedural knowledge. Ask one how to clean a spill and it produces a sensible multi-step narrative. But that narrative was learned from text, not from physical interaction. The model has no concept of what the robot's gripper can reach, which objects are currently visible, or whether the suggested tool even exists in the scene.
On the other side, robots equipped with learned visuomotor policies can reliably execute short skills — "pick up the sponge", "go to the table" — but they cannot decompose a vague instruction like "I just worked out, can you bring me a drink and a snack to recover?" into an ordered sequence of those skills. They lack the world knowledge and commonsense reasoning that language models provide.
The fundamental tension is between two complementary weaknesses: the language model knows what to do but not what is possible; the robot knows what is possible but not what to do. SayCan resolves this by treating the problem as a product of two probabilities.
The SayCan framework: Say × Can
SayCan stands for two complementary functions working together. The Say component uses a large language model as a scoring function. Given a high-level instruction like "I spilled my drink, can you help?", the LLM is presented with each available skill description (e.g. "find a sponge", "pick up the coke can", "go to the trash") and outputs the probability that each skill is a useful next step toward completing the instruction. This is the task-grounding: it connects abstract language to concrete skill names.
The Can component uses a learned value function — an affordance function — that captures the probability of successfully executing a skill in the current physical state. For instance, "pick up the sponge" has a high affordance value when the sponge is visible and reachable, but near zero when the robot is navigating an empty hallway. This is the world-grounding: it connects skill names to physical reality.
The combined probability of each skill making real progress on the instruction is the product of these two terms. At each step, the robot selects the skill with the highest combined score, executes it, observes the new state, and repeats until the LLM outputs a termination token. The result is an interpretable, step-by-step plan expressed entirely in natural language.
The Say side: scoring with a language model
Instead of asking the LLM to generate free-form text and hoping it names valid skills, SayCan uses the LLM as a scoring model. The key shift: rather than sampling open-ended completions, the system feeds each skill description as a candidate completion and reads the probability the LLM assigns to it. This constrains outputs to the set of skills the robot actually has.
The interaction is structured as a dialog. The user provides a high-level instruction ("How would you clean up after my lunch?") and the LLM produces a numbered step plan ("1. find a sponge, 2. pick up the sponge, 3. go to the table, 4. ..."). provides a few examples of this dialog structure so the LLM knows the format. At each step, SayCan scores all available skills, picks the top one, appends it to the running plan, and queries the LLM again for the next step.
This scoring approach has a crucial advantage over generation: it is interpretable. The user can inspect the LLM's probability for every candidate skill and understand why the robot chose its plan.
Simplified to show the idea — not the real implementation.
# Simplified SayCan loop
plan = []
state = observe_environment()
while True:
best_score = -1
best_skill = None
for skill in available_skills:
# "Say": LLM scores how useful this skill is
p_useful = llm.score(skill.description, instruction, plan)
# "Can": value function scores feasibility
p_feasible = skill.value_function(state)
# Combined: useful AND feasible
combined = p_useful * p_feasible
if combined > best_score:
best_score = combined
best_skill = skill
if best_skill.description == "done":
break
# Execute and observe new state
best_skill.execute()
plan.append(best_skill.description)
state = observe_environment()The Can side: affordance functions as world-grounding
The affordance function answers a simple question: "If I ask the robot to do this skill right now, will it succeed?" This is exactly what a value function provides in the framework. The robot receives a reward of 1.0 if it successfully completes a skill and 0.0 otherwise, making the value function a direct estimate of the probability of completion.
The authors trained language-conditioned value functions via temporal-difference reinforcement learning. A single multi-task model accepts both the current camera image and a language of the skill description, then outputs a Q-value between 0 and 1. When the sponge is right in front of the robot, "pick up the sponge" gets a high Q-value. When the robot is in an empty hallway, every "pick up" action gets a near-zero Q-value.
Think of the value function as the robot's spatial awareness: it forms a map of what is possible here and now across all 551 skills. This map changes with every step the robot takes, providing continuous feedback without needing explicit object detectors or scene descriptions. The SayCan system used both policies (for better execution success) and RL-trained value functions (for better affordance estimation), combining the strengths of both approaches.
Step-by-step: how SayCan builds a plan
SayCan builds plans iteratively, not all at once. At each decision point, it scores all skills, selects the best, executes it, and then asks the LLM again — this time including the skill just executed in the conversation history. This iterative approach has two advantages. First, the LLM can reason about what has already been done: after delivering a drink, it knows to move on to the snack rather than fetching another drink. Second, the affordance values update after each action because the robot's physical state has changed — it may have moved to a new location or picked up an object.
Consider the instruction: "I just worked out, can you bring me a drink and a snack to recover?" SayCan must infer that "recover from a workout" means something healthy — water and an apple, not soda and chips. It must also sequence the subtasks: navigate to the water, pick it up, bring it, put it down, then navigate to the apple, pick it up, bring it, put it down, and terminate. At each step the LLM provides commonsense reasoning about what to do next, and the value functions ensure the chosen skill is currently feasible.
Building the skill library: 551 skills across 7 families
SayCan's skill repertoire spans seven families of robotic behaviors: picking up objects, placing objects, going to locations, finding objects, putting objects down, knocking objects over, and placing objects in specific configurations. Across 17 common kitchen objects and 5 locations (two counters, a table, a trash can, and the user location), this produces 551 distinct skills.
Skills were trained using two complementary methods. Behavioral cloning (BC) used 68,000 human teleoperated demonstrations collected over 11 months with a fleet of 10 robots. Reinforcement learning (RL) used the MT-Opt framework in simulation with RetinaGAN . The BC policies achieved higher execution success rates, while the RL-trained value functions provided better affordance estimation — so SayCan used BC for execution and RL for scoring.
Both policy types are conditioned on language through a pre-trained Universal Sentence Encoder. The language embedding specifies which skill to perform, and a single multi-task network handles all skills within a family. This means adding a new skill requires only collecting demonstrations and adding a language label — no architecture changes.
Results: grounding nearly doubles performance
SayCan was evaluated on 101 real-world instructions spanning 7 families: single primitives, abstract nouns, abstract verbs, structured language, embodiment-dependent tasks, crowd-sourced queries, and long-horizon plans. Using PaLM (540B parameters) as the LLM, PaLM-SayCan achieved 84% planning success and 74% execution success in a mock kitchen, and 81% planning / 60% execution in a real office kitchen.
The ablation studies reveal the critical importance of both components. Without value function grounding (using only the LLM's top skill), planning success dropped to 67%. Without the LLM (using only behavioral cloning with the raw instruction), success was 0% on all multi-step tasks. This confirms the core thesis: neither component alone is sufficient, but together they produce a capable robotic system.
Particularly notable is the LLM scaling result: switching from the 137B FLAN model to the 540B PaLM model improved planning from 70% to 84% and execution from 61% to 74%. This was the first demonstration that improving a language model directly improves robot performance — an exciting bridge between NLP research and robotics.
New capabilities: drawers, chain-of-thought, and multilingual
SayCan demonstrated several extensions beyond basic planning. Adding new skills — such as drawer manipulation (open, close, put in, take out) — required only adding the skill descriptions as LLM options and providing value functions. With just a single new prompt example, PaLM-SayCan achieved 100% planning accuracy on drawer tasks.
A known weakness of vanilla SayCan was handling negation ("bring me a snack that is not an apple") and reasoning-heavy queries. By integrating chain-of-thought prompting, where the LLM first generates an explicit explanation before scoring skills, these problems were substantially improved. For example: "Can you bring a fruit-flavored drink without caffeine?" triggers the explanation "The user wants a drink that is fruit-flavored and has no caffeine, I will bring the lime soda" — correctly resolving the constraint.
Finally, since PaLM was trained on multilingual data, SayCan handled queries in Chinese, French, and Spanish with almost no performance drop — despite never being explicitly designed for multilingual operation.
Limitations and failure modes
SayCan inherits the limitations of both its components. From the LLM side, it struggles with negation, ambiguous references (e.g. "something with caffeine"), and occasionally terminates long plans prematurely — bringing one object but forgetting the second. Of all planning errors, 65% were traced to LLM failures and 35% to affordance misclassification.
From the skill side, the system is fundamentally limited by the repertoire of available skills. It cannot improvise actions outside its training distribution. If no skill matches what the instruction requires, no amount of LLM reasoning will help.
The system also lacks closed-loop replanning: if a skill fails mid-execution, SayCan has no mechanism to detect the failure and retry or replan. The Inner Monologue paper later addressed this limitation by incorporating environment feedback into the LLM's planning loop.
Legacy: from SayCan to foundation models for robotics
2022
SayCan (this paper)
Combined LLM scoring with learned affordance functions to ground language in robotic capabilities. Demonstrated 101 real-world kitchen tasks on a mobile manipulator with 84% planning and 74% execution success.
2022
Inner Monologue
Extended SayCan with closed-loop feedback — success detectors and scene descriptions fed back into the LLM — enabling replanning when skills fail.
2022
Code as Policies
Used LLMs to generate executable Python code for robot control rather than scoring fixed skill sets, greatly expanding the action space.
2022
RT-1
Robotics Transformer 1 trained a single end-to-end Transformer on 130k real demonstrations to map language instructions directly to robot actions, building on the large-scale data collection infrastructure pioneered in SayCan.
2023
RT-2
Merged a vision-language model with robotic action outputs, creating a single model that reasons about both visual scenes and motor commands — the vision-language-action model paradigm.
2024
Foundation models for robotics
The "LLM + grounding" paradigm SayCan established has evolved into vision-language-action models, where grounding, planning, and control merge into a single foundation model trained at scale.
SayCan's most enduring contribution may not be the specific algorithm, but the conceptual bridge it built. It showed that the vast knowledge in language models is not locked behind a text interface — it can be extracted, grounded, and made physically actionable. The idea that improving NLP directly improves robotics opened a research trajectory that continues to accelerate, from RT-1's data-driven approach to RT-2's unified vision-language-action model and beyond.
CitationAhn, Brohan, Brown, Chebotar, Cortes, David, Finn, Fu, Gopalakrishnan, Hausman, Herzog, Ho, Hsu, Ibarz, Ichter, Irpan, Jang, et al.. Do As I Can, Not As I Say: Grounding Language in Robotic Affordances. CoRL, 2022.
Terms in this paper
- Groundingالتأسيس
- Value Functionدالة تقييم العوائد
- Large Language Modelالنموذج اللغوي الكبير
- Planningالتخطيط
- Reinforcement Learningالتعلم المعزز
- Behavioral Cloningالاستنساخ السلوكي
- Visuomotor Policyالسياسة البصرية-الحركية
- Prompt Engineeringهندسة التحفيز
- Chain of Thoughtسلسلة التفكير
- Action Spaceفضاء الأفعال
- Reward Functionدالة صياغة المكافآت
- Sim-to-Real Transferالنقل من المحاكاة إلى الواقع
- Zero-Shotالنمط الصفري
- Multi-Task Learningالتعلّم متعدد المهام