Vision-Language-Action Model
نموذج الرؤية-اللغة-الفعل
نموذج يأخذ مدخلات بصرية وتعليمات لغوية ويُخرِج مباشرةً أفعالاً روبوتية منخفضة المستوى، موحِّداً الإدراك وفهم اللغة والتحكم الحركي في بنية واحدة من طرف إلى طرف.
A model that takes visual input and language instructions and directly outputs low-level robot actions — unifying perception, language understanding, and motor control in a single end-to-end architecture.
Also translated asنموذج VLA، نموذج الرؤية والفعل اللغوي
First appears in this corpus in: RT-1: Robotics Transformer for Real-World Control at Scale (2022)
Appears in these papers
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models2024in the sky ✦
- π₀: A Vision-Language-Action Flow Model for General Robot Control2024in the sky ✦
- RT-1: Robotics Transformer for Real-World Control at Scale2022in the sky ✦
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control2023in the sky ✦