Guided Policy Search
البحث الموجَّه عن السياسة
خوارزمية بحث عن السياسة تُحوِّل المسألة إلى تعلّم بإشراف عبر التناوب بين: (1) أمثَلة مسارات بمتحكمات بسيطة تصل إلى الحالة الكاملة، و(2) تدريب سياسة شبكة عصبية على تقليد تلك المسارات من الملاحظات فقط.
A policy search algorithm that converts the problem into supervised learning by alternating between: (1) trajectory optimization with simple controllers that access full state, and (2) training a neural network policy to imitate those trajectories from observations alone.
Also translated asالبحث الإرشادي عن السياسة، GPS
First appears in this corpus in: End-to-End Training of Deep Visuomotor Policies (2016)
Appears in these papers