Glossary

Guided Policy Search

البحث الموجَّه عن السياسة

خوارزمية بحث عن السياسة تُحوِّل المسألة إلى تعلّم بإشراف عبر التناوب بين: (1) أمثَلة مسارات بمتحكمات بسيطة تصل إلى الحالة الكاملة، و(2) تدريب سياسة شبكة عصبية على تقليد تلك المسارات من الملاحظات فقط.

A policy search algorithm that converts the problem into supervised learning by alternating between: (1) trajectory optimization with simple controllers that access full state, and (2) training a neural network policy to imitate those trajectories from observations alone.

Also translated asالبحث الإرشادي عن السياسة، GPS

First appears in this corpus in: End-to-End Training of Deep Visuomotor Policies (2016)

Appears in these papers