Asymmetric Actor-Critic
الممثِّل-الناقد غير المتماثل
منهج تدريب يتلقى فيه الناقد (شبكة القيمة) معلومات مميَّزة متاحة فقط في المحاكاة (كالزوايا والسرعات الحقيقية) بينما الممثِّل (شبكة السياسة) لا يرى إلا المراقبات المُشوَّشة المتاحة على الروبوت الحقيقي.
A training approach where the critic (value network) receives privileged information available only in simulation (like true joint angles and velocities) while the actor (policy network) sees only the noisy observations available on the physical robot.
Also translated asالفاعل-الناقد اللامتماثل
First appears in this corpus in: Learning Dexterous In-Hand Manipulation (2019)
Appears in these papers