Adaptive Heuristic Critic
الناقد التكيّفي الحدسي
بنية تعلم معزز مبكرة صممها سوتون (1984) تستخدم خطأ الفارق الزمني المُخصَّم لتقييم السياسة — وهي السلف المباشر لبنى الممثل-الناقد الحديثة.
An early reinforcement learning architecture designed by Sutton (1984) that uses discounted TD error for policy evaluation — the direct ancestor of modern actor-critic architectures.
Also translated asالناقد الاستدلالي التكيفي
First appears in this corpus in: Learning to Predict by the Methods of Temporal Differences (1988)
Appears in these papers