Return-to-Go
العائد المتبقّي
مجموع المكافآت المستقبلية من خطوة زمنية معيّنة حتى نهاية الحلقة. يُستخدم في محوِّل القرار بدلاً من المكافأة الآنية لتمكين الاشتراط على الأداء المستقبلي المرغوب عند الاختبار.
The sum of future rewards from a given timestep until the end of the episode. Used in Decision Transformer instead of immediate rewards to enable conditioning on desired future performance at test time.
Also translated asمجموع المكافآت المتبقّية، العائد المستقبلي
First appears in this corpus in: Decision Transformer: Reinforcement Learning via Sequence Modeling (2021)
Appears in these papers