Glossary

V-trace

V-trace

خوارزمية تصحيح خارج السياسة تستخدم أوزان أخذ عيّنات الأهمية المقطوعة لتقدير دالة قيمة الحالة من بيانات وُلِّدت بسياسة سلوك مختلفة عن السياسة المستهدفة. تتيح التدريب الموزَّع المستقر على نطاق واسع.

An off-policy correction algorithm that uses truncated importance sampling weights to estimate a state-value function from data generated by a behavior policy different from the target policy. Enables stable distributed training at scale.

Also translated asتتبّع القيمة، هدف V-trace

First appears in this corpus in: Grandmaster Level in StarCraft II Using Multi-Agent Reinforcement Learning (2019)

Appears in these papers