المعجم

العائد متعدد الخطوات

Multi-Step Return

هدف في التعلُّم المعزَّز يجمع مكافآت فعلية مخصومة من عدة خطوات زمنية قبل الاسترجاع من دالة القيمة المُقدَّرة. يُحقق توازناً بين انحياز الاسترجاع بخطوة واحدة وتباين عائد مونت كارلو الكامل.

A reinforcement learning target that accumulates actual discounted rewards over multiple time steps before bootstrapping from the estimated value function. Achieves a balance between the bias of one-step bootstrapping and the variance of full Monte Carlo returns.

تُرجم أيضاًالعائد من n خطوة