Baseline
الخط المرجعي
دالة تُطرح من العائد عند تقدير متّجه ميل السياسة لتخفيض التباين دون إدخال انحياز. عادةً ما تكون تقديراً لقيمة الحالة V(s).
A function subtracted from the return when estimating the policy gradient to reduce variance without introducing bias. Typically an estimate of the state value V(s).
Also translated asخط الأساس، المرجع
First appears in this corpus in: Policy Gradient Methods for Reinforcement Learning with Function Approximation (1999)
Appears in these papers
- Are Transformers Effective for Time Series Forecasting?2023in the sky ✦
- GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding2018in the sky ✦
- GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding2018in the sky ✦
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour2017in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation2014in the sky ✦
- A Unified Approach to Interpreting Model Predictions2017in the sky ✦
- End-to-End Training of Deep Visuomotor Policies2016in the sky ✦