KL Divergence
تباعد KL
مقياس لمدى اختلاف توزيعين احتماليين، يُستخدم في التعلم المعزز كحدّ عقوبة يمنع السياسة المُحدَّثة من الانحراف بشدة عن السياسة المرجعية.
A measure of how much two probability distributions differ, used in RL as a penalty term preventing the updated policy from diverging drastically from the reference policy.
Also translated asالافتراق KL، الفجوة الاحتمالية النسبية، انحراف KL، انحراف كولباك-لايبلر، تباعد كولباك-لايبلر، مقياس الهدر المعلوماتي
First appears in this corpus in: On Information and Sufficiency (1951)
Appears in these papers
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware2023in the sky ✦
- Atlas: Few-shot Learning with Retrieval Augmented Language Models2023in the sky ✦
- beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework2017in the sky ✦
- beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework2017in the sky ✦
- DALL·E: Zero-Shot Text-to-Image Generation2021in the sky ✦
- DALL·E: Zero-Shot Text-to-Image Generation2021in the sky ✦
- Denoising Diffusion Implicit Models2021in the sky ✦
- Denoising Diffusion Probabilistic Models2020in the sky ✦
- Training Data-Efficient Image Transformers & Distillation Through Attention2021in the sky ✦
- DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter2019in the sky ✦
- Direct Preference Optimization: Your Language Model Is Secretly a Reward Model2023in the sky ✦
- Mastering Diverse Domains Through World Models2023in the sky ✦
- Mastering Diverse Domains Through World Models2023in the sky ✦
- High-Dimensional Continuous Control Using Generalized Advantage Estimation2016in the sky ✦
- Glow: Generative Flow with Invertible 1×1 Convolutions2018in the sky ✦
- Informer: Beyond Efficient Transformer for Long Sequence Time Series Forecasting2021in the sky ✦
- Informer: Beyond Efficient Transformer for Long Sequence Time Series Forecasting2021in the sky ✦
- Training Language Models to Follow Instructions with Human Feedback2022in the sky ✦
- Training Language Models to Follow Instructions with Human Feedback2022in the sky ✦
- On Information and Sufficiency1951in the sky ✦
- On Information and Sufficiency1951in the sky ✦
- Distilling the Knowledge in a Neural Network2015in the sky ✦
- Latent Dirichlet Allocation2003in the sky ✦
- Learning to Summarize from Human Feedback2020in the sky ✦
- Learning to Summarize from Human Feedback2020in the sky ✦
- Natural Gradient Works Efficiently in Learning1998in the sky ✦
- Variational Inference with Normalizing Flows2015in the sky ✦
- Proximal Policy Optimization Algorithms2017in the sky ✦
- RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback2023in the sky ✦
- RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback2023in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Visualizing Data Using t-SNE2008in the sky ✦
- Trust Region Policy Optimization2015in the sky ✦
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction2018in the sky ✦
- Auto-Encoding Variational Bayes2013in the sky ✦
- Auto-Encoding Variational Bayes2013in the sky ✦
- Neural Discrete Representation Learning2017in the sky ✦
- WebGPT: Browser-Assisted Question-Answering with Human Feedback2021in the sky ✦
- Wasserstein GAN2017in the sky ✦