Softmax
سوفت ماكس
دالة تحوّل متجه درجات خام إلى توزيع احتمالي حيث جميع القيم موجبة ومجموعها 1، مع تضخيم القيمة الأكبر.
A function that converts a vector of raw scores into a probability distribution where all values are positive and sum to 1, amplifying the largest value.
Also translated asالحد الأقصى اللين، دالة التطبيع الأُسِّية، دالة تعظيم الفروق اللينة، دالة سوفت ماكس الترجحيية
First appears in this corpus in: The Regression Analysis of Binary Sequences (1958)
Appears in these papers
- ImageNet Classification with Deep Convolutional Neural Networks2012in the sky ✦
- Mastering the Game of Go with Deep Neural Networks and Tree Search2016in the sky ✦
- A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play2018in the sky ✦
- Neural Machine Translation by Jointly Learning to Align and Translate2014in the sky ✦
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting2021in the sky ✦
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting2021in the sky ✦
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding2018in the sky ✦
- BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer2019in the sky ✦
- Unsupervised Visual Representation Learning by Context Prediction2015in the sky ✦
- Representation Learning with Contrastive Predictive Coding2018in the sky ✦
- DALL·E: Zero-Shot Text-to-Image Generation2021in the sky ✦
- DeBERTa: Decoding-Enhanced BERT with Disentangled Attention2020in the sky ✦
- DeBERTa: Decoding-Enhanced BERT with Disentangled Attention2020in the sky ✦
- DeepSeek-V3 Technical Report2024in the sky ✦
- DeepWalk: Online Learning of Social Representations2014in the sky ✦
- Training Data-Efficient Image Transformers & Distillation Through Attention2021in the sky ✦
- Training Data-Efficient Image Transformers & Distillation Through Attention2021in the sky ✦
- Densely Connected Convolutional Networks2017in the sky ✦
- End-to-End Object Detection with Transformers2020in the sky ✦
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion2023in the sky ✦
- Deep Interest Network for Click-Through Rate Prediction2018in the sky ✦
- Deep Interest Network for Click-Through Rate Prediction2018in the sky ✦
- Emerging Properties in Self-Supervised Vision Transformers2021in the sky ✦
- Emerging Properties in Self-Supervised Vision Transformers2021in the sky ✦
- DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter2019in the sky ✦
- DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter2019in the sky ✦
- ELECTRA: Pre-Training Text Encoders as Discriminators Rather Than Generators2020in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- Fast R-CNN2015in the sky ✦
- Fast R-CNN2015in the sky ✦
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks2015in the sky ✦
- Explaining and Harnessing Adversarial Examples2015in the sky ✦
- Explaining and Harnessing Adversarial Examples2015in the sky ✦
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning2023in the sky ✦
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness2022in the sky ✦
- Graph Attention Networks2018in the sky ✦
- Graph Attention Networks2018in the sky ✦
- Gaussian Processes for Machine Learning2006in the sky ✦
- Semi-Supervised Classification with Graph Convolutional Networks2017in the sky ✦
- Going Deeper with Convolutions2014in the sky ✦
- Improving Language Understanding by Generative Pre-Training2018in the sky ✦
- Do Transformers Really Perform Bad for Graph Representation?2021in the sky ✦
- Do Transformers Really Perform Bad for Graph Representation?2021in the sky ✦
- Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset2017in the sky ✦
- Informer: Beyond Efficient Transformer for Long Sequence Time Series Forecasting2021in the sky ✦
- Distilling the Knowledge in a Neural Network2015in the sky ✦
- When Does Label Smoothing Help?2019in the sky ✦
- When Does Label Smoothing Help?2019in the sky ✦
- "Why Should I Trust You?": Explaining the Predictions of Any Classifier2016in the sky ✦
- LLM.int8(): 8-Bit Matrix Multiplication for Transformers at Scale2022in the sky ✦
- The Regression Analysis of Binary Sequences1958in the sky ✦
- Mask R-CNN2017in the sky ✦
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism2019in the sky ✦
- Mixtral of Experts2024in the sky ✦
- Momentum Contrast for Unsupervised Visual Representation Learning2020in the sky ✦
- Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer2022in the sky ✦
- A Neural Probabilistic Language Model2003in the sky ✦
- A Neural Probabilistic Language Model2003in the sky ✦
- Neural Turing Machines2014in the sky ✦
- node2vec: Scalable Feature Learning for Networks2016in the sky ✦
- PaLM: Scaling Language Modeling with Pathways2022in the sky ✦
- Pixel Recurrent Neural Networks2016in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation2014in the sky ✦
- Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning1992in the sky ✦
- RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback2023in the sky ✦
- Recurrent Neural Networks (RNNs): A Gentle Introduction and Overview2019in the sky ✦
- Unsupervised Representation Learning by Predicting Image Rotations2018in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Sequence to Sequence Learning with Neural Networks2014in the sky ✦
- Show and Tell: A Neural Image Caption Generator2015in the sky ✦
- Show and Tell: A Neural Image Caption Generator2015in the sky ✦
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention2015in the sky ✦
- Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows2021in the sky ✦
- Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows2021in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Attention Is All You Need2017in the sky ✦
- VQA: Visual Question Answering2015in the sky ✦
- VQA: Visual Question Answering2015in the sky ✦
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations2020in the sky ✦
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations2020in the sky ✦
- WaveNet: A Generative Model for Raw Audio2016in the sky ✦
- Distributed Representations of Words and Phrases and Their Compositionality2013in the sky ✦
- Understanding the Difficulty of Training Deep Feedforward Neural Networks2010in the sky ✦
- YOLOv3: An Incremental Improvement2018in the sky ✦
- Deep Neural Networks for YouTube Recommendations2016in the sky ✦
- Deep Neural Networks for YouTube Recommendations2016in the sky ✦