Training
التدريب
عملية مواءمة معلمات النموذج وأوزانه عبر مراجعة البيانات لتقليل نسبة الخطأ.
Training
Also translated asمواءمة النموذج، التكييف الرقمي
First appears in this corpus in: Méthode Générale pour la Résolution des Systèmes d'Équations Simultanées (1847)
Appears in these papers
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- A Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting1997in the sky ✦
- Adaptive Subgradient Methods for Online Learning and Stochastic Optimization2011in the sky ✦
- Adam: A Method for Stochastic Optimization2014in the sky ✦
- Decoupled Weight Decay Regularization2019in the sky ✦
- Parameter-Efficient Transfer Learning for NLP2019in the sky ✦
- Intriguing Properties of Neural Networks2014in the sky ✦
- AI Safety via Debate2018in the sky ✦
- ImageNet Classification with Deep Convolutional Neural Networks2012in the sky ✦
- ALIGN: Scaling Up Visual and Vision-Language Representation Learning with Noisy Text Supervision2021in the sky ✦
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware2023in the sky ✦
- Alpaca: A Strong, Replicable Instruction-Following Model2023in the sky ✦
- Mastering the Game of Go Without Human Knowledge2017in the sky ✦
- Mastering the Game of Go with Deep Neural Networks and Tree Search2016in the sky ✦
- A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play2018in the sky ✦
- Atlas: Few-shot Learning with Retrieval Augmented Language Models2023in the sky ✦
- Learning Representations by Back-Propagating Errors1986in the sky ✦
- Bagging Predictors1996in the sky ✦
- BART: Denoising Sequence-to-Sequence Pre-Training for Natural Language Generation, Translation, and Comprehension2019in the sky ✦
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift2015in the sky ✦
- Practical Bayesian Optimization of Machine Learning Algorithms2012in the sky ✦
- BEiT: BERT Pre-Training of Image Transformers2021in the sky ✦
- Learning Long-Term Dependencies with Gradient Descent is Difficult1994in the sky ✦
- BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer2019in the sky ✦
- beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework2017in the sky ✦
- Neural Networks and the Bias/Variance Dilemma1992in the sky ✦
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model2022in the sky ✦
- Okapi at TREC-31994in the sky ✦
- A Learning Algorithm for Boltzmann Machines1985in the sky ✦
- BPR: Bayesian Personalized Ranking from Implicit Feedback2009in the sky ✦
- C4.5: Programs for Machine Learning1993in the sky ✦
- Dynamic Routing Between Capsules2017in the sky ✦
- Classification and Regression Trees1984in the sky ✦
- Training Compute-Optimal Large Language Models2022in the sky ✦
- Chronos: Learning the Language of Time Series2024in the sky ✦
- Classifier-Free Diffusion Guidance2022in the sky ✦
- Evaluating Large Language Models Trained on Code2021in the sky ✦
- Concrete Problems in AI Safety2016in the sky ✦
- Unsupervised Visual Representation Learning by Context Prediction2015in the sky ✦
- Contriever: Unsupervised Dense Information Retrieval with Contrastive Learning2022in the sky ✦
- A ConvNet for the 2020s2022in the sky ✦
- Representation Learning with Contrastive Predictive Coding2018in the sky ✦
- Conservative Q-Learning for Offline Reinforcement Learning2020in the sky ✦
- Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data2001in the sky ✦
- Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks2017in the sky ✦
- Cyclical Learning Rates for Training Neural Networks2017in the sky ✦
- DALL·E: Zero-Shot Text-to-Image Generation2021in the sky ✦
- A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence1956in the sky ✦
- Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks2015in the sky ✦
- Denoising Diffusion Implicit Models2021in the sky ✦
- Denoising Diffusion Probabilistic Models2020in the sky ✦
- DeBERTa: Decoding-Enhanced BERT with Disentangled Attention2020in the sky ✦
- Decision Transformer: Reinforcement Learning via Sequence Modeling2021in the sky ✦
- A Fast Learning Algorithm for Deep Belief Nets2006in the sky ✦
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding2016in the sky ✦
- Deep Learning2015in the sky ✦
- Deep Speech 2: End-to-End Speech Recognition in English and Mandarin2015in the sky ✦
- DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks2020in the sky ✦
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning2025in the sky ✦
- DeepSeek-V3 Technical Report2024in the sky ✦
- DeepWalk: Online Learning of Social Representations2014in the sky ✦
- Deformable Convolutional Networks2017in the sky ✦
- Training Data-Efficient Image Transformers & Distillation Through Attention2021in the sky ✦
- Extracting and Composing Robust Features with Denoising Autoencoders2008in the sky ✦
- K-SVD: An Algorithm for Designing Overcomplete Dictionaries for Sparse Representation2006in the sky ✦
- Diffusion Models Beat GANs on Image Synthesis2021in the sky ✦
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion2023in the sky ✦
- Deep Interest Network for Click-Through Rate Prediction2018in the sky ✦
- Emerging Properties in Self-Supervised Vision Transformers2021in the sky ✦
- DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter2019in the sky ✦
- Learning Distributed Representations of Concepts1986in the sky ✦
- Are Transformers Effective for Time Series Forecasting?2023in the sky ✦
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World2017in the sky ✦
- Reconciling Modern Machine-Learning Practice and the Classical Bias–Variance Trade-Off2019in the sky ✦
- Deep Reinforcement Learning with Double Q-Learning2016in the sky ✦
- Human-Level Control Through Deep Reinforcement Learning2015in the sky ✦
- Mastering Diverse Domains Through World Models2023in the sky ✦
- Dropout: A Simple Way to Prevent Neural Networks from Overfitting2014in the sky ✦
- Dueling Network Architectures for Deep Reinforcement Learning2016in the sky ✦
- EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks2019in the sky ✦
- Face Recognition Using Eigenfaces1991in the sky ✦
- ELIZA — A Computer Program for the Study of Natural Language Communication1966in the sky ✦
- Deep Contextualized Word Representations2018in the sky ✦
- Emergent Abilities of Large Language Models2022in the sky ✦
- FaceNet: A Unified Embedding for Face Recognition and Clustering2015in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- Fast R-CNN2015in the sky ✦
- Enriching Word Vectors with Subword Information2017in the sky ✦
- Fully Convolutional Networks for Semantic Segmentation2015in the sky ✦
- Explaining and Harnessing Adversarial Examples2015in the sky ✦
- Fusion-in-Decoder: Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering2021in the sky ✦
- Flamingo: a Visual Language Model for Few-Shot Learning2022in the sky ✦
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning2023in the sky ✦
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness2022in the sky ✦
- High-Dimensional Continuous Control Using Generalized Advantage Estimation2016in the sky ✦
- A Generalist Agent2022in the sky ✦
- Gaussian Processes for Machine Learning2006in the sky ✦
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering2023in the sky ✦
- Semi-Supervised Classification with Graph Convolutional Networks2017in the sky ✦
- Gaussian Error Linear Units (GELUs)2016in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- Gemma: Open Models Based on Gemini Research and Technology2024in the sky ✦
- How Powerful Are Graph Neural Networks?2019in the sky ✦
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models2021in the sky ✦
- GloVe: Global Vectors for Word Representation2014in the sky ✦
- Glow: Generative Flow with Invertible 1×1 Convolutions2018in the sky ✦
- Google's Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation2016in the sky ✦
- The Graph Neural Network Model2009in the sky ✦
- First Return, Then Explore2021in the sky ✦
- Going Deeper with Convolutions2014in the sky ✦
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher2022in the sky ✦
- Language Models Are Unsupervised Multitask Learners2019in the sky ✦
- Language Models Are Few-Shot Learners2020in the sky ✦
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints2023in the sky ✦
- Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization2017in the sky ✦
- Greedy Function Approximation: A Gradient Boosting Machine2001in the sky ✦
- On the Difficulty of Training Recurrent Neural Networks2013in the sky ✦
- Méthode Générale pour la Résolution des Systèmes d'Équations Simultanées1847in the sky ✦
- Inductive Representation Learning on Large Graphs2017in the sky ✦
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets2022in the sky ✦
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection2023in the sky ✦
- Group Normalization2018in the sky ✦
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling2014in the sky ✦
- Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification2015in the sky ✦
- The Organization of Behavior: A Neuropsychological Theory1949in the sky ✦
- Hindsight Experience Replay2017in the sky ✦
- Highway Networks2015in the sky ✦
- A Tutorial on Hidden Markov Models and Selected Applications in Speech Recognition1989in the sky ✦
- Histograms of Oriented Gradients for Human Detection2005in the sky ✦
- Neural Networks and Physical Systems with Emergent Collective Computational Abilities1982in the sky ✦
- Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization2018in the sky ✦
- Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset2017in the sky ✦
- The Mathematics of Statistical Machine Translation: Parameter Estimation1993in the sky ✦
- Curiosity-Driven Exploration by Self-Supervised Prediction2017in the sky ✦
- Induction of Decision Trees1986in the sky ✦
- ImageBind: One Embedding Space to Bind Them All2023in the sky ✦
- Imagen: Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding2022in the sky ✦
- Training Language Models to Follow Instructions with Human Feedback2022in the sky ✦
- Isolation Forest2008in the sky ✦
- Jukebox: A Generative Model for Music2020in the sky ✦
- Optimizing Neural Networks with Kronecker-Factored Approximate Curvature2015in the sky ✦
- On Information and Sufficiency1951in the sky ✦
- Distilling the Knowledge in a Neural Network2015in the sky ✦
- Foundations of the Theory of Probability1933in the sky ✦
- When Does Label Smoothing Help?2019in the sky ✦
- Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes2019in the sky ✦
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour2017in the sky ✦
- Regression Shrinkage and Selection via the Lasso1996in the sky ✦
- High-Resolution Image Synthesis with Latent Diffusion Models2022in the sky ✦
- Layer Normalization2016in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- LightGBM: A Highly Efficient Gradient Boosting Decision Tree2017in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- LLaMA: Open and Efficient Foundation Language Models2023in the sky ✦
- Visual Instruction Tuning2023in the sky ✦
- Visualizing the Loss Landscape of Neural Nets2018in the sky ✦
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks2019in the sky ✦
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks2019in the sky ✦
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces2023in the sky ✦
- Extension of the Law of Large Numbers to Dependent Quantities1906in the sky ✦
- A Logical Calculus of the Ideas Immanent in Nervous Activity1943in the sky ✦
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism2019in the sky ✦
- Mistral 7B2023in the sky ✦
- Mixed Precision Training2018in the sky ✦
- Mixtral of Experts2024in the sky ✦
- mixup: Beyond Empirical Risk Minimization2018in the sky ✦
- On the Mathematical Foundations of Theoretical Statistics1922in the sky ✦
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications2017in the sky ✦
- The Monte Carlo Method1949in the sky ✦
- Neural Message Passing for Quantum Chemistry2017in the sky ✦
- Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer2022in the sky ✦
- MusicLM: Generating Music From Text2023in the sky ✦
- MuZero: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model2020in the sky ✦
- MuZero: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model2020in the sky ✦
- N-BEATS: Neural Basis Expansion Analysis for Interpretable Time Series Forecasting2019in the sky ✦
- Natural Gradient Works Efficiently in Learning1998in the sky ✦
- Nearest Neighbor Pattern Classification1967in the sky ✦
- Neocognitron: A Self-Organizing Neural Network Model for Pattern Recognition1980in the sky ✦
- NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis2020in the sky ✦
- A Method for Solving the Convex Programming Problem with Convergence Rate O(1/k²)1983in the sky ✦
- A Neural Probabilistic Language Model2003in the sky ✦
- No Free Lunch Theorems for Optimization1997in the sky ✦
- Variational Inference with Normalizing Flows2015in the sky ✦
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models2024in the sky ✦
- OPT: Open Pre-Trained Transformer Language Models2022in the sky ✦
- Optimal Brain Damage1989in the sky ✦
- PaLM: Scaling Language Modeling with Pathways2022in the sky ✦
- A Time Series Is Worth 64 Words: Long-Term Forecasting with Transformers2023in the sky ✦
- Perceiver: General Perception with Iterative Attention2021in the sky ✦
- The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain1958in the sky ✦
- Perceptrons: An Introduction to Computational Geometry1969in the sky ✦
- π₀: A Vision-Language-Action Flow Model for General Robot Control2024in the sky ✦
- PinSage: Graph Convolutional Neural Networks for Web-Scale Recommender Systems2018in the sky ✦
- Image-to-Image Translation with Conditional Adversarial Networks2017in the sky ✦
- Pixel Recurrent Neural Networks2016in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Proximal Policy Optimization Algorithms2017in the sky ✦
- Prioritized Experience Replay2015in the sky ✦
- Progressive Growing of GANs for Improved Quality, Stability, and Variation2018in the sky ✦
- QLoRA: Efficient Finetuning of Quantized LLMs2023in the sky ✦
- RAFT: Recurrent All-Pairs Field Transforms for Optical Flow2020in the sky ✦
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks2020in the sky ✦
- Rainbow: Combining Improvements in Deep Reinforcement Learning2018in the sky ✦
- Random Features for Large-Scale Kernel Machines2007in the sky ✦
- Random Forests2001in the sky ✦
- Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation2014in the sky ✦
- REALM: Retrieval-Augmented Language Model Pre-Training2020in the sky ✦
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2022in the sky ✦
- Representation Learning: A Review and New Perspectives2013in the sky ✦
- Focal Loss for Dense Object Detection2017in the sky ✦
- RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback2023in the sky ✦
- RMSProp: Divide the Gradient by a Running Average of Its Recent Magnitude2012in the sky ✦
- Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation2014in the sky ✦
- RoBERTa: A Robustly Optimized BERT Pretraining Approach2019in the sky ✦
- Unsupervised Representation Learning by Predicting Image Rotations2018in the sky ✦
- RT-1: Robotics Transformer for Real-World Control at Scale2022in the sky ✦
- Efficiently Modeling Long Sequences with Structured State Spaces2022in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Some Studies in Machine Learning Using the Game of Checkers1959in the sky ✦
- Scaling Laws for Neural Language Models2020in the sky ✦
- Generative Modeling by Estimating Gradients of the Data Distribution2019in the sky ✦
- SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers2021in the sky ✦
- Self-Consistency Improves Chain of Thought Reasoning in Language Models2022in the sky ✦
- Self-Instruct: Aligning Language Models with Self-Generated Instructions2022in the sky ✦
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection2023in the sky ✦
- Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks2019in the sky ✦
- A Stochastic Approximation Method1951in the sky ✦
- SGDR: Stochastic Gradient Descent with Warm Restarts2017in the sky ✦
- Shampoo: Preconditioned Stochastic Tensor Optimization2018in the sky ✦
- Sharpness-Aware Minimization for Efficiently Improving Generalization2021in the sky ✦
- Show and Tell: A Neural Image Caption Generator2015in the sky ✦
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention2015in the sky ✦
- Signature Verification Using a "Siamese" Time Delay Neural Network1993in the sky ✦
- Distinctive Image Features from Scale-Invariant Keypoints2004in the sky ✦
- A Simple Framework for Contrastive Learning of Visual Representations2020in the sky ✦
- Video Generation Models as World Simulators2024in the sky ✦
- Sparks of Artificial General Intelligence: Early Experiments with GPT-42023in the sky ✦
- Emergence of Simple-Cell Receptive Field Properties by Learning a Sparse Code for Natural Images1996in the sky ✦
- Spatial Transformer Networks2015in the sky ✦
- SSD: Single Shot MultiBox Detector2016in the sky ✦
- A Style-Based Generator Architecture for Generative Adversarial Networks2019in the sky ✦
- Analyzing and Improving the Image Quality of StyleGAN2020in the sky ✦
- Support-Vector Networks1995in the sky ✦
- SwAV: Unsupervised Learning of Visual Features by Contrasting Cluster Assignments2020in the sky ✦
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?2024in the sky ✦
- Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows2021in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions2018in the sky ✦
- Temporal Difference Learning and TD-Gammon1995in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦
- A Decoder-Only Foundation Model for Time Series Forecasting2024in the sky ✦
- Toolformer: Language Models Can Teach Themselves to Use Tools2023in the sky ✦
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models2023in the sky ✦
- Trust Region Policy Optimization2015in the sky ✦
- Approximation by Superpositions of a Sigmoidal Function1989in the sky ✦
- Auto-Encoding Variational Bayes2013in the sky ✦
- On the Uniform Convergence of Relative Frequencies of Events to Their Probabilities1971in the sky ✦
- Very Deep Convolutional Networks for Large-Scale Image Recognition2014in the sky ✦
- Rapid Object Detection Using a Boosted Cascade of Simple Features2001in the sky ✦
- An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale2020in the sky ✦
- End-to-End Training of Deep Visuomotor Policies2016in the sky ✦
- Neural Discrete Representation Learning2017in the sky ✦
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations2020in the sky ✦
- WaveNet: A Generative Model for Raw Audio2016in the sky ✦
- The Strength of Weak Learnability1990in the sky ✦
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision2023in the sky ✦
- WebGPT: Browser-Assisted Question-Answering with Human Feedback2021in the sky ✦
- Wide & Deep Learning for Recommender Systems2016in the sky ✦
- Distributed Representations of Words and Phrases and Their Compositionality2013in the sky ✦
- World Models2018in the sky ✦
- Understanding the Difficulty of Training Deep Feedforward Neural Networks2010in the sky ✦
- XGBoost: A Scalable Tree Boosting System2016in the sky ✦
- You Only Look Once: Unified, Real-Time Object Detection2015in the sky ✦
- YOLOv3: An Incremental Improvement2018in the sky ✦
- Deep Neural Networks for YouTube Recommendations2016in the sky ✦
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models2019in the sky ✦