Context Window
نافذة السياق
العدد الأقصى من الرموز التي يستطيع النموذج معالجتها في مرة واحدة. FlashAttention مكّن نمو نوافذ السياق من 2 ألف إلى أكثر من 128 ألف رمز.
The maximum number of tokens a model can process at once. FlashAttention enabled context windows to grow from 2K to 128K+ tokens.
Also translated asالمدى السياقي المتاح، حجم نافذة المدخلات، طول السياق، نافذة السياق النصي، نطاق الاستيعاب، نطاق السياق
First appears in this corpus in: Distributed Representations of Words and Phrases and Their Compositionality (2013)
Appears in these papers
- Evaluating Large Language Models Trained on Code2021in the sky ✦
- Decision Transformer: Reinforcement Learning via Sequence Modeling2021in the sky ✦
- DeepSeek-V3 Technical Report2024in the sky ✦
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness2022in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- Generative Agents: Interactive Simulacra of Human Behavior2023in the sky ✦
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher2022in the sky ✦
- Language Models Are Unsupervised Multitask Learners2019in the sky ✦
- Language Models Are Few-Shot Learners2020in the sky ✦
- GPT-4 Technical Report2023in the sky ✦
- HuggingGPT: Solving AI Tasks with ChatGPT and Its Friends in Hugging Face2023in the sky ✦
- Jukebox: A Generative Model for Music2020in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- LLaMA: Open and Efficient Foundation Language Models2023in the sky ✦
- Mixtral of Experts2024in the sky ✦
- Mixtral of Experts2024in the sky ✦
- node2vec: Scalable Feature Learning for Networks2016in the sky ✦
- OPT: Open Pre-Trained Transformer Language Models2022in the sky ✦
- RETRO: Improving Language Models by Retrieving from Trillions of Tokens2022in the sky ✦
- RoFormer: Enhanced Transformer with Rotary Position Embedding2021in the sky ✦
- Scaling Laws for Neural Language Models2020in the sky ✦
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?2024in the sky ✦
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?2024in the sky ✦
- Distributed Representations of Words and Phrases and Their Compositionality2013in the sky ✦
- XLNet: Generalized Autoregressive Pretraining for Language Understanding2019in the sky ✦
- XLNet: Generalized Autoregressive Pretraining for Language Understanding2019in the sky ✦