News / #long-context Tag Long Context 500 articles archived under #long-context · RSS Sign in to follow arXiv — NLP / Computation & Language research 7h ago MoSAR: Mixture of Semantic Attention Regimes for Learning Adaptive and Approximable Attention Geometries arXiv:2609.31261v1 Announce Type: new Abstract: The quadratic complexity of dense self-attention remains a central bottleneck for long-context language modeling. Many efficient alternatives address this cost by deciding in advance where attention should be sparse or local. We… 13 arXiv — NLP / Computation & Language research 7h ago Highlight-Then-Summarize: Learning to Compress Evidence for Long-Context Understanding arXiv:2609.31382v1 Announce Type: new Abstract: Long-context understanding requires large language models (LLMs) to reason over lengthy documents, conversations, and code, yet task-relevant evidence is often sparse and scattered amid substantial irrelevant and redundant content.… 26 arXiv — NLP / Computation & Language research 7h ago UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification arXiv:2605.06221v2 Announce Type: replace Abstract: As large language models (LLMs) continue to advance rapidly, they are becoming increasingly capable while simultaneously demanding ever-longer context lengths. To improve the inference efficiency of long-context processing,… 30 r/LocalLLaMA community 16h ago Naive-N0.5-Flash - 309B-A15.5B https://naive.ai/en/research/ Built for coding and AI R&D 1M context Hybrid SWA/DSA   submitted by   /u/nullmove [link]   [comments] 35 arXiv — NLP / Computation & Language research 3d ago ChunkRank: Model-Aware Text Chunking and Abstention-Aware Answer Selection for LLM Pipelines arXiv:2609.29828v1 Announce Type: new Abstract: We present ChunkRank, an open-source Python library that derives chunk boundaries from a target model's tokenizer and context window, and selects an answer among candidates produced independently per chunk. It ships a validated… 38 r/LocalLLaMA community 4d ago model : add Ling 3.0 VL support by aetherbird · Pull Request #29151 · ggml-org/llama.cpp Model Overview Ling-3.0-flash-VL inherits the language, reasoning, and long-context capabilities of Ling-3.0-flash, while extending them with native image and video understanding. The model has 124B total parameters, with only 5.5B parameters activated per token, and supports a… 31 arXiv — NLP / Computation & Language research 4d ago When Learned Context Planning Fails to Beat Strong Retrieval: A Controlled Study of Planning, Routing, and Reranking for Long-Context QA arXiv:2609.26976v1 Announce Type: new Abstract: Learned context planning selects evidence atoms before an answer model reasons over them. We test whether this learned selection improves long-context multiple-choice QA after strong retrieval, routing, budgeted-selector, and… 27 arXiv — NLP / Computation & Language research 4d ago MWE-ECL: Recoverable Long-Range Context Does Not Always Override Local Lexical Priors arXiv:2609.27590v1 Announce Type: new Abstract: Long-context evaluations often test whether a model can recover distant evidence, but recoverability does not guarantee behavioral influence. We test the prediction that a distant discourse anchor can remain explicitly recoverable… 21 arXiv — NLP / Computation & Language research 4d ago DeltaS: Reading the Gated Linear Attention State for KV Cache Eviction in Streaming Video arXiv:2609.27470v1 Announce Type: cross Abstract: Recent video-language models increasingly adopt hybrid architectures that interleave linear and full attention layers for efficient long-context processing. While the recurrent state of linear attention remains fixed in size, the… 9 arXiv — Machine Learning research 5d ago CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference arXiv:2609.26300v1 Announce Type: new Abstract: Despite their strong performance, large language models (LLMs) are bottlenecked by KV cache memory traffic during long-context inference. Sparse attention is widely used to accelerate LLM inference by computing exact attention over… 10 arXiv — NLP / Computation & Language research 5d ago Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference arXiv:2609.25537v1 Announce Type: new Abstract: Large language model (LLM) inference is constrained by the quadratic scaling of self-attention and the linear scaling of the KV cache, increasing latency, energy consumption, and GPU memory demand as context length scales. Existing… 24 arXiv — NLP / Computation & Language research 5d ago HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing arXiv:2609.26368v1 Announce Type: new Abstract: Long-horizon and multi-turn agents typically generate short actions and process long observations from tools and environments. This growing context demands efficient prefill, compact KV-cache storage, and accurate long-context… 38 r/LocalLLaMA community 6d ago Is llama.cpp meant to be slow at long context, even when you aren't using that context? I am trying out a few fine-tunes of Qwen 3.5 9B @ IQ4_XS @ 131K context and trying to go mostly local (free beats cheap, after all). However, it is much slower than at, say 16K context, even when I am not actually using 131K tokens in the first place. Anyone know why this is? I… 30 arXiv — NLP / Computation & Language research 6d ago Context Poisoning as Extreme-Value Attention Interference in Long-Context Language Models arXiv:2609.22101v1 Announce Type: new Abstract: Large language models can process increasingly long prompts, yet their ability to locate and use decisive evidence may degrade as irrelevant or confusable context is added. We formulate this phenomenon, which we call context… 8 arXiv — NLP / Computation & Language research 6d ago Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models arXiv:2609.22452v1 Announce Type: new Abstract: Long-context understanding remains a fundamental challenge for large language models, as excessively long inputs often lead models to forget salient information. This issue is even more pronounced in the speech domain, where audio,… 33 arXiv — NLP / Computation & Language research 6d ago Block-Sparse Attention with Semantic-Geometric Decoupled Routing arXiv:2609.22884v1 Announce Type: new Abstract: Long-context inference has become a defining capability of large language models, but exact dense attention remains costly due to its quadratic scaling with sequence length. Block-sparse attention offers a hardware-friendly… 36 arXiv — Machine Learning research 7d ago Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding arXiv:2609.20888v1 Announce Type: new Abstract: Massive KV caches can cause severe memory-bandwidth bottlenecks during long-context decoding. Sparse attention methods mitigate this via selective loading, but that comes at a cost: rigid heuristics drop necessary context, leading… 38 arXiv — Machine Learning research 7d ago TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching arXiv:2609.21172v1 Announce Type: new Abstract: Large language models (LLMs) are moving onto mobile devices for increasingly diverse workloads over text, images, video, and audio. These applications often require long contexts, making the Key-Value (KV) cache a dominant memory… 28 Vercel — AI dev-tools 7d ago Grok 4.7 now available and 40% off on AI Gateway, fx, and eve Grok 4.7 from SpaceXAI is now available on AI Gateway and 40% off through September 27. The discount applies automatically when you call spacexai/grok-4.7 . Grok 4.7 has a 500K token context window and supports low, medium, high, and xhigh reasoning levels, giving you control… 31 r/LocalLLaMA community 8d ago Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0) Hi. I saw some feedback that halogen was degrading at context depth. So I fixed that. Served through the image, same machine, same session, same prompts, 0.11.10 vs 0.12.0: decode at 1,004,581 tokens of context: 27.3 to 38.3 tok/s (default speculative drafter) decode at 258,794:… 18 The Information — AI news-outlet 9d ago Communicate Technical Topics to a Non-Technical Audience with Google Gemini In a room full of IT professionals, words and phrases like “tokenization,” “retrieval-augmented generation,” and “context window” need no explanation. But when technology leaders present to their CEOs and other company leaders about topics like AI, they often find that they need… 10 arXiv — Machine Learning research 10d ago Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training arXiv:2609.19242v1 Announce Type: new Abstract: Block diffusion language models (BDLMs) combine autoregressive dependencies across blocks with parallel denoising within blocks, but long-context training is constrained by distributed attention communication and activation memory.… 30 arXiv — NLP / Computation & Language research 10d ago DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression arXiv:2609.19969v1 Announce Type: new Abstract: The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and… 36 arXiv — NLP / Computation & Language research 10d ago On-Demand Attention: Language Models Know When to Recall arXiv:2609.20734v1 Announce Type: new Abstract: Reasoning and agentic workloads increasingly demand efficient long-context inference. Yet full-attention decoding reads the growing history at every step, regardless of its benefit to the next prediction. We show that a pretrained… 30 Hugging Face Daily Papers research 11d ago Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training Abstract Training a Mixture-of-Experts (MoE) model at long context or large batch size fails as soon as any one component's peak allocation exceeds device memory, so the target is every peak at once, not the average footprint. Four are left unbounded by the parallelism plans in… 5 Hugging Face Daily Papers research 11d ago SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization Abstract Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context modeling. However, existing context extension approaches typically apply continued pretraining directly without modifying these layers, overlooking… 17 arXiv — Machine Learning research 11d ago Beyond Static RAG: An Adaptive, Tri-Metric Routing Framework for Efficient Long-Context Inference on Commodity GPUs arXiv:2609.17564v1 Announce Type: new Abstract: Deploying retrieval-augmented generation (RAG) on commodity GPUs such as the NVIDIA T4 (16 GB VRAM) exposes a practical failure mode we call the Compression Paradox: neural prompt compression can add key-value (KV) cache contention… 7 arXiv — NLP / Computation & Language research 11d ago Long-Context Demonstration Selection Using State Space Models arXiv:2609.17888v1 Announce Type: cross Abstract: We study the problem of demonstration selection, which involves selecting a subset of examples for prepending to a query to a language model. This problem is closely related to in-context learning and language model inference.… 17 arXiv — NLP / Computation & Language research 11d ago ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference arXiv:2609.17943v1 Announce Type: cross Abstract: Long-context LLM inference is bottlenecked by attention, whose repeated KV-cache reads make decoding memory-bound. Self-speculative decoding alleviates this by drafting tokens with sparse attention and verifying them with full… 8 r/LocalLLaMA community 11d ago Qwen 3.8 27B Running for 63 hours on a RTX 3090 to solve the Riemann hypothesis I let Qwen 3.8 27B 4bit quantized with 100K context window run autonomously for 63 hours (50 million+ tokens) to try to solve the RH. Of course it did not solve it, but the experiment still shows it's internal work, memory organization, strategies used and more. The interesting… 11 arXiv — NLP / Computation & Language research 12d ago Can LLMs Follow the Pulse of a Crisis? Evaluating Crisis Sentiment in Bangladesh's July Uprising arXiv:2609.16997v1 Announce Type: new Abstract: Crisis sentiment analysis is especially challenging for low-resource languages such as Bangla, where language, context, and public reaction shift rapidly. We introduce UNRESTSENT200K, a Bangla crisis sentiment dataset with… 8 arXiv — NLP / Computation & Language research 12d ago Where Should a Document Live: Context, Representations, or Parameters? arXiv:2609.17346v1 Announce Type: new Abstract: To answer questions outside of their pre-training data, large language models (LLMs) need access to new information, which can be presented in the context window as documents, encoded into the model's parameters, or injected as… 21 arXiv — NLP / Computation & Language research 12d ago VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs arXiv:2609.16722v1 Announce Type: cross Abstract: Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates context windows and incurs prohibitive costs. Current solutions predominantly… 28 r/LocalLLaMA community 12d ago Qwen3.8-27B-NVFP4 1M context. So far so good. I am a beginner, Took a while to get started, get everything right. This setup is native not container. Still not sure if I did this right, or if I can tune this more. Environment=HF_HUB_OFFLINE=1 Environment=VLLM_LOGGING_LEVEL=INFO Environment=VLLM_ALLOW_LONG_MAX_MODEL_LEN=1… 15 arXiv — Machine Learning research 13d ago Tabby: An Open Pretraining Recipe for Time Series Foundation Models arXiv:2609.13956v1 Announce Type: new Abstract: In this report, we release Tabby, a long context probabilistic time series foundation model, together with a complete and open recipe of how it was built. Tabby adopts an encoder-only patch Transformer architecture and concentrates… 35 arXiv — NLP / Computation & Language research 13d ago SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization arXiv:2609.14320v1 Announce Type: new Abstract: Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context modeling. However, existing context extension approaches typically apply continued pretraining directly without… 26 Hugging Face Daily Papers research 13d ago ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Abstract ZGCM-1 is a 7B open foundation model that combines internal reasoning with external tool use, trained via efficient architecture-system co-design, progressive long-context scaling, and autonomous agent workflows to achieve strong reasoning and efficiency. Generated by… 26 arXiv — NLP / Computation & Language research 14d ago Doc2FRC: Length-Consistent Document-Level Machine Translation via Fixed-Range Chunking arXiv:2609.12674v1 Announce Type: new Abstract: Advanced large language models (LLMs) with long context windows can substantially reduce input truncation in document-level machine translation (DocMT). However, direct Doc2Doc translation remains prone to n-gram repetition and… 31 arXiv — NLP / Computation & Language research 14d ago Residual Vector-based Reconstruction as Long-Context Recall Regardless of Context Window Size arXiv:2609.12686v1 Announce Type: cross Abstract: Large language models (LLMs) process long contexts, including long documents and lengthy conversations, but face token-level memory usage that increases proportionally to input length. Although model optimization and lossy prompt… 35 r/LocalLLaMA community 15d ago This draft model is OP on 16 GB cards for Qwen 3.8 27b https://huggingface.co/HermiHg/Qwen3.8-27B-DFlash2-Q2_K_S-MIX-GGUF I used this draft model with https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with the IQ3_XXS with 128k context and I saw it averaging about 60 tokens per second tg speed on the 16 GB RX 9070 XT. This… 16 r/LocalLLaMA community 16d ago Agnes-AI/Agnes-3.0-Flash 33B Multimodal, AA score: 36 I find this new model at HF: "Built for demanding work. A 262 144-token context window, adjustable reasoning effort, tool calling, and text, image and video understanding. Architecture Agnes-3.0-Flash is a hybrid-attention decoder: three of every four layers run a gated delta… 26 arXiv — NLP / Computation & Language research 17d ago Larger Context Window, Fewer Overcorrections: Optimizing Prompts and Batching for Minimal-Edit Grammatical Error Correction arXiv:2609.10810v1 Announce Type: new Abstract: Minimal-edit Grammatical Error Correction (GEC) is a challenging task for zero- and few-shot prompted Large Language Models (LLMs), which systematically overcorrect and degrade $F_{0.5}$ by rewriting well-formed spans. While… 38 r/LocalLLaMA community 17d ago GigaChat-3.5-Reasoning Hey y'all! We've released a new model in our lineup: GigaChat-3.5 Reasoning. It's a 432B-A28B MoE with Gated DeltaNet for long-context efficiency. We trained domain experts (code, math, general, etc.) with CISPO and then distilled them into a single model via on-policy… 18 arXiv — NLP / Computation & Language research 18d ago ConvMem: Convolutional Memory for Long-Context Reasoning arXiv:2609.10441v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have demonstrated impressive capabilities, they often struggle with extremely long contexts due to fixed context limits. To address this, sequential approaches like MemAgent extend the effective… 38 r/LocalLLaMA community 18d ago GLM 5.3 Flash Q4 @ 60tps / 550tps on M3 Ultra I have been a dwarfstar fan for awhile and I really liked glm 5.3 flash but needed it to be materially faster to feel good using it. In the screenshot you can see the outcome of using the model with a claude code harness at ~200k depth, with many tool calls and averaging over… 29 Hugging Face Daily Papers research 19d ago Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training Abstract A system for large-scale online draft co-training accelerates speculative decoding in RL post-training by extending context-parallel attention and adding cross-stage feature transport. Generated by thinkingmachines/Inkling-Small Speculative decoding accelerates rollout… 36 r/LocalLLaMA community 19d ago Qwen3.8-Flash-Next on MLX-serve, 1m context is released! Hi, I'm the co-creator of this Qwen3.8-Flash-Next engine support in MLX-serve. I've been tuning this one to run both fast, efficient and correct up 1m context using kv cache 8 bits in M5 Max 128GB. Qwen is working well at very long context as showed in the video (a snapshot at… 17 r/LocalLLaMA community 19d ago inclusionAI/Ling-3.0-flash-VL · Hugging Face Ling-3.0-flash-VL inherits the language, reasoning, and long-context capabilities of Ling-3.0-flash, while extending them with native image and video understanding. The model has 124B total parameters, with only 5.5B parameters activated per token, and supports a context window… 27 arXiv — Machine Learning research 21d ago KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU arXiv:2609.04852v1 Announce Type: new Abstract: Modern LLM agents operate in persistent workspaces whose accumulated history can exceed both GPU KV capacity and the model's native context window. Existing systems typically compact older context into summaries or retrieve it… 34 r/LocalLLaMA community 21d ago LayerStoRm open-source expert streaming: 1M context GLM-5.3-Flash [UD-Q4_K_XL] at 24.5 tok/s @8k on just 2× RTX 5090 + 2× RTX 5080 (186 GiB MoE on 96 GB VRAM) LayerStoRm: GLM-5.3-Flash UD-Q4_K_XL (186 GiB) at 1M context on 2× RTX 5090 + 2× RTX 5080 (96 GB VRAM total) using RAM for the pinned experts. LayerStoRm is a (still experimental) MIT-licensed continuous expert-streaming inference engine: it runs MoE models far larger than your… 33 Page 1 of 10 · 500 articles Older →