News / #long-context Tag Long Context 500 articles archived under #long-context · RSS Sign in to follow r/LocalLLaMA community 1mo ago Tesla V100 Qwen3.6 27B Performance Looking for V100 users to share your config and it's performance. GPU: Tesla V100 PCIE 32Gb Qwen3.6 27B Q4_K_M + Q8_0 MTP 128K context length Pi coding agent llama.cpp model preset: [*] spec-default = 1 ctx-size = 131072 mmap = 1 kv-unified = 1 n-gpu-layers = 999 threads = 18… 14 arXiv — NLP / Computation & Language research 1mo ago QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding arXiv:2608.05326v1 Announce Type: cross Abstract: Autoregressive large language model inference is increasingly constrained by the memory footprint of the Key-Value (KV) cache. A dominant line of work reduces this footprint by evicting tokens that appear unimportant under… 28 arXiv — Machine Learning research 1mo ago Is Self-Pretraining really useful to improve diagnosis in medical Time Series? arXiv:2608.06122v1 Announce Type: new Abstract: Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal, multivariate, and even simple univariate medical… 23 r/LocalLLaMA community 1mo ago Final optimization: from ~10 tok/s to ~15 tok/s on DeepSeek-V4-Flash-0731 at 128K ctx - 1 RTX 3090 J'ai consacré beaucoup de temps à l'optimisation de DeepSeek-V4-Flash-0731 GGUF sur une seule RTX 3090. Mon exigence absolue pour chaque configuration était la suivante : Le modèle doit rester utilisable avec une fenêtre de contexte de 128 000 jetons. J'ai testé les différentes… 4 arXiv — Machine Learning research 1mo ago Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms arXiv:2608.04074v1 Announce Type: new Abstract: Long-context LLM decoding reads the key-value (KV) cache at every step. Loading it takes longer than computing attention over it, so throughput is bandwidth-bound. Hence, reducing the cache size can raise both decoding speed and… 13 arXiv — NLP / Computation & Language research 1mo ago Training-Free Hashing-Based Attention via Binary Principal Components arXiv:2608.04405v1 Announce Type: cross Abstract: Long-context large language models (LLMs) are increasingly deployed in real-world applications, yet self-attention remains a major efficiency bottleneck -- especially during decoding -- due to the necessity of repeatedly… 24 arXiv — NLP / Computation & Language research 1mo ago Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression arXiv:2608.04569v1 Announce Type: new Abstract: Hard prompt compression reduces long-context inference cost by independently scoring tokens, sentences, or chunks and retaining the highest-scoring units under a budget. We identify a structural failure in this procedure:… 17 arXiv — NLP / Computation & Language research 1mo ago Chained Recursive Language Models for Multi-Iteration Reasoning arXiv:2608.05124v1 Announce Type: new Abstract: Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference trajectory has to simultaneously explore the context, store intermediate state, verify evidence, and produce the… 20 Vercel — AI dev-tools 1mo ago Ling 3.0 Tiny is now available on AI Gateway Ling 3.0 Tiny from ANT Group is now on AI Gateway, free to use till 8:00am PT on 8/14. Ling 3.0 Tiny takes the free slot from Ling 3.0 Flash . Ling 3.0 Tiny is a MOE model with 7.9B total parameters and about 1.3B active per token, a 256K token context window, and up to 32K… 6 r/LocalLLaMA community 1mo ago LFM2.5-2.6B on a OnePlus 13 at 17 tok/s ~ Pure CPU As you all know the model is 2.69B parameters with a 128K context window and purpose-built for multi-step agent workflows. What you are seeing is the Q4_K_M GGUF running on my own inference engine built from scratch. The TUI is my own device probe suite running through ADB… 27 Hugging Face Daily Papers research 1mo ago Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements Abstract Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding numerical precision and multi-step logic over long contexts, is an ideal testbed. Existing benchmarks fail to capture… 10 arXiv — Machine Learning research 1mo ago Output-Aware Rotation for INT2 KV-Cache Quantization arXiv:2608.02691v1 Announce Type: new Abstract: The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-low-bit quantization increasingly important. However, existing rotation-based INT2 methods… 23 arXiv — NLP / Computation & Language research 1mo ago AnchorKV: Anchor-Residual KV Cache Compression arXiv:2608.02901v1 Announce Type: cross Abstract: The key-value (KV) cache is the primary memory bottleneck in long-context LLM inference. Existing approaches attack it from opposite ends: eviction methods permanently discard tokens, degrading performance whenever a discarded… 37 arXiv — Machine Learning research 1mo ago SAKI: Score-Aware Low-Rank Key Indexing for Long-Context KV Retrieval arXiv:2608.03228v1 Announce Type: new Abstract: Existing low rank KV cache methods preserve either model weights or key variance, neither of which directly reflects the attention scores used during inference. We derive the expected attention score distortion caused by rank r key… 11 arXiv — Machine Learning research 1mo ago TimeRLM: Recursive Language Models Enable Precise Anomaly Localization in Long-Context Time-Series arXiv:2608.03391v1 Announce Type: new Abstract: Precise anomaly localization over long-context time series is a crucial task in monitoring applications across clinical care, industrial operations, financial services, and logistics, where brief evidence may hide inside long spans… 31 arXiv — NLP / Computation & Language research 1mo ago PI-Mem: Pushing Long-Context Reasoning to 3.6M Tokens with Parallel-Iterative Memory arXiv:2608.03048v1 Announce Type: new Abstract: Long-context reasoning remains a critical bottleneck for large language models, as recent recurrent-memory approaches face two inherent challenges: sequential chunk-wise updates can overwrite early critical evidence with later… 15 r/LocalLLaMA community 1mo ago A 2.6B model with tool calling and 128K context now runs at 30 tok/s on a phone Liquid AI released LFM2.5-2.6B today, and this might be more relevant to local AI than another massive model most people cannot run. The model is only 2.69B parameters, has 128K context, supports tool calling and was post-trained specifically for multi-step agent workflows. The… 23 llama.cpp releases dev-tools 1mo ago b10273 sampler : remove "full-context windows" from history-based samplers ( #26524 ) Resolve -1 to 1024 instead of ctx-len for samplers Because of backend-sampling we initialize samplers before the complete llama_context is there. Therefore, we cannot infer the resolved context length… 36 r/LocalLLaMA community 1mo ago [Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding] First of all, obviously I took some help from AI to type this post and this is the topic that enabled me to accomplish all that: https://old.reddit.com/r/LocalLLaMA/comments/1veow4b/deepseek_v4flash_284b_moe_at_33_toks_single_68/ This post of mine is based on the link above. My… 11 arXiv — NLP / Computation & Language research 1mo ago AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI Agents arXiv:2608.00009v1 Announce Type: new Abstract: Long-term memory remains a critical bottleneck for conversational AI agents, whose finite context windows cannot support coherent recall across thousands of turns. We present AgentMemBench, a unified, reproducible benchmark… 32 arXiv — NLP / Computation & Language research 1mo ago SeDeM: Selective Decompression of Hidden-State Memories for Long-Context Question Answering arXiv:2608.00311v1 Announce Type: new Abstract: Long-context inference with large language models (LLMs) is costly: self-attention during prefill scales quadratically with sequence length, and the key-value (KV) cache grows with the number of processed tokens. Larger context… 15 arXiv — NLP / Computation & Language research 1mo ago S$^4$R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching arXiv:2608.00528v1 Announce Type: new Abstract: The growth of context window lengths in Large Language Models (LLMs) significantly enhances their long-context capabilities but incurs prohibitive memory costs due to the Key-Value (KV) cache. Although low-rank compression of KV… 7 arXiv — NLP / Computation & Language research 1mo ago LongChart VQA: A Comprehensive Benchmark for MLLMs with Complex Multi-Chart Reasoning arXiv:2608.01328v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are rapidly evolving with expanded context windows and stronger reasoning capabilities, enabling multi-chart understanding and multi-step inference. These abilities are increasingly… 10 arXiv — NLP / Computation & Language research 1mo ago Learning What to Remember: Test-Time Training via Context Distillation arXiv:2608.01672v1 Announce Type: new Abstract: Effective long-context modeling is not merely about retaining more of the past, but about preserving the information that may prove relevant later. Test-time training (TTT) is an appealing approach that performs online parameter… 28 arXiv — NLP / Computation & Language research 1mo ago Understanding Sparse Attention Selectivity in Long-Context Foundation Models via Counterfactual Evaluation arXiv:2608.01676v1 Announce Type: new Abstract: Sparse attention is widely deployed in long-context serving stacks, yet no framework audits how discarding blocks changes the influence of specific content on model output. We first establish that the phenomenon is real and causal:… 18 arXiv — NLP / Computation & Language research 1mo ago ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning arXiv:2607.28642v1 Announce Type: cross Abstract: Long chain-of-thought reasoning improves performance on complex problems, but it also introduces redundancy accumulation, context overflow, and error anchoring. We argue that under bounded context windows, the core bottleneck is… 31 arXiv — NLP / Computation & Language research 1mo ago Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements arXiv:2607.28661v1 Announce Type: new Abstract: Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding numerical precision and multi-step logic over long contexts, is an ideal… 30 arXiv — NLP / Computation & Language research 1mo ago ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression arXiv:2607.29591v1 Announce Type: new Abstract: KV cache compression is essential for efficient long-context inference. Existing eviction methods permanently discard unselected tokens and consequently remove their aggregate contribution to attention. Merging-based alternatives… 10 Vercel — AI dev-tools 1mo ago Qwen 3.8 Max now available on Vercel AI Gateway Qwen 3.8 Max is now available on AI Gateway. Qwen 3.8 Max handles text-only and vision-language work in one model, with 2.4 trillion parameters and a context window of up to 1 million tokens. The model is suited for software engineering and office productivity, along with visual… 13 r/LocalLLaMA community 1mo ago LongCat-Flash-Lite-Sparse Is Now Available for Download The weights have now been added to the repo an hour ago. This model is built upon LongCat-Flash-Lite , the differences are that LongCat-Flash-Lite-Sparse : Replaces dense MLA with LongCat Sparse Attention (LSA) Natively supports context lengths of up to 1M tokens (vs 256k for… 13 r/LocalLLaMA community 1mo ago What speeds are everyone getting with deepseek v4 flash 0731? What speeds are everyone getting with deepseek v4 flash 0731? I’m getting~200 tps prompt processing / ~11 tps token gen, on 4x5060ti16gb with ddr4 3200 ram at 4-channel, via llamacpp, with context window of 128000, -ub/-b at 4096, “q8” unsloth’s lossless quant   submitted by… 36 NVIDIA Developer Blog official-blog 1mo ago Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference As agentic and long-context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1). Because... 20 llama.cpp releases dev-tools 1mo ago b10201 ggml-webgpu: improve flash_attn_vec for quantized KV at long contexts ( #25956 ) improve fa of quantized kv cache Fix some bugs and some comments. fix v type check and some comments Fix build error caused by rebasing editorconfig checking pass Website: https://llama.app… 13 arXiv — Machine Learning research 1mo ago Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding arXiv:2607.27269v1 Announce Type: new Abstract: Multi-head latent attention (MLA) is increasingly important for long-context LLM inference because compact latent states replace the growing key-value (KV) cache and reduce decoding memory traffic. Yet most capable open checkpoints… 16 arXiv — NLP / Computation & Language research 1mo ago Recall Before You Rank: Similarity-Guided Top-$K$ Reuse for Efficient Long-Context Attention arXiv:2607.27692v1 Announce Type: new Abstract: Top-$K$ sparse attention reduces the cost of Softmax and value aggregation by attending to only a small subset of key--value (KV) entries. However, identifying this subset still requires scoring the current query against the full… 13 arXiv — NLP / Computation & Language research 1mo ago PCAP-LM: An LLM-Native Text Representation for TLS Bulk Traffic Analysis arXiv:2607.28100v1 Announce Type: cross Abstract: Large language models (LLMs) offer powerful reasoning capabilities for network traffic analysis, but standard capture formats and their textual equivalents are prohibitively verbose, overflowing LLM context windows by two orders… 37 r/LocalLLaMA community 1mo ago Inkling-Small by thinkingmachines 276B total parameters, 12B active, 1M context window. Blog post: https://thinkingmachines.ai/news/inkling-small/ NVFP4: https://huggingface.co/thinkingmachines/Inkling-Small-NVFP4 GGUF's by Unsloth: https://huggingface.co/unsloth/Inkling-Small-GGUF   submitted by  … 20 arXiv — NLP / Computation & Language research 2mo ago Mergeable Model-Side Aggregation States for Long-Context Language Models arXiv:2607.26448v1 Announce Type: new Abstract: A known limitation of long-context language models is their increasingly unreliable performance in non-additive, set-based aggregation as context length grows. Examples include cardinality estimation, set relationships, and grouped… 22 arXiv — NLP / Computation & Language research 2mo ago DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search arXiv:2607.27178v1 Announce Type: new Abstract: State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to… 21 arXiv — NLP / Computation & Language research 2mo ago MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent arXiv:2507.02259v2 Announce Type: replace Abstract: Despite improvements by length extrapolation, efficient attention and memory modules, handling infinitely long documents with linear complexity without performance degradation during extrapolation remains the ultimate challenge… 22 Vercel — AI dev-tools 2mo ago AI Gateway: GPT-5.6 pricing and speed updates On AI Gateway , GPT-5.6 Luna and GPT-5.6 Terra are now cheaper and GPT-5.6 Sol is faster. AI Gateway adds no markup on token pricing, so these changes reach you at the upstream rate. The changes apply to both short and long context pricing. Model Change Input: Short context (per… 4 arXiv — NLP / Computation & Language research 2mo ago CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention arXiv:2607.25291v1 Announce Type: new Abstract: The quadratic cost of self-attention makes long-context inference prohibitively expensive, and proxy-based block-sparse attention has become a practical remedy. Existing methods typically rely on a proxy to predict a binary sparse… 32 arXiv — NLP / Computation & Language research 2mo ago GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference arXiv:2607.24788v1 Announce Type: cross Abstract: As Large Language Models scale to increasingly long contexts, the memory I/O and computational overhead of the Key-Value (KV) cache during decoding emerges as the primary throughput bottleneck. To address this, we propose GLIDE,… 36 arXiv — NLP / Computation & Language research 2mo ago Addressable Recall Compaction for Long Context-Window Control in AI Agents arXiv:2607.25066v1 Announce Type: cross Abstract: Long-horizon LLM agents accumulate reasoning traces, actions, and tool observations that can eventually exceed a model's fixed context window. Existing compaction methods address this limitation by discarding, summarizing, or… 30 arXiv — NLP / Computation & Language research 2mo ago HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following arXiv:2607.25398v1 Announce Type: cross Abstract: Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows. Existing… 6 Hugging Face official-blog 2mo ago LFM2.5-Encoders for Fast Long-Context Inference on CPU Back to Articles a]:hidden"> LFM2.5-Encoders for Fast Long-Context Inference on CPU Team Article Published July 28, 2026 Upvote 14 Fernando Fernandes Neto fernandofernandes LiquidAI Edoardo Mosca EdoardoMosca LiquidAI Maxime Labonne mlabonne LiquidAI Leonie Monigatti iamleonie… 18 arXiv — Machine Learning research 2mo ago Variational-Ising-Attention (VIA):TailoredAttentionMattersfor Science arXiv:2607.23634v1 Announce Type: new Abstract: Attention enables context modeling via query-key scoring with softmax normalization. Driven by industrial long-context demands, mainstream research has converged toward sparsity and efficiency--yet softmax's independence assumption… 36 arXiv — NLP / Computation & Language research 2mo ago INS-ActBench: A Comprehensive Benchmark for Assessing Professional Actuarial Capability of Large Language Models arXiv:2607.24273v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong potential in financial reasoning, but existing benchmarks often evaluate domain knowledge, numerical reasoning, long-context understanding, and tool use in separate settings. This… 32 arXiv — NLP / Computation & Language research 2mo ago Kimi K3: Open Frontier Intelligence arXiv:2607.24653v1 Announce Type: new Abstract: We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention… 20 Hugging Face Daily Papers research 2mo ago Kimi K3: Open Frontier Intelligence Abstract We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow… 7 Page 4 of 10 · 500 articles ← Newer Older →