News / #long-context Tag Long Context 500 articles archived under #long-context · RSS Sign in to follow r/LocalLLaMA community 2mo ago First evidence of a pending qwen3.7 open weights release. Qwen3.7-flash is on open router. They referred to Qwen3.6-35b-a3b as Qwen3.6 flash so this is likely a small MoE. The prices are substantially cheaper than 3.6 flash with a native 1M context window.   submitted by   /u/fulgencio_batista [link]   [comments] 37 arXiv — NLP / Computation & Language research 2mo ago Learning What Matters: Supervising Sparse Attention Routing with Causal Evidence Sets arXiv:2607.21692v1 Announce Type: cross Abstract: Sparse attention reduces the cost of long contexts by allowing each query to read only selected parts of the input. These selectors are often trained by distilling the attention patterns of a dense teacher, assuming that… 31 arXiv — Machine Learning research 2mo ago RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention arXiv:2607.21927v1 Announce Type: new Abstract: Full self-attention in large language models scales as O(N^2), which limits long-context document analysis to 65,536 tokens and requires costly GPU clusters. The Reduced Interaction Sampling (RIS) inference engine addresses this… 31 r/LocalLLaMA community 2mo ago Mobile Offline LLMs: What do you use them for? I've spent the last year or so playing around with open source MLX and GGUF models on iPhone hardware. Given the limitations in memory, GPU/CPU/ANE, and in turn the context window I've been trying to figure out the best use cases for them. I've also done a lot of testing with… 13 r/LocalLLaMA community 2mo ago Kimi Linear 48B A3B? Just noticed this exists, 1M context MOE with 48B par seems just like what Ive been looking for - it runs pretty damn fast too compared to Qwen 3.6 35B. after some testing it seems capable of producing *not terrible* results but it always tries to go for the minimun possible… 24 r/LocalLLaMA community 2mo ago My GX10 died Everything ran fine, I was using UD 3.6 Q6 for 35 and 27B, each 4 concurrent requests at 200K context. I had Dify and Mastra to play around with, Unsloth studio to get around to and vLLM ready for whenever I decided to do some more testing. LLama-swap above lama.cpp and liteLLM… 30 r/LocalLLaMA community 2mo ago DKV: Open-source KV-cache compression framework for local LLM inference (CLI + technical report) Hi everyone! Over the past five months I've been working on DKV (DifferentialKV), an open-source project exploring KV-cache compression for long-context local LLM inference. The goal is to reduce KV-cache memory requirements through anchor-based representations, joint low-rank… 21 arXiv — Machine Learning research 2mo ago Codec-Gauge: Learning Compression-Friendly Gauges for Transformer KV Caches arXiv:2607.20538v1 Announce Type: new Abstract: Long-context Transformer inference increasingly relies on KV-cache compression or quantization. Prior rotation and transform-coding results suggest that the channel basis of each key/value vector affects how faithfully a fixed… 14 arXiv — NLP / Computation & Language research 2mo ago news-crawler-LM: A Small Long-Context Model For High-Quality News Crawling arXiv:2607.21284v1 Announce Type: new Abstract: Extracting structured content from news pages remains challenging due to heterogeneous HTML layouts, inconsistent markup, and substantial boilerplate such as navigation elements and advertisements. Rule-based news crawlers can… 13 Vercel — AI dev-tools 2mo ago Ling 3.0 Flash is now available on AI Gateway Ling 3.0 Flash from Ant Group is now available on AI Gateway. The model is free to use for the next three weeks, through August 3rd. Ling 3.0 Flash is a Mixture-of-Experts model with 124B total parameters and about 5.1B active per token. It has a 256K token context window and… 12 r/LocalLLaMA community 2mo ago Tokenizer Expansion: Upgrading a Model's Tokenizer in Place - LFM2.5-8B-A1B Today, we're sharing the recipe behind the new tokenizer in LFM2.5-8B-A1B . It upgrades a pre-trained model's tokenizer in place , without retraining from scratch. We doubled the vocabulary from 65K to 128K to fix the languages our original tokenizer split too finely. Blog:… 34 arXiv — NLP / Computation & Language research 2mo ago Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning arXiv:2607.19345v1 Announce Type: new Abstract: Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to long-context settings has emerged as an important frontier. However, we identify a critical… 16 arXiv — Machine Learning research 2mo ago High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration arXiv:2607.16248v1 Announce Type: new Abstract: Long-context large language model inference relies on the KV cache to avoid redundant attention computation, but incurs high memory and bandwidth overheads. Low-bit KV-cache quantization reduces this cost, yet it severely degrade… 17 arXiv — Machine Learning research 2mo ago More Than Memory: Task-Conditioned Signed FFN Writes in Long-Context Retrieval arXiv:2607.16254v1 Announce Type: new Abstract: FFNs are often treated as parametric memories. In long-context retrieval, however, the sharper question is not only what they store, but whether their native residual writes push the current retrieval state toward or away from the… 10 arXiv — NLP / Computation & Language research 2mo ago C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference arXiv:2607.17715v1 Announce Type: new Abstract: Long-context inference is central to modern large language model (LLM) applications such as retrieval-augmented generation and multi-document reasoning. To mitigate the growing inference cost, recent work has explored key-value… 20 arXiv — NLP / Computation & Language research 2mo ago SWE-Pruner Pro: The Coder LLM Already Knows What to Prune arXiv:2607.18213v1 Announce Type: new Abstract: Pruning long context for coding agents has been a vital technology for efficient context management. While existing context pruning methods such as SWE-Pruner realize this by attaching a separate code classifier, we find the agent… 8 arXiv — NLP / Computation & Language research 2mo ago Is Progressive Disclosure All You Need for Long-Context Agents? arXiv:2607.17598v1 Announce Type: cross Abstract: Long-document question answering usually forces a choice between loading the whole document into the context window and bolting on a separate retriever. Agentic AI suggests a broader option, giving the agent the document path and… 30 Hugging Face Daily Papers research 2mo ago SWE-Pruner Pro: The Coder LLM Already Knows What to Prune Abstract Pruning long context for coding agents has been a vital technology for efficient context management. While existing context pruning methods such as SWE-Pruner realize this by attaching a separate code classifier, we find the agent itself encodes internal representations… 22 Vercel — AI dev-tools 2mo ago Laguna S 2.1 is now available on AI Gateway Laguna S 2.1 from Poolside is now available on AI Gateway. There are 2 versions of the model available: Free version (256K context window): poolside/laguna-s-2.1-free Paid version (1M context window): poolside/laguna-s-2.1 Laguna S 2.1 is an open-weight Mixture-of-Experts model… 17 arXiv — NLP / Computation & Language research 2mo ago VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs arXiv:2607.15498v1 Announce Type: new Abstract: The key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) inference. Two leading training-free families are both structurally limited: token-selection methods (SnapKV, Ada-KV) score… 31 r/LocalLLaMA community 2mo ago Fractale-350M-base: memory as trained behaviour instead of long context, a fully open research release Some of you may remember my post about the research project behind this: a trained fast-weight memory, with the paper and the full research log at github.com/kkuette/thought-bank. This is the follow-up. The first public model of the series is out. Quick context: solo researcher,… 11 Hugging Face Daily Papers research 2mo ago LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Abstract A growing gap separates inference context lengths from RL post-training: inference systems are approaching million-token contexts, while post-training workloads often remain at 256K tokens or below and rely on length generalization at deployment. The gap is especially… 16 arXiv — Machine Learning research 2mo ago LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget arXiv:2607.14952v1 Announce Type: new Abstract: A growing gap separates inference context lengths from RL post-training: inference systems are approaching million-token contexts, while post-training workloads often remain at 256K tokens or below and rely on length generalization… 14 arXiv — NLP / Computation & Language research 2mo ago PReM: Learning What to Preserve and When to Refresh for Context Compression arXiv:2607.14327v1 Announce Type: new Abstract: Efficient long-context inference is not only about reducing memory cost, but also about keeping useful contextual evidence accessible as generation proceeds. However, existing compression-oriented approaches, such as key-value (KV)… 32 r/LocalLLaMA community 2mo ago Anyone else completely tuning out these massive "open weight" drops? Tbh the benchmarks on stuff like GLM-5.2 look insane. 753B params, 1M context, MIT license... everyone is throwing a party on the front page right now. But like... what is actually "local" about this anymore? A 700B+ MoE isn't fitting on anyone's home rig. Even if you absolutely… 29 Smol AI News news-outlet 2mo ago not much happened today **Moonshot AI** launched **Kimi K3**, a frontier-class open-weights model with **2.8T parameters**, **1M-token context window**, and **native multimodal input**. It features novel **Kimi Delta Attention (KDA)** enabling up to **6.3x faster decoding** and **Attention Residuals**… 18 r/LocalLLaMA community 2mo ago Qwen 3.6 27B is solid up to 262K context. How high have you guys gone above that using Rope/Yarn scaling? Stack: i7 12700K | RTX 3090 TI | 96GB RAM Qwen 3.6 27B Q3/Q5 KXL UD I've been pushing Qwen 3.6 27B above 200K ctx all week, and it handles it like a champ. I'm impressed. Today I hit the ceiling at 262K and it's still functioning well and coherent. I'm planning on trying out… 22 arXiv — Machine Learning research 2mo ago Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference arXiv:2607.13205v1 Announce Type: cross Abstract: Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of a long-context model by ranking tokens on accumulated attention mass, treated here as signal energy, and keeping the heaviest.… 11 Vercel — AI dev-tools 2mo ago Kimi K3 is now available on AI Gateway Kimi K3 from Moonshot AI is now available on AI Gateway. K3 is an open-source model with a 1M-token context window and native visual understanding, accepting text, image, and video inputs. Built for long-horizon software engineering, knowledge work, and deep reasoning, K3 is… 35 Hugging Face Daily Papers research 2mo ago SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Abstract Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc. However, real-world documents combine multiple factors such as length, layout complexity, modality, and question… 23 arXiv — Machine Learning research 2mo ago Semidirect Fourier Delta Attention: Phase-Controlled Delta Memory with Constructive Chunk-WY Kernels arXiv:2607.11897v1 Announce Type: new Abstract: Linear attention replaces softmax attention's growing KV cache with a fixed recurrent state, but this compression limits exact state tracking and long-context memory. We introduce \emph{Semidirect Fourier Delta Attention} (SFDA), a… 38 arXiv — Machine Learning research 2mo ago LiteTopK: Exploiting the Curse of Dimensionality for a Fused Indexer-TopK Kernel in Long-Context Sparse Attention arXiv:2607.11976v1 Announce Type: new Abstract: Indexer-TopK, the operation to compute the scores and select the top-k candidates, is widely used by sparse attention kernels in large language models and vector retrieval in recommendation systems and vector databases. However,… 34 arXiv — NLP / Computation & Language research 2mo ago A JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs arXiv:2607.12550v1 Announce Type: cross Abstract: The key-value (KV) cache has become the dominant memory cost of transformer inference. It grows with batch size, context length, and depth, and at long context it, rather than the model weights, sets the ceiling on throughput.… 31 arXiv — Machine Learning research 2mo ago What Context Does a Coding Agent Actually Need to Act? arXiv:2607.09691v1 Announce Type: new Abstract: A modern coding agent can hold an entire repository in its context window. Most of its reading is wasted -- and the interesting question is not how much context an agent can use, but what it actually \emph{needs}. We study that… 38 arXiv — NLP / Computation & Language research 2mo ago GEIS: A Generation-Evaluation-Improvement Loop of Agent Skills for Long-Form Article Generation arXiv:2607.11503v1 Announce Type: new Abstract: Long-form article generation remains difficult for large language models because it combines long context, long instructions, and long outputs. Existing multi-agent pipelines such as STORM improve information coverage by simulating… 19 Hugging Face Daily Papers research 2mo ago Self-Guided Test-Time Training for Long-Context LLMs Abstract Long-context processing has become increasingly important for large language models (LLMs), but simply extending the context window does not guarantee effective utilization of long inputs. As input length grows, accuracy often degrades, indicating that models still… 4 arXiv — NLP / Computation & Language research 2mo ago WILDTRACE: Benchmarking Natural Evidence Trails in Long-Context Reasoning arXiv:2607.09328v1 Announce Type: new Abstract: Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across distant passages. In an incident report, the operating condition, design flaw, and missed… 16 arXiv — NLP / Computation & Language research 2mo ago Self-Guided Test-Time Training for Long-Context LLMs arXiv:2607.09415v1 Announce Type: new Abstract: Long-context processing has become increasingly important for large language models (LLMs), but simply extending the context window does not guarantee effective utilization of long inputs. As input length grows, accuracy often… 31 arXiv — NLP / Computation & Language research 2mo ago Evaluating Retrieval-Augmented Generation vs. Long-Context Input for Clinical Reasoning over EHRs arXiv:2508.14817v2 Announce Type: replace Abstract: Objective: To evaluate whether retrieval-augmented generation (RAG) can serve as an efficient alternative to long-context prompting for clinical reasoning over electronic health records (EHRs). Methods: We defined three… 27 arXiv — NLP / Computation & Language research 2mo ago REAL: REtrieval-reAsoning and Logic-constructed Attention Behaviors for Long-Context KV Cache Compression arXiv:2508.15806v2 Announce Type: replace Abstract: The growing sequence length of large language models poses significant challenges for key-value (KV) caches. Existing state-of-the-art cache eviction methods primarily analyze the inference behavior of attention heads in… 4 r/LocalLLaMA community 2mo ago Running Qwen3.5-122B on Mac Studio 96GB: Fixed 3 bugs that made long-context inference usable Hey everyone, I recently switched from DS4 Flash to Qwen3.5-122B on my M3 Ultra Mac Studio for long-context agentic coding. While the model fit better, I hit a wall where follow-up messages took 3-5 minutes to start generating (cold fills) despite having a "warm" context. Turns… 22 r/LocalLLaMA community 2mo ago CTX: How far can you reasonably go with Qwen 3.6 27B? How far can i stretch the context window with Qwen 3.6 27B (using Q8_0) before it gets too unreliable? I am at 100k right now and i am not quite statisfied. Other than not quantizing KV cache, is there anything else that can be done to make the model more stable over longer CTX?… 18 r/LocalLLaMA community 2mo ago tencent/HiLS-Attention-7B · Hugging Face HiLS-Attention is a chunk-wise sparse attention mechanism that learns chunk selection end-to-end under the language-modeling loss, enabling native sparse training for efficient long-context modeling. This repository hosts the 7B checkpoint continued-trained on top of an… 7 Hugging Face Daily Papers research 2mo ago Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE Abstract A novel zero-shot method called Jet-Long enables efficient long-context processing for large language models by dynamically adapting rescaling factors and utilizing a bifocal attention mechanism that maintains high performance across varying sequence lengths. Generated… 36 r/LocalLLaMA community 2mo ago [Paper] Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity Linear attention models allow a fixed state size and a fixed amount of compute per token. However, due to their limited state size, linear attention models fall behind in long-context recall compared to softmax-attention-based transformer architectures. Increasing the state size… 9 arXiv — NLP / Computation & Language research 2mo ago Uncertainty-gated selection for block-sparse attention arXiv:2607.07724v1 Announce Type: cross Abstract: Block-sparse attention scales long-context language models by replacing the O(N^2) softmax with a per-query top-k selection over key blocks. This cutoff is myopic: when the k-th and (k+1)-th blocks are nearly tied in score, the… 7 arXiv — Machine Learning research 2mo ago Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE arXiv:2607.07740v1 Announce Type: new Abstract: Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces routinely push the input an order… 37 arXiv — Machine Learning research 2mo ago Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing arXiv:2607.07953v1 Announce Type: new Abstract: Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length limits training and inference at long context. This paper presents a comparative study of softmax attention and… 10 r/LocalLLaMA community 2mo ago Locally run assistant on a w-10 board on a local Xiaozhi server Hello, Locallama! "long-time" member here, from the days when the max context window doubled from 2K to a whooping 4K! Now I feel like I am living in a dream world, doing things I could not have imagined back then. This is Belochka, or Bella, running on a Mac and projected to an… 31 r/LocalLLaMA community 2mo ago Stripping terminal noise from agent context via a lazy-loaded local CLI layer. Looking for brutal feedback on this heuristic. When building long-running coding agents, terminal output is one of the fastest ways to poison a context window. If an agent runs an intensive build, an install command, or a massive search ( grep / find ), it easily generates hundreds of lines of raw log noise. The agent reads… 7 Page 5 of 10 · 500 articles ← Newer Older →