News / #long-context Tag Long Context 500 articles archived under #long-context · RSS Sign in to follow arXiv — NLP / Computation & Language research 1mo ago SCOPE: A Generative Approach for LLM Prompt Compression arXiv:2508.15813v2 Announce Type: replace Abstract: A big issue in modern LLM applications is they tend to feed long context to LLM, which results in high inference cost and latency, and may exceed the context limit. Prompt compression addresses this issue by reducing the length… 11 r/LocalLLaMA community 1mo ago Single RTX 5090: Qwen3.8-27B NVFP4 at a real 262K context in vLLM — 77 tok/s short-context, 64.7 tok/s at 128K This is the Qwen3.8-27B setup I actually use every day on one RTX 5090. I wanted to write it down with enough detail that another 5090 owner can reproduce it instead of guessing which memory knobs I used. The short version: the full 262,144-token window fits together with… 25 r/MachineLearning community 1mo ago I developed my own quantized LLM from scratch, trained on 30B tokens, deploys in 60 MB [R] I trained a 250M parameter model from scratch on 30B tokens of fineweb. It’s quantized to under 2 bits so the whole deployment is 60 MB and it needs about 80 MB of RAM to run. Runs around 400 tok/s on a normal laptop CPU, no GPU needed. How the long context works: the most… 10 Hugging Face Daily Papers research 1mo ago FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving Abstract FlashPrefill V2 improves long-context serving via mean-corrected sparse attention, optimized GPU operators, and framework integration, achieving large speedups over dense baselines. Generated by thinkingmachines/Inkling-Small Long-context modeling is a pivotal… 9 arXiv — Machine Learning research 1mo ago Inadvertent Context Leakage in Language Models arXiv:2608.19857v1 Announce Type: new Abstract: For AI agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. We study whether the mere presence of such secrets in a model's context window… 13 arXiv — Machine Learning research 1mo ago HYDRA: A Heterogeneous Chiplet DSE Framework for Serving Dynamic Hybrid LLM Workloads arXiv:2608.19395v1 Announce Type: cross Abstract: Hybrid Transformer-Mamba large language models (LLMs) enhance long-context efficiency, but their heterogeneous computation and communication patterns complicate efficient hardware acceleration. Chiplet-based architectures offer a… 7 arXiv — NLP / Computation & Language research 1mo ago FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving arXiv:2608.19758v1 Announce Type: new Abstract: Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. Our previous work,… 33 arXiv — NLP / Computation & Language research 1mo ago Learning how to Forget: Fine-tuning for Long-Context Sparse Attention arXiv:2608.19920v1 Announce Type: new Abstract: A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets. We provide a new method for… 18 arXiv — NLP / Computation & Language research 1mo ago Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning arXiv:2608.19181v1 Announce Type: cross Abstract: On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-context tasks, however, token-level teacher support can favor locally plausible responses that… 36 arXiv — NLP / Computation & Language research 1mo ago LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization arXiv:2608.18082v1 Announce Type: new Abstract: Although context windows have expanded significantly in recent years, hallucinations in long-context summarization remain a challenge. Long novels are better suited than news or papers for researching these hallucinations, due to… 25 arXiv — NLP / Computation & Language research 1mo ago Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale arXiv:2608.19026v1 Announce Type: new Abstract: Released in 2025, Institutional Books: Harvard Library (IB-HL) is a collection of 983,004 volumes (242B o200k_base tokens), originally digitized through Harvard Library's participation in the Google Books Library project. As… 17 Hugging Face Daily Papers research 1mo ago CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing Abstract A new dataset, benchmark, and 22B model enable compositional instruction-guided video editing with multi-region attention and temporal coherence. Generated by thinkingmachines/Inkling-Small The quality and diversity of instruction-based video editing datasets are… 11 arXiv — Machine Learning research 1mo ago Dynamic Compression in Recurrent Networks arXiv:2608.17896v1 Announce Type: new Abstract: Recurrent models process long contexts efficiently by compressing their history into a fixed-size state, but modern architectures typically do so in a single causal pass over the sequence. Each input must therefore be compressed… 13 arXiv — NLP / Computation & Language research 1mo ago Token Optimization and Context Window Management in Multi-Agent AI Workflows arXiv:2608.17188v1 Announce Type: new Abstract: Multi-agent AI workflows are limited not only by model quality but by token cost, latency, and context-window quality. This paper presents a practitioner framework for token optimization and context-window management, grounded in… 23 arXiv — NLP / Computation & Language research 1mo ago Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds arXiv:2608.17950v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate remarkable multi-hop reasoning capabilities over long contexts, yet the internal mechanisms enabling these distant cognitive leaps remain poorly understood. Traditional attention-based… 12 arXiv — NLP / Computation & Language research 1mo ago MoNe: Modular Neural Memory for Efficient Long Context Inference arXiv:2608.17616v1 Announce Type: cross Abstract: We present MoNe, a lightweight modular neural memory that attaches to any frozen pretrained Transformer to enable long-context inference without retraining. MoNe reads context in fixed-size segments via test-time learning of… 8 r/LocalLLaMA community 1mo ago Ling-3.0-tiny is a very interesting model. Run on NVIDIA Orin Nano Super 8GB at 128K context with IQ4_NL quant. I have been searching for suitable model to run on my 8GB RAM toy, NVIDIA Orin Nano Super 8GB. This little toy was priced at $249 earlier this year (not any more), and pulls very little power when idle. It was an interesting device that suitable for an agent to host on. It is… 31 r/LocalLLaMA community 1mo ago Running DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GB I managed to run the 143–144 GiB DeepSeek-V4-Flash-0731 UD-Q4_K_XL GGUF on four RTX 3060 12GB cards while keeping a 360k–376k context window. Hardware: CPU: Intel Core i9-10920X, 12C/24T RAM: 128 GB DDR4-3200, quad-channel GPU: 4× NVIDIA RTX 3060 12GB Total VRAM: 48 GB Storage:… 25 arXiv — NLP / Computation & Language research 1mo ago SEER: Long-Context Reasoning via Selective Visual-Text Compression arXiv:2608.15962v1 Announce Type: new Abstract: Long-context reasoning remains computationally expensive for large language models due to the quadratic complexity of attention over text tokens. Visual-text compression offers a promising alternative by rendering text into images… 12 Hugging Face Daily Papers research 1mo ago MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling Abstract MegaParts scales part-aware 3D generation via token-efficient vector-quantized part tokens and structured autoregressive sequence modeling with long-context training. Generated by thinkingmachines/Inkling-Small Part-aware 3D object generation is essential for graphics… 21 r/LocalLLaMA community 1mo ago Made this game in two prompts with Q4, Qwen 3.8 is amazing This took one prompt to build, and another follow up prompt to fix two issues (player got stuck with the bomb and broken enemies path-finding), this is only html, css and js, no external assets, all done by Qwen. Using UD-Q4_K_XL in llama.cpp with 128k context and k5_0/v4_1… 36 Hugging Face Daily Papers research 1mo ago SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning Abstract On-policy distillation from a long-context reasoning teacher to short-context students improves mathematical proof reasoning and generalizes to science benchmarks by aligning token spans, constraining length growth, and stabilizing training. Generated by… 26 arXiv — Machine Learning research 1mo ago The Query Knows What to Forget: A Second Erase Direction for Linear Attention arXiv:2608.13668v1 Announce Type: new Abstract: Linear attention keeps a state of fixed size. At long context, many stored items share this state, and interference between them degrades retrieval. Gated DeltaNet-2 (GDN-2), like every delta-rule model before it, derives its erase… 28 arXiv — NLP / Computation & Language research 1mo ago KV Cache Compression Through the Lens of Transform Coding arXiv:2608.14191v1 Announce Type: cross Abstract: The key-value (KV) cache stores information from past tokens and is a major memory bottleneck in long-context inference. Existing quantization methods address this bottleneck by representing the KV cache uniformly with… 29 arXiv — NLP / Computation & Language research 1mo ago SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning arXiv:2608.14277v1 Announce Type: new Abstract: On-policy distillation (OPD) offers a promising way to transfer reasoning capabilities from stronger teacher models, but applying it to long-context reasoning teachers and short-context students introduces practical challenges,… 11 r/MachineLearning community 1mo ago How can we solve long-range recall in linear attention? [D] Recently, I started working on DNA sequence modeling and decided to explore linear attention , mainly because DNA sequences can easily reach 1M tokens , making standard softmax attention extremely expensive in terms of memory and computation. The model performed reasonably well… 5 Hugging Face Daily Papers research 1mo ago Maglev: Sliding Recurrent Memory Abstract A recurrent Transformer with fixed-size memory and coupled prefiller-decoder training improves long-context modeling while enabling efficient parallel training and reduced inference cost. Generated by thinkingmachines/Inkling-Small We introduce , a recurrent Transformer… 38 r/LocalLLaMA community 1mo ago Qwen3.8-27B Q6_K at 128K on a single 32GB GPU Fresh download today. Really quick: my impression of Qwen3.8-27B: I ran the Q6_K GGUF locally on a 32GB R9700 through llama.cpp + OpenCode, with 128K context. Q8 would not fit with enough left over for context. I gave it my real, years-old swimming pool-controller repository and… 16 arXiv — Machine Learning research 1mo ago MARCH: Scaling Recurrent Memory with Content-Routed State Anchors arXiv:2608.12435v1 Announce Type: new Abstract: Transformers owe much of their strong long-context retrieval capability to a token-level memory that grows with context length. This flexibility, however, incurs a quadratic computation complexity during training and a key--value… 15 r/LocalLLaMA community 1mo ago Qwen 30b MoE - 30tps - 6GB vram - Done! So, I have been dreaming of getting 17 tokens per second using my RTX 3050 6GB version on a decent context window for Hermes needed above 60k. The hope is that has was a 22GB of DDR 4, hoping they can take some of those experts and give me room for context. What did I get 10 or… 33 r/LocalLLaMA community 1mo ago You could purchase a Desktop with 2TB of DDR5 - It only sets you back some $200k+ Just watched Wendell's (level1 techs) latest video on the HP Z8 Fury desktop workstation and was curious how you could configure it. And oh boy, there's an option for 2TB which costs some $211k just for the RAM alone. But the real interesting part with the latest price hikes for… 21 arXiv — Machine Learning research 1mo ago Disentangling the Expressivity of RoPE arXiv:2608.11909v1 Announce Type: new Abstract: Two accounts recur in explanations of the success of rotary position embeddings (RoPE). Expressivity studies associate periodic position information with modular predicates, whereas mechanistic and long-context studies emphasize… 16 arXiv — NLP / Computation & Language research 1mo ago Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction arXiv:2608.11242v1 Announce Type: new Abstract: When the context window is under pressure, LLM systems compact prior context to continue ongoing tasks. We identify a class of user-issued instructions, Session Constraints (SCs), such as "do not delete any emails until I confirm,"… 27 arXiv — NLP / Computation & Language research 1mo ago Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge arXiv:2608.12218v1 Announce Type: new Abstract: Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction histories. This scaling reflects the implicit assumption that training on longer contexts will… 32 Hugging Face Daily Papers research 1mo ago CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG Abstract CoinRAG improves retrieval-augmented generation efficiency and accuracy by reusing fine-grained semantic nugget caches instead of full chunks. Generated by thinkingmachines/Inkling-Small Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited… 4 Vercel — AI dev-tools 1mo ago GLM 5.2 free for eve agents through August 27 via Blackbox on AI Gateway GLM 5.2 , the open-weights coding model from Z.ai with a 1M-token context window, is free for eve agents through August 27, served by Blackbox AI on AI Gateway . New eve agents come with GLM 5.2 as their default model. Use npx eve@latest init my-agent to get started. Existing… 35 arXiv — NLP / Computation & Language research 1mo ago Position Encoding in Transformers: From Absolute and Relative Methods to Rotary Position Embeddings and Long-Context Scaling arXiv:2608.10021v1 Announce Type: new Abstract: Self-attention models content-dependent interactions between tokens but does not by itself encode token order. Position encoding addresses this limitation by introducing absolute coordinates, relative distances, or… 11 arXiv — NLP / Computation & Language research 1mo ago Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension arXiv:2608.10296v1 Announce Type: new Abstract: One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that this is not the case in the long context setting. Specifically, we show that a… 32 Vercel — AI dev-tools 1mo ago Grok 4.6 now available on AI Gateway Grok 4.6 from SpaceXAI is now available on AI Gateway . The model has a 500K token context window and accepts text and image inputs. Grok 4.6 supports low, medium, high, and xhigh reasoning levels and defaults to high. To use Grok 4.6, set model to xai/grok-4.6 in the AI SDK :… 29 r/LocalLLaMA community 1mo ago I ran Muse Glimmer @ 1M context - All tests passed. Heeeey all! I just completed some fun tests with Muse Glimmer, I thought I'd let you know. In fact, the summary below was written by Muse itself! I ran a 2× DGX Spark cluster and got Meta's day-old Muse Glimmer 30B running the day after release — then pushed its context from the… 36 arXiv — Machine Learning research 1mo ago RippleKV: Cross-Layer KV Cache Allocation via Perturbation Propagation arXiv:2608.08684v1 Announce Type: new Abstract: Long-context LLM inference is bottlenecked by KV cache memory, yet distributing a limited cache budget across layers remains challenging. Existing methods rely on proxies such as layer depth, attention statistics, or representation… 34 arXiv — Machine Learning research 1mo ago DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference arXiv:2608.08878v1 Announce Type: new Abstract: Transformer-based large language models (LLMs) achieve strong performance across many tasks, but their Key-Value (KV) cache grows linearly with sequence length, creating a severe memory bottleneck for long-context inference.… 32 Hugging Face Daily Papers research 1mo ago Motif 3: Technical Report Abstract Motif 3 is a large sparse mixture-of-experts language model using grouped differential latent attention and specialized training techniques to achieve strong reasoning, coding, and long-context performance. Generated by thinkingmachines/Inkling-Small We introduce Motif… 16 r/LocalLLaMA community 1mo ago Please Share Your Experience About Muse Glimmer I have a classic test for local LLM's. I asked for 8 ball pool game with only one HTML file and Muse Glimmer spend 21k Token(I m using full context so 128k) and only created a 220 lines of HTML and said its done. With my experience its not even close to Qwen 3.6 27B and we are… 19 NVIDIA Developer Blog official-blog 1mo ago Run Local Agentic AI Workflows with Meta’s Muse Glimmer on NVIDIA Meta returns to the open source ecosystem with the release of Muse Glimmer, a 30B open-weight dense model with a 120K+ context window built for local AI... 37 r/LocalLLaMA community 1mo ago 1M context with 17 GB model in 24 GB VRAM: "for the first time I was able to load a context of almost 1M tokens and extract 7 needles from various parts of the text" https://preview.redd.it/xxjh11f38jih1.png?width=1852&format=png&auto=webp&s=76850ed51e29a8bc86c2ca718d4320075eed4363 Just wanted to share a user report that I found to be very interesting. Some person with an intriguing name manu69x managed to run 1M context on a single RTX 3090… 30 arXiv — Machine Learning research 1mo ago Every Cache Entry Earns Its Place: Global Allocation of Resolution and Coverage for KV Cache Compression arXiv:2608.07001v1 Announce Type: new Abstract: As large language models (LLMs) process increasingly long contexts, KV cache storage and repeated access have become a major bottleneck. Existing KV cache compression methods rely on predefined, fixed compression rules and are… 30 arXiv — NLP / Computation & Language research 1mo ago Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry arXiv:2608.06849v1 Announce Type: new Abstract: Long-context LLM inference is bottlenecked by quadratic attention computation and growing KV-cache costs. Existing sparse attention and KV-compression methods typically decide which tokens or heads to preserve from runtime… 27 arXiv — NLP / Computation & Language research 1mo ago CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG arXiv:2608.07458v1 Announce Type: new Abstract: Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise… 35 r/LocalLLaMA community 1mo ago 128GB vs 256gb of ram Imagine you have 128gb of VRAM. what accompanying ram capacity you would choose (DDR4 8channel)? For example Deepseek v4 flash in q8 takes around 170GB + 12GB Dflash + ~10GB per 1m context so it’s under 200gb. so 128 + 128 should be good But for something like MiMo… 5 Page 3 of 10 · 500 articles ← Newer Older →