News / #long-context Tag Long Context 500 articles archived under #long-context · RSS Sign in to follow r/LocalLLaMA community 21d ago vllm + p2p driver hack + qwen 3.8 27B vs llamacpp + qwen flash next ? Hi everyone I'm running four rtx 4090, 64GB ram, on a threadripper pro motherboard so all PCIe x16 ports, as a homelab machine for coding. I was migrating from vllm + qwen 3.8 27B (fp8+256k kv cache) to llamacpp + qwen flash next iq4xs + 8 bit cache 200k kv cache... until… 34 r/MachineLearning community 21d ago I built a local-first hybrid router for AI Agent Skills (sub-20ms, zero tokens, runs on CPU) [P] If you use agentic workflows with custom skills or rules (Cursor rules, Claude Code slash commands, OpenCode, etc.), you have probably run into the routing trade-off: Stuff every skill definition into the system prompt (destroys your context window and degrades… 14 r/LocalLLaMA community 22d ago Block KV cache streaming: bound VRAM at long context via a shared CUDA phase arena by giveen · Pull Request #357 · TheTom/llama-cpp-turboquant So after all my work, yeah, Raymond did it better, so I ported his work over, extended it turboX, extended it multiple other models (he had only Qwen models), and benchmarked the crap out of it to make sure it was worth it still. So really the credit goes to Raymond (… 9 r/LocalLLaMA community 22d ago NInfer vs llama.cpp vs vLLM: quality + speed comparison for Qwen3.8-27B NVFP4 on RTX 5090 I've been running Qwen3.8-27B as a local inference server for a production content intelligence pipeline (HVAC industry stuff, lots of long-context retrieval and structured extraction). I have been watching other redditors post their custom configurations, and I wanted to share… 9 arXiv — NLP / Computation & Language research 24d ago SGD-KV: Summarization Guided KV Cache Compression arXiv:2609.03235v1 Announce Type: new Abstract: Large language models (LLMs) face severe memory bottlenecks in long-context inference due to the linearly growing size of key-value (KV) caches. Existing KV cache compression techniques typically rely on simple heuristics,… 5 Hugging Face Daily Papers research 24d ago Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Abstract Fully quantizing hybrid LLMs—including recurrent Gated DeltaNet layers—to 4-bit NVFP4 preserves accuracy across long-context and reasoning benchmarks by localizing outliers and exploiting robust delta-rule dynamics. Generated by thinkingmachines/Inkling-Small Hybrid… 30 r/LocalLLaMA community 24d ago Ling-3.0-flash-Fin weights released 124B total parameters, 5.1B activated parameters, and a 256K context window   submitted by   /u/Bestlife73 [link]   [comments] 12 r/LocalLLaMA community 25d ago KV cache might be a bigger problem for local models than parameter count Everyone keeps talking about fitting larger models into local hardware, but parameter count isn’t the whole story. For long context inference, KV cache can become the real memory bottleneck. Every new token adds key and value states that need to stay available, so a model that… 8 Hugging Face Daily Papers research 25d ago CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing Abstract CRISP improves long-context sparse attention by replacing indirect routing proxies with a direct structural metric and using a sink-aware threshold to eliminate background noise, achieving large speedups and better retrieval accuracy. Generated by… 26 arXiv — NLP / Computation & Language research 25d ago CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing arXiv:2609.01925v1 Announce Type: cross Abstract: The attention prefilling phase of long-context LLM inference scales quadratically, making self-attention a severe computational bottleneck. Traditional sparse attention methods mitigate this through fixed patterns or offline… 18 arXiv — NLP / Computation & Language research 25d ago Trace as State: Reasoning Traces as Conditional States for Long-Context Transformers arXiv:2609.02702v1 Announce Type: new Abstract: Transformers process information causally, but long-context reasoning may depend on task state discovered only later. We formalize this mismatch through conditional state update tasks. For causal state update processors, providing… 22 arXiv — Machine Learning research 26d ago Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding arXiv:2609.00097v1 Announce Type: new Abstract: The development of long-context Large Language Models (LLMs) is constrained by the memory bandwidth bottleneck and quadratic complexity of the attention mechanism during decoding. To overcome the inherent trade-offs between the… 6 arXiv — Machine Learning research 26d ago Context Window Failures in Relational Foundation Models arXiv:2609.00460v1 Announce Type: new Abstract: Recent Relational Deep Learning architectures have been proposed as foundation models for multi-table relational data, yet they impose constrained neighborhood budgets that force row truncation when an entity has many related… 15 arXiv — NLP / Computation & Language research 26d ago Neurosymbolics for Data Engineering: Achieving Long Context Token Reduction Without Finetuning arXiv:2609.00367v1 Announce Type: new Abstract: Large Language Models are increasingly deployed for sophisticated data engineering tasks such as generating structured queries from natural language, Text-to-SQL, and automating complex spreadsheet operations. However, maximizing… 21 Vercel — AI dev-tools 26d ago Gemini 3.8 Flash now available on AI Gateway Gemini 3.8 Flash from Google is now available on AI Gateway. The model is 50% off through December 31st. It has a 1M token context window, accepts text, image, PDF, and video input, returns text, and supports tool calling and web search. Maximum output is 65,536 tokens. Gemini… 26 Vercel — AI dev-tools 26d ago Muse Spark 1.3 now available on AI Gateway Muse Spark 1.3 from Meta is now available on AI Gateway, in both the standard and contributor pricing tiers. This model improves on prior Muse Spark models at agent work and coding, with a 1M token context window and text, image, and PDF input. On coding it takes fewer turns and… 33 r/LocalLLaMA community 26d ago Update: llama.cpp for Radeon VII / MI50 / MI60 — +14% PP, +9% long-context fill vs upstream + adaptive Flash Attention I posted a new gfx906 based llama.cpp fork a few days ago. One of the main points of critique was that i did not provide sufficient numbers for the gains to be achieved. -- TL;DR: After switching our Qwen 3.8 27B production setup to DFlash2, several of the old gfx906… 21 arXiv — Machine Learning research 27d ago SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference arXiv:2608.28911v1 Announce Type: new Abstract: The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length. We show that uniform KV quantization on a fractional-bit grid does not degrade… 11 arXiv — Machine Learning research 27d ago Higher-Dimensional Rotary Position Embedding arXiv:2608.29715v1 Announce Type: new Abstract: Transformers rely on position embedding mechanisms in long context modeling in most cases. Rotary Position Embedding (RoPE) embeds positional information with independent 2D rotations, forming relative position terms in… 12 arXiv — NLP / Computation & Language research 27d ago RouteSparse: Input-Conditional Pattern Routing for Budgeted Long-Context Prefilling arXiv:2608.29058v1 Announce Type: new Abstract: Dynamic sparse attention can reduce the quadratic cost of long-context prefilling without changing model weights. MInference assigns each attention head one pattern offline and estimates that pattern's sparse indices for every… 31 r/MachineLearning community 27d ago Sliding-window attention beats linear on long-context reasoning [R] Sliding Window Attention with sinks, one of the simplest existing fixes for the quadratic-cost problem in LLMs, holds up as well or better than the linear-attention variants labs have been spending post-training compute to produce. That is the claim of a [new arXiv preprint](… 29 Hugging Face Daily Papers research 27d ago Sliding-window beats linear attention Abstract Sliding window attention with sinks outperforms post-trained linear attention on long-context tasks without requiring retraining, offering a cheaper and more reliable inference solution. Generated by thinkingmachines/Inkling-Small Due to the nature of quadratic… 5 r/LocalLLaMA community 28d ago R9V: A designer set of kernels I've been working on for R9700s/RDNA4. Qwen3.8-Flash-Next Unsloth IQ4_XS (w/ TP on 2 R9700s, MTP, SSD n-gram, 128k ctx, vision): TG256 of *78 tok/s* (~3x increase), PP8192 of *1510 tok/s* (~30x increase). TL;DR: Ninfer/DS4 but for RDNA4 Highly custom kernels built for RDNA4, applied to vLLM-Radiance to greatly improve Qwen3.8 Flash Next speeds. This is mostly for dual R9700s with preferably 48GB of RAM or higher, but feel free to tinker. SOTA-Scan/DeepGit report in repo. For… 21 Simon Willison community 29d ago Introducing Hy4 Preview Introducing Hy4 Preview New open weight text input (no vision) LLM from Chinese company Tencent today: 770B total parameters, 49B active parameters, 1M token context window, 1.56TB on Hugging Face . This is a big size increase from their previous Hy3 in July, which was 295B, 21B… 9 r/LocalLLaMA community 29d ago (NInfer Fork) I wanted to have a 1M context Qwen-3.8 27B, tp2, dual 5090s Hey! I forked NInfer (a from-scratch C++20/CUDA inference engine for Qwen models) and added two things: tensor-parallel across two GPUs, and YaRN ×4 rope scaling. Together they let Qwen3.8-27B NVFP4 run a 1,048,576-token context on two consumer 5090s — 27.4 GB per card, no… 21 r/LocalLLaMA community 29d ago I’ve pushed llama.cpp pretty far for Qwen3.8-Flash-Next — is there any reason not to move to vLLM for 200K+ context? I'm currently running Qwen3.8-Flash-Next on a CMP 170HX 64GB + RTX 3090 24GB, with 80GB system RAM. With llama.cpp I've already spent quite a bit of time tuning it: layer split across the two GPUs, PLE on CPU, q8 KV, Flash Attention, detached MTP draft on the 170HX, and some… 29 r/LocalLLaMA community 29d ago Qwen 3.8 27B at 50 tok/s with 100k Context on a 16GB GPU! (beellama.cpp) I wanted to share my successful setup for running a Qwen 3.8 27B model with a massive context window on a consumer 16GB GPU (RTX 4070 Ti SUPER). The goal was to fit everything into VRAM without sacrificing quality or speed. 🧠 Key Components Model:… 38 arXiv — NLP / Computation & Language research 1mo ago TwinKV: A Composable Repair Pass for KV Cache Eviction via Pairwise Key Redundancy arXiv:2608.27128v1 Announce Type: new Abstract: Long-context inference is bottlenecked by the memory footprint of the key-value (KV) cache, especially for small models under tight resource budgets. Existing KV cache eviction methods score tokens using the model's attention… 14 r/LocalLLaMA community 1mo ago yall are sleeping on qwen 3.8 27b q2 + q2 dflash + q5 kv ok bit more context: it's actually a QAT Q2 for Qwen 3.8 27 B: https://huggingface.co/sdkyuan/qwen3.8-27B-qat-q2_0-gguf QAT Q2 for DFlash model: https://huggingface.co/HermiHg/Qwen3.8-27B-DFlash2-Q2_K_S-MIX-GGUF Q5 KV seems to cause 0 problems for me; I've used it up to 200K… 36 Vercel — AI dev-tools 1mo ago Hy4 Preview now available on AI Gateway Hy4 Preview from Tencent is now available on AI Gateway. Hy4 Preview is an open-source Mixture-of-Experts model with 770B total parameters aimed at long-horizon coding, document analysis, game development, and scientific reasoning. It serves a context window of 1M tokens. To use… 32 r/LocalLLaMA community 1mo ago Over 200k context on 16GB VRAM with Qwen 3.8 27B UD-IQ3_XXS I was using UD-Q3_K_XL until now with more than 140000 context. Quality wise it's very good, very few erroneous tool calls. Then I saw many others here reporting good results with IQ3_XXS, so I gave it a try. The downside is prompt processing speed went down from 700-800 tk/s to… 7 r/LocalLLaMA community 1mo ago GLM-5.3-Flash @ DGX Station GB300: ~206 tok/s (single stream), 1M context Hey all! I'm finally doing some cool stuff with my "thinking heater" (h/t u/-TV-Stand- ). I'm still experimenting with GLM-5.2 (in anticipation of 5.3 coming tomorrow, I hope!) and things are very cool so far. With the release of GLM-5.3-flash, I decided to play with it on the… 25 arXiv — NLP / Computation & Language research 1mo ago A Storage-Retrieval Gap in Parametric Knowledge Graph Memory arXiv:2608.25489v1 Announce Type: cross Abstract: Graph retrieval-augmented generation places retrieved subgraphs into the model's context window at query time, paying a recurring token cost and exposing source data on every call. We study an alternative: compiling a knowledge… 21 arXiv — Machine Learning research 1mo ago Beyond Scaling: Self-Evolving LLM Agents for Hardware Kernel Optimization via an Experience-Driven Workflow and Experience Graph Memory arXiv:2608.25570v1 Announce Type: new Abstract: Hardware kernel optimization requires repeated compilation, correctness testing, profiling, and revision. LLM agents can automate parts of this process, and stronger foundation models, longer context windows, and longer execution… 19 arXiv — NLP / Computation & Language research 1mo ago ClueWeaver: Reward-Guided Dual-Agent Evidence Reasoning for Compact LLMs on Literary Long Narratives arXiv:2608.25531v1 Announce Type: new Abstract: Humanities and social science research requires close reading of long narrative materials such as novels, scripts, archives, and case reports, yet many users have limited access to costly proprietary long-context models. Compact,… 29 arXiv — NLP / Computation & Language research 1mo ago Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context arXiv:2608.25655v1 Announce Type: new Abstract: Conversations with chat assistants increasingly span many topics in a single long-running thread, challenging memory systems. Existing long-context and memory benchmarks often expose session or topic boundaries, or probe direct… 25 Vercel — AI dev-tools 1mo ago Ling 3.0 Flash Fin now available on AI Gateway for free Ling 3.0 Flash Fin from Inclusion AI is now available on AI Gateway, free to use through September 25. Ling 3.0 Flash Fin is a finance-focused version of Ling 3.0 Flash . It has a 256K token context window, produces up to 32K output tokens, and supports reasoning and function… 29 r/LocalLLaMA community 1mo ago First serious confirmation. Ox Alpha is GLM-5.3-Flash https://x.com/romanchernin/status/2092488160680751437?s=20 - Multimodal (Vision) - 1M Tokens Context Window - DeepSWE ~63%   submitted by   /u/MrWidmoreHK [link]   [comments] 14 Smol AI News news-outlet 1mo ago not much happened today **Z.ai** launched **GLM-5.3-Flash**, a natively multimodal model with a **1M-token context window**, **320B total parameters / 18B active parameters**, under the **MIT License**. It is positioned as a price-competitive successor to GLM-5.2 and claims performance on par with… 29 arXiv — Machine Learning research 1mo ago PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression arXiv:2608.23843v1 Announce Type: new Abstract: Long-context inference in large language models (LLMs) is increasingly limited by the memory required for the key-value (KV) cache. KV cache compression addresses this problem by reducing the storage cost of previous tokens. Among… 9 Vercel — AI dev-tools 1mo ago GLM 5.3 Flash now available on AI Gateway GLM 5.3 Flash from Z.ai is now available on AI Gateway. The model is a faster, cheaper sibling of GLM 5.3 built for coding and agent tasks that run across many steps. GLM-5.3 Flash is a multimodal model that supports text and vision input, with a 1M token context window and a… 4 Vercel — AI dev-tools 1mo ago Qwen 3.8 Flash now available on AI Gateway Qwen3.8-Flash from Alibaba is now available on AI Gateway. It takes text and images as input, serves a context window of 1 million tokens, and can return up to 65k tokens in a response. Alibaba points it at coding, tool use, and multi-step agent work. To use Qwen3.8-Flash, set… 24 r/LocalLLaMA community 1mo ago Peak Portable Personal Datacenter Portable rig for Qwen3.8-27B-BF16 200K+ token prompts. My work Panasonic Toughbook + the T1 + power brick + headphones all fit in my lunchbox. Need the BF16 for huge context highly sensitive document OCR, image analysis, aggregation and summarization. I've done a ton of testing… 35 Hugging Face Daily Papers research 1mo ago RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA Modeling Abstract RIBOSPAN is a large bidirectional RNA foundation model pretrained on up to 10,240 nucleotides that enables high-resolution full-transcript modeling, strong long-context representations, and discrete-diffusion-based mRNA generation and redesign. Generated by… 29 Hugging Face Daily Papers research 1mo ago TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration Abstract TileMix routes attention score tiles to mixed FP16 or INT8 precision within fused dense attention, recovering long-context accuracy while improving prefill throughput without retraining. Generated by thinkingmachines/Inkling-Small Long-context prefill in large language… 5 NVIDIA Developer Blog official-blog 1mo ago How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubin NVIDIA Groq 3 LPX is the interactive AI inference accelerator for the NVIDIA Vera Rubin platform. At the core of the platform is NVIDIA Vera Rubin NVL72, the... 15 Smol AI News news-outlet 1mo ago not much happened today **Z.ai** released the **GLM-5.3** open-weight model family, optimized for **agentic coding** and **cyber defense**, with impressive specs like **744B total / 40B active parameters**, **1M context window**, and a **239GB 2-bit** variant retaining **81% accuracy**. **Tencent**… 28 arXiv — Machine Learning research 1mo ago BF1: A Causal Dyadic Sparse-Attention Retrofit for Efficient Long-Context Transformers arXiv:2608.20427v1 Announce Type: new Abstract: Dense causal attention remains expensive at long context even when implemented with highly optimized exact kernels. We study BF1, a deterministic block-aligned dyadic sparse-attention route that combines a small exact local… 16 arXiv — Machine Learning research 1mo ago Rethinking Expressivity and Efficiency in Test-Time Training arXiv:2608.21308v1 Announce Type: new Abstract: Test-Time Training (TTT) enables long-context processing via continuous weight updates during inference, but current methods struggle to balance the expressivity of per-token update dynamics with the hardware efficiency of… 19 arXiv — NLP / Computation & Language research 1mo ago Inhibitory Attention for Clinical Long-Context Reasoning: Characterizing and Mitigating Lost-in-the-Middle Effects in EHR Processing arXiv:2608.20348v1 Announce Type: new Abstract: Electronic health records now routinely exceed 100,000 tokens per patient. Yet large language models exhibit the lost-in-the-middle (LitM) effect: information near the center of a long context is retrieved less reliably than… 19 Page 2 of 10 · 500 articles ← Newer Older →