HyQuant: Hybrid-Precision Quantization for LLM Attention (EMNLP 2026 Main)</p>\n<p>Low-bit attention / KV-cache quantization often hurts long-context reasoning,<br>because the error concentrates on a small set of tokens. Across Qwen3, Llama-3,<br>Gemma 4 and Qwen3.5, we observe persistent \"vertical lines\" in attention maps:<br>the top-5% key positions plus a 128-token local window capture 82-86% of the<br>attention mass.</p>\n<p>HyQuant keeps these vertical-line tokens and the local window in FP16 and<br>quantizes the rest to 4-bit (K4V4), with fused hybrid-precision kernels for both<br>prefill and decode, saving memory by preventing KV cache materialization. Identifying the vertical lines costs only 3-5% of runtime.</p>\n<p>Highlights (single H100):</p>\n<ul>\n<li>Decode kernel up to 3.58x faster than FlashAttention-2 at 32K context;<br>end-to-end decode 1.04-1.17x faster, while KIVI / KVTuner end up slower than FA2 (0.69-0.80x)</li>\n<li>Near-full-precision accuracy: LongBench avg 45.04 vs. 44.59 for FA2 on Qwen3-8B<br>(thinking mode), vs. 37.7-40.5 for KIVI / SageAttention / KVTuner; results hold<br>on Qwen3-32B, Llama-3.1-8B and GLM-4-9B</li>\n<li>Only method that still runs at batch 16 with a 32K prefix (231.6 tok/s);<br>FA2, KIVI and KVTuner run out of memory</li>\n</ul>\n<p>Code: <a href=\"https://github.com/jerrysfls/HyQuant\" rel=\"nofollow\">https://github.com/jerrysfls/HyQuant</a><br>Happy to answer questions!</p>\n","updatedAt":"2026-09-11T03:15:36.455Z","author":{"_id":"6847c20bb510d25b2322fa5c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6847c20bb510d25b2322fa5c/LWov6B3_1BVbkYWewUyL-.png","fullname":"Jiatong Ding","name":"jerrysfls","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8074182868003845},"editors":["jerrysfls"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6847c20bb510d25b2322fa5c/LWov6B3_1BVbkYWewUyL-.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.27875","authors":[{"_id":"6a9e97efaf127efa95389b39","user":{"_id":"6847c20bb510d25b2322fa5c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6847c20bb510d25b2322fa5c/LWov6B3_1BVbkYWewUyL-.png","isPro":false,"fullname":"Jiatong Ding","user":"jerrysfls","type":"user","name":"jerrysfls"},"name":"Jiatong Ding","status":"claimed_verified","statusLastChangedAt":"2026-09-11T00:45:04.234Z","hidden":false},{"_id":"6a9e97efaf127efa95389b3a","name":"Bingxin Xing","hidden":false},{"_id":"6a9e97efaf127efa95389b3b","name":"Yu Zhang","hidden":false},{"_id":"6a9e97efaf127efa95389b3c","name":"Dian Ding","hidden":false},{"_id":"6a9e97efaf127efa95389b3d","name":"Xiaodong Yi","hidden":false},{"_id":"6a9e97efaf127efa95389b3e","name":"Xianbin Ouyang","hidden":false},{"_id":"6a9e97efaf127efa95389b3f","name":"Feihu Zhou","hidden":false},{"_id":"6a9e97efaf127efa95389b40","name":"Kun Zhang","hidden":false},{"_id":"6a9e97efaf127efa95389b41","name":"Zhenyu Guo","hidden":false},{"_id":"6a9e97efaf127efa95389b42","name":"Hao Pan","hidden":false},{"_id":"6a9e97efaf127efa95389b43","name":"Guangtao Xue","hidden":false},{"_id":"6a9e97efaf127efa95389b44","name":"Yiming Zhang","hidden":false}],"publishedAt":"2026-08-28T00:00:00.000Z","submittedOnDailyAt":"2026-09-11T00:00:00.000Z","title":"HyQuant: Hybrid-Precision Quantization for LLM Attention","submittedOnDailyBy":{"_id":"6847c20bb510d25b2322fa5c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6847c20bb510d25b2322fa5c/LWov6B3_1BVbkYWewUyL-.png","isPro":false,"fullname":"Jiatong Ding","user":"jerrysfls","type":"user","name":"jerrysfls"},"summary":"Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the attention module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods mainly rely on smoothing techniques to handle outliers, while we propose a hybrid quantization design to better balance accuracy and efficiency. Specifically, we propose HyQuant, an efficient hybrid quantization framework for LLM attention. HyQuant quantizes most attention states into low-bit formats while retaining a small set of vertical-line tokens and local-window states in high precision. These accuracy-critical regions are selected using lightweight vertical-line-aware attention-pattern signals, reducing quantization error with limited overhead. In the Prefill stage, HyQuant uses a hybrid-precision quantized attention operator that preserves vertical-line tokens and a local sliding window in full precision while quantizing the remaining context. In the Decode stage, HyQuant applies the same principle to KV-cache compression and fuses KV dequantization with attention computation to improve memory and hardware efficiency. Across diverse tasks, models, and datasets, HyQuant maintains nearly lossless accuracy with an extremely simple design, demonstrating the efficiency and practical feasibility of hybrid quantization for LLM attention. Code is available at: https://github.com/jerrysfls/HyQuant .","upvotes":3,"discussionId":"6a9e97f0af127efa95389b4c","githubRepo":"https://github.com/jerrysfls/HyQuant","githubRepoAddedBy":"user","ai_summary":"HyQuant improves low-bit LLM attention quantization by preserving critical vertical-line tokens and local windows in high precision while quantizing the rest, maintaining accuracy with low overhead.","ai_keywords":["hybrid quantization","attention quantization","low-bit quantization","vertical-line tokens","local-window states","KV-cache compression","prefill stage","decode stage"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":8,"organization":{"_id":"63e5ef7bf2e9a8f22c515654","name":"SJTU","fullname":"Shanghai Jiao Tong University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676013394657-63e5ee22b6a40bf941da0928.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6847c20bb510d25b2322fa5c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6847c20bb510d25b2322fa5c/LWov6B3_1BVbkYWewUyL-.png","isPro":false,"fullname":"Jiatong Ding","user":"jerrysfls","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63e5ef7bf2e9a8f22c515654","name":"SJTU","fullname":"Shanghai Jiao Tong University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676013394657-63e5ee22b6a40bf941da0928.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.27875.md","query":{}}">
HyQuant: Hybrid-Precision Quantization for LLM Attention
Abstract
HyQuant improves low-bit LLM attention quantization by preserving critical vertical-line tokens and local windows in high precision while quantizing the rest, maintaining accuracy with low overhead.
Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the attention module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods mainly rely on smoothing techniques to handle outliers, while we propose a hybrid quantization design to better balance accuracy and efficiency. Specifically, we propose HyQuant, an efficient hybrid quantization framework for LLM attention. HyQuant quantizes most attention states into low-bit formats while retaining a small set of vertical-line tokens and local-window states in high precision. These accuracy-critical regions are selected using lightweight vertical-line-aware attention-pattern signals, reducing quantization error with limited overhead. In the Prefill stage, HyQuant uses a hybrid-precision quantized attention operator that preserves vertical-line tokens and a local sliding window in full precision while quantizing the remaining context. In the Decode stage, HyQuant applies the same principle to KV-cache compression and fuses KV dequantization with attention computation to improve memory and hardware efficiency. Across diverse tasks, models, and datasets, HyQuant maintains nearly lossless accuracy with an extremely simple design, demonstrating the efficiency and practical feasibility of hybrid quantization for LLM attention. Code is available at: https://github.com/jerrysfls/HyQuant .
Community
HyQuant: Hybrid-Precision Quantization for LLM Attention (EMNLP 2026 Main)
Low-bit attention / KV-cache quantization often hurts long-context reasoning,
because the error concentrates on a small set of tokens. Across Qwen3, Llama-3,
Gemma 4 and Qwen3.5, we observe persistent "vertical lines" in attention maps:
the top-5% key positions plus a 128-token local window capture 82-86% of the
attention mass.
HyQuant keeps these vertical-line tokens and the local window in FP16 and
quantizes the rest to 4-bit (K4V4), with fused hybrid-precision kernels for both
prefill and decode, saving memory by preventing KV cache materialization. Identifying the vertical lines costs only 3-5% of runtime.
Highlights (single H100):
- Decode kernel up to 3.58x faster than FlashAttention-2 at 32K context;
end-to-end decode 1.04-1.17x faster, while KIVI / KVTuner end up slower than FA2 (0.69-0.80x)
- Near-full-precision accuracy: LongBench avg 45.04 vs. 44.59 for FA2 on Qwen3-8B
(thinking mode), vs. 37.7-40.5 for KIVI / SageAttention / KVTuner; results hold
on Qwen3-32B, Llama-3.1-8B and GLM-4-9B
- Only method that still runs at batch 16 with a 32K prefix (231.6 tok/s);
FA2, KIVI and KVTuner run out of memory
Code: https://github.com/jerrysfls/HyQuant
Happy to answer questions!
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.27875 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.27875 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.27875 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.