Hugging Face Daily Papers · · 5 min read

HyQuant: Hybrid-Precision Quantization for LLM Attention

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

HyQuant: Hybrid-Precision Quantization for LLM Attention (EMNLP 2026 Main)</p>\n<p>Low-bit attention / KV-cache quantization often hurts long-context reasoning,<br>because the error concentrates on a small set of tokens. Across Qwen3, Llama-3,<br>Gemma 4 and Qwen3.5, we observe persistent \"vertical lines\" in attention maps:<br>the top-5% key positions plus a 128-token local window capture 82-86% of the<br>attention mass.</p>\n<p>HyQuant keeps these vertical-line tokens and the local window in FP16 and<br>quantizes the rest to 4-bit (K4V4), with fused hybrid-precision kernels for both<br>prefill and decode, saving memory by preventing KV cache materialization. Identifying the vertical lines costs only 3-5% of runtime.</p>\n<p>Highlights (single H100):</p>\n<ul>\n<li>Decode kernel up to 3.58x faster than FlashAttention-2 at 32K context;<br>end-to-end decode 1.04-1.17x faster, while KIVI / KVTuner end up slower than FA2 (0.69-0.80x)</li>\n<li>Near-full-precision accuracy: LongBench avg 45.04 vs. 44.59 for FA2 on Qwen3-8B<br>(thinking mode), vs. 37.7-40.5 for KIVI / SageAttention / KVTuner; results hold<br>on Qwen3-32B, Llama-3.1-8B and GLM-4-9B</li>\n<li>Only method that still runs at batch 16 with a 32K prefix (231.6 tok/s);<br>FA2, KIVI and KVTuner run out of memory</li>\n</ul>\n<p>Code: <a href=\"https://github.com/jerrysfls/HyQuant\" rel=\"nofollow\">https://github.com/jerrysfls/HyQuant</a><br>Happy to answer questions!</p>\n","updatedAt":"2026-09-11T03:15:36.455Z","author":{"_id":"6847c20bb510d25b2322fa5c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6847c20bb510d25b2322fa5c/LWov6B3_1BVbkYWewUyL-.png","fullname":"Jiatong Ding","name":"jerrysfls","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8074182868003845},"editors":["jerrysfls"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6847c20bb510d25b2322fa5c/LWov6B3_1BVbkYWewUyL-.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.27875","authors":[{"_id":"6a9e97efaf127efa95389b39","user":{"_id":"6847c20bb510d25b2322fa5c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6847c20bb510d25b2322fa5c/LWov6B3_1BVbkYWewUyL-.png","isPro":false,"fullname":"Jiatong Ding","user":"jerrysfls","type":"user","name":"jerrysfls"},"name":"Jiatong Ding","status":"claimed_verified","statusLastChangedAt":"2026-09-11T00:45:04.234Z","hidden":false},{"_id":"6a9e97efaf127efa95389b3a","name":"Bingxin Xing","hidden":false},{"_id":"6a9e97efaf127efa95389b3b","name":"Yu Zhang","hidden":false},{"_id":"6a9e97efaf127efa95389b3c","name":"Dian Ding","hidden":false},{"_id":"6a9e97efaf127efa95389b3d","name":"Xiaodong Yi","hidden":false},{"_id":"6a9e97efaf127efa95389b3e","name":"Xianbin Ouyang","hidden":false},{"_id":"6a9e97efaf127efa95389b3f","name":"Feihu Zhou","hidden":false},{"_id":"6a9e97efaf127efa95389b40","name":"Kun Zhang","hidden":false},{"_id":"6a9e97efaf127efa95389b41","name":"Zhenyu Guo","hidden":false},{"_id":"6a9e97efaf127efa95389b42","name":"Hao Pan","hidden":false},{"_id":"6a9e97efaf127efa95389b43","name":"Guangtao Xue","hidden":false},{"_id":"6a9e97efaf127efa95389b44","name":"Yiming Zhang","hidden":false}],"publishedAt":"2026-08-28T00:00:00.000Z","submittedOnDailyAt":"2026-09-11T00:00:00.000Z","title":"HyQuant: Hybrid-Precision Quantization for LLM Attention","submittedOnDailyBy":{"_id":"6847c20bb510d25b2322fa5c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6847c20bb510d25b2322fa5c/LWov6B3_1BVbkYWewUyL-.png","isPro":false,"fullname":"Jiatong Ding","user":"jerrysfls","type":"user","name":"jerrysfls"},"summary":"Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the attention module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods mainly rely on smoothing techniques to handle outliers, while we propose a hybrid quantization design to better balance accuracy and efficiency. Specifically, we propose HyQuant, an efficient hybrid quantization framework for LLM attention. HyQuant quantizes most attention states into low-bit formats while retaining a small set of vertical-line tokens and local-window states in high precision. These accuracy-critical regions are selected using lightweight vertical-line-aware attention-pattern signals, reducing quantization error with limited overhead. In the Prefill stage, HyQuant uses a hybrid-precision quantized attention operator that preserves vertical-line tokens and a local sliding window in full precision while quantizing the remaining context. In the Decode stage, HyQuant applies the same principle to KV-cache compression and fuses KV dequantization with attention computation to improve memory and hardware efficiency. Across diverse tasks, models, and datasets, HyQuant maintains nearly lossless accuracy with an extremely simple design, demonstrating the efficiency and practical feasibility of hybrid quantization for LLM attention. Code is available at: https://github.com/jerrysfls/HyQuant .","upvotes":3,"discussionId":"6a9e97f0af127efa95389b4c","githubRepo":"https://github.com/jerrysfls/HyQuant","githubRepoAddedBy":"user","ai_summary":"HyQuant improves low-bit LLM attention quantization by preserving critical vertical-line tokens and local windows in high precision while quantizing the rest, maintaining accuracy with low overhead.","ai_keywords":["hybrid quantization","attention quantization","low-bit quantization","vertical-line tokens","local-window states","KV-cache compression","prefill stage","decode stage"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":8,"organization":{"_id":"63e5ef7bf2e9a8f22c515654","name":"SJTU","fullname":"Shanghai Jiao Tong University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676013394657-63e5ee22b6a40bf941da0928.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6847c20bb510d25b2322fa5c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6847c20bb510d25b2322fa5c/LWov6B3_1BVbkYWewUyL-.png","isPro":false,"fullname":"Jiatong Ding","user":"jerrysfls","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63e5ef7bf2e9a8f22c515654","name":"SJTU","fullname":"Shanghai Jiao Tong University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676013394657-63e5ee22b6a40bf941da0928.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.27875.md","query":{}}">
Papers
arxiv:2608.27875

HyQuant: Hybrid-Precision Quantization for LLM Attention

Published on Aug 28
· Submitted by
Jiatong Ding
on Sep 11
Authors:

Abstract

HyQuant improves low-bit LLM attention quantization by preserving critical vertical-line tokens and local windows in high precision while quantizing the rest, maintaining accuracy with low overhead.

Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the attention module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods mainly rely on smoothing techniques to handle outliers, while we propose a hybrid quantization design to better balance accuracy and efficiency. Specifically, we propose HyQuant, an efficient hybrid quantization framework for LLM attention. HyQuant quantizes most attention states into low-bit formats while retaining a small set of vertical-line tokens and local-window states in high precision. These accuracy-critical regions are selected using lightweight vertical-line-aware attention-pattern signals, reducing quantization error with limited overhead. In the Prefill stage, HyQuant uses a hybrid-precision quantized attention operator that preserves vertical-line tokens and a local sliding window in full precision while quantizing the remaining context. In the Decode stage, HyQuant applies the same principle to KV-cache compression and fuses KV dequantization with attention computation to improve memory and hardware efficiency. Across diverse tasks, models, and datasets, HyQuant maintains nearly lossless accuracy with an extremely simple design, demonstrating the efficiency and practical feasibility of hybrid quantization for LLM attention. Code is available at: https://github.com/jerrysfls/HyQuant .

Community

Paper author Paper submitter about 11 hours ago

HyQuant: Hybrid-Precision Quantization for LLM Attention (EMNLP 2026 Main)

Low-bit attention / KV-cache quantization often hurts long-context reasoning,
because the error concentrates on a small set of tokens. Across Qwen3, Llama-3,
Gemma 4 and Qwen3.5, we observe persistent "vertical lines" in attention maps:
the top-5% key positions plus a 128-token local window capture 82-86% of the
attention mass.

HyQuant keeps these vertical-line tokens and the local window in FP16 and
quantizes the rest to 4-bit (K4V4), with fused hybrid-precision kernels for both
prefill and decode, saving memory by preventing KV cache materialization. Identifying the vertical lines costs only 3-5% of runtime.

Highlights (single H100):

  • Decode kernel up to 3.58x faster than FlashAttention-2 at 32K context;
    end-to-end decode 1.04-1.17x faster, while KIVI / KVTuner end up slower than FA2 (0.69-0.80x)
  • Near-full-precision accuracy: LongBench avg 45.04 vs. 44.59 for FA2 on Qwen3-8B
    (thinking mode), vs. 37.7-40.5 for KIVI / SageAttention / KVTuner; results hold
    on Qwen3-32B, Llama-3.1-8B and GLM-4-9B
  • Only method that still runs at batch 16 with a 32K prefix (231.6 tok/s);
    FA2, KIVI and KVTuner run out of memory

Code: https://github.com/jerrysfls/HyQuant
Happy to answer questions!

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.27875
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.27875 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.27875 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.27875 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers