<a href=\"https://cdn-uploads.huggingface.co/production/uploads/65bf7ddf0038ecd97149a44c/KByAixwGUDqMlxYyL5ngt.jpeg\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/65bf7ddf0038ecd97149a44c/KByAixwGUDqMlxYyL5ngt.jpeg\" alt=\"method_overview\"></a><br>We introduce BeaconKV, a training-free KV cache compression method for large reasoning models, based on our observation that queries revisiting distant context form a small number of clusters. BeaconKV combines representative \"beacon queries\" with recent queries to identify and preserve KV pairs that may be needed again during extended reasoning.</p>\n<p>Paper: <a href=\"https://arxiv.org/abs/2609.04971\" rel=\"nofollow\">https://arxiv.org/abs/2609.04971</a><br>Code: <a href=\"https://github.com/aiha-lab/BeaconKV\" rel=\"nofollow\">https://github.com/aiha-lab/BeaconKV</a></p>\n","updatedAt":"2026-09-09T05:04:22.752Z","author":{"_id":"65bf7ddf0038ecd97149a44c","avatarUrl":"/avatars/7a49e7d6763ad609d6cb92af570f3d89.svg","fullname":"jhk","name":"kkt20","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7752936482429504},"editors":["kkt20"],"editorAvatarUrls":["/avatars/7a49e7d6763ad609d6cb92af570f3d89.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.04971","authors":[{"_id":"6a9e1cadde5ea82090db656f","user":{"_id":"65bf7ddf0038ecd97149a44c","avatarUrl":"/avatars/7a49e7d6763ad609d6cb92af570f3d89.svg","isPro":false,"fullname":"jhk","user":"kkt20","type":"user","name":"kkt20"},"name":"Janghyeon Kim","status":"claimed_verified","statusLastChangedAt":"2026-09-07T09:57:06.741Z","hidden":false},{"_id":"6a9e1cadde5ea82090db6570","name":"Minsoo Kim","hidden":false},{"_id":"6a9e1cadde5ea82090db6571","name":"Kyuhong Shim","hidden":false},{"_id":"6a9e1cadde5ea82090db6572","name":"Jungwook Choi","hidden":false}],"publishedAt":"2026-09-04T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference","submittedOnDailyBy":{"_id":"65bf7ddf0038ecd97149a44c","avatarUrl":"/avatars/7a49e7d6763ad609d6cb92af570f3d89.svg","isPro":false,"fullname":"jhk","user":"kkt20","type":"user","name":"kkt20"},"summary":"Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generation, but the resulting key-value (KV) cache grows linearly with sequence length and creates severe memory bottlenecks, often exceeding GPU capacity for long reasoning traces. Existing KV cache compression methods rely on recent queries to estimate future token importance, implicitly assuming these serve as reliable proxies for future attention patterns. We demonstrate that this assumption fails in long-horizon reasoning: certain decoding steps generate Thought Revisiting Tokens (TRT) that re-attend to distant previous context, such as task-solving plans formulated early in the trace. Through systematic analysis, we discover that queries corresponding to the TRT cluster into a small number of similarity groups in the embedding space. Based on this insight, we propose BeaconKV, a training-free KV cache compression method that maintains beacon queries, compact representatives for each global query cluster, to anticipate which KV pairs will be revisited without storing the entire query history. Across four open-source LRMs and diverse reasoning benchmarks, BeaconKV generally outperforms existing compression methods, achieving up to 5.8times memory reduction while nearly preserving full cache accuracy and improving throughput by over 4.3times.","upvotes":29,"discussionId":"6a9e1cadde5ea82090db6573","githubRepo":"https://github.com/aiha-lab/BeaconKV","githubRepoAddedBy":"user","ai_summary":"BeaconKV improves memory efficiency for long reasoning traces by using compact beacon queries to predict which past key-value pairs will be revisited, reducing cache size without sacrificing accuracy.","ai_keywords":["Large Reasoning Models","Chain-of-Thought","KV cache compression","Thought Revisiting Tokens","beacon queries","BeaconKV"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"687f48a44e05dc7a8dc1fe72","avatarUrl":"/avatars/4dedb9a6308d680c759effc6d1803e26.svg","isPro":false,"fullname":"Jongwon Lee","user":"jowlee","type":"user"},{"_id":"67106f425bc64dd54976f486","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/2eNbWy7uWBDSTwyVreLhE.png","isPro":false,"fullname":"HSK","user":"momarom","type":"user"},{"_id":"628d111530de7f00af86bd65","avatarUrl":"/avatars/468eaed7a19781f1325d5a8c414f6be6.svg","isPro":false,"fullname":"Sihwa Lee","user":"macto","type":"user"},{"_id":"6970a0f98674a8dc932c6e6b","avatarUrl":"/avatars/7a141b9c7129532d8ee691f5aa0f153e.svg","isPro":false,"fullname":"SEOBIN SONG","user":"manononon","type":"user"},{"_id":"66597bdf2b6af3cf61aac170","avatarUrl":"/avatars/11ac7490b8e231241f0b5908a9983bd2.svg","isPro":false,"fullname":"Lee","user":"Woongkyu","type":"user"},{"_id":"66003e8b6aa2a661ddfafba2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66003e8b6aa2a661ddfafba2/XpGxIx7H1JAg3YkBQn6Tq.jpeg","isPro":false,"fullname":"Kyungmo, Koo","user":"9beans","type":"user"},{"_id":"6a2fd306114dc7ae7df93cc8","avatarUrl":"/avatars/ff95975a7248b71b170f8b345152982b.svg","isPro":false,"fullname":"Kyungmo Koo","user":"kyungmokooTT","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"665c132faeeaa9618772ae2b","avatarUrl":"/avatars/bed669a77fc264da8a33e3e9b4aae596.svg","isPro":false,"fullname":"YuJun Heo","user":"ordidxzero","type":"user"},{"_id":"6aa0f20e50271eb539c5ff1a","avatarUrl":"/avatars/8bfefb49e08260db34bca7bf7bbfbf87.svg","isPro":false,"fullname":"Chanhee Chung","user":"blind-ez","type":"user"},{"_id":"646c866ae77eaac8f8d26b7f","avatarUrl":"/avatars/1d56d3a7cbe4428fe495bd720f014edc.svg","isPro":false,"fullname":"Kim","user":"hmkim97","type":"user"},{"_id":"68773cff86eab77f79c488a9","avatarUrl":"/avatars/114f79dc4a90a43b360d7bd6e6e12037.svg","isPro":false,"fullname":"LeeJaeWon","user":"J1One","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.04971.md","query":{}}">
BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference
Published on Sep 4
· Submitted by jhk on Sep 9 Abstract
BeaconKV improves memory efficiency for long reasoning traces by using compact beacon queries to predict which past key-value pairs will be revisited, reducing cache size without sacrificing accuracy.
Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generation, but the resulting key-value (KV) cache grows linearly with sequence length and creates severe memory bottlenecks, often exceeding GPU capacity for long reasoning traces. Existing KV cache compression methods rely on recent queries to estimate future token importance, implicitly assuming these serve as reliable proxies for future attention patterns. We demonstrate that this assumption fails in long-horizon reasoning: certain decoding steps generate Thought Revisiting Tokens (TRT) that re-attend to distant previous context, such as task-solving plans formulated early in the trace. Through systematic analysis, we discover that queries corresponding to the TRT cluster into a small number of similarity groups in the embedding space. Based on this insight, we propose BeaconKV, a training-free KV cache compression method that maintains beacon queries, compact representatives for each global query cluster, to anticipate which KV pairs will be revisited without storing the entire query history. Across four open-source LRMs and diverse reasoning benchmarks, BeaconKV generally outperforms existing compression methods, achieving up to 5.8times memory reduction while nearly preserving full cache accuracy and improving throughput by over 4.3times.
Community

We introduce BeaconKV, a training-free KV cache compression method for large reasoning models, based on our observation that queries revisiting distant context form a small number of clusters. BeaconKV combines representative "beacon queries" with recent queries to identify and preserve KV pairs that may be needed again during extended reasoning.
Paper: https://arxiv.org/abs/2609.04971
Code: https://github.com/aiha-lab/BeaconKV
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.04971 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.04971 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.04971 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.