Hugging Face Daily Papers · · 4 min read

CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2607.17715\">C2KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.23642\">Improving Long-Context Retrieval with Multi-Prefix Embedding</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.19667\">CacheWeaver: Cache-Aware Evidence Ordering for Efficient Grounded RAG Inference</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.18508\">MCompassRAG: Topic Metadata as a Semantic Compass for Paragraph-Level Retrieval</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.18381\">SproutRAG: Attention-Guided Tree Search with Progressive Embeddings for Long-Document RAG</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.17034\">KVEraser: Learning to Steer KV Cache for Efficient Localized Context Erasing</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.03148\">Lightweight Chunk Selection for Mobile Retrieval-Augmented Generation</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-08-13T01:35:24.231Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7154533863067627},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.07458","authors":[{"_id":"6a7d043f0ac8bee77474edc0","name":"Gyuwan Kim","hidden":false},{"_id":"6a7d043f0ac8bee77474edc1","name":"Cheoneum Park","hidden":false},{"_id":"6a7d043f0ac8bee77474edc2","name":"Tao Yang","hidden":false}],"publishedAt":"2026-08-07T00:00:00.000Z","submittedOnDailyAt":"2026-08-12T00:00:00.000Z","title":"CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG","submittedOnDailyBy":{"_id":"5ebe1c3bed25d76864d553e4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1639810620051-5ebe1c3bed25d76864d553e4.jpeg","isPro":false,"fullname":"Gyuwan Kim","user":"gyuwankim","type":"user","name":"gyuwankim"},"summary":"Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise still remain in the coarse-grained chunks. This paper optimizes the Pareto frontier under low prefill latency constraints while maximizing accuracy by proposing CoinRAG (Contextualized Information Nugget KV Cache Reuse for Long-Context RAG). The name metaphorically reflects our core mechanism: much like assembling small tokens (or \"coins\") to accumulate a larger value, CoinRAG compositionally reuses offline-computed, fine-grained nugget caches to form a learned contextual representation efficiently in a more semantically relevant but compact manner. Specifically, instead of full-chunk encoding, CoinRAG identifies query-relevant semantic units within retrieved chunks through two-stage retrieval and seamlessly assembles their sliced KV representations with a chunk-level context. Extensive evaluations on LongBench multi-hop question answering tasks demonstrate that CoinRAG significantly reduces operational costs and outperforms the other baselines with a new Pareto frontier and an average 5.3% relative improvement in answer quality (F1) under a standard fast prefill latency budget.","upvotes":4,"discussionId":"6a7d04400ac8bee77474edc3","ai_summary":"CoinRAG improves retrieval-augmented generation efficiency and accuracy by reusing fine-grained semantic nugget caches instead of full chunks.","ai_keywords":["Retrieval-Augmented Generation","KV cache reuse","CoinRAG","nugget caches","two-stage retrieval","multi-hop question answering","Pareto frontier"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"691d9d63284266ada1eb632a","name":"UCSantaBarbara","fullname":"University of California, Santa Barbara","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/AoPyKoRF9O8PHOoaJOi7h.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a6a92b662b8078b79bd4eb1","avatarUrl":"/avatars/47d17fa1bb148ae3fdc4e077954359ea.svg","isPro":false,"fullname":"Linda Taylor","user":"linda-taylor","type":"user"},{"_id":"6a6c83ab626b1eb992a0842b","avatarUrl":"/avatars/3f75008ac54daff76f8cf5ac20f4cd0c.svg","isPro":false,"fullname":"Michael Anderson","user":"EmberMichael","type":"user"},{"_id":"6a6d4372a5a4538841b20b30","avatarUrl":"/avatars/cca9481b6e02179085d0574d464d7299.svg","isPro":false,"fullname":"Linda Perez","user":"vectorDawn","type":"user"},{"_id":"6a6da7e008e6705013bc9e63","avatarUrl":"/avatars/c3d8bc73df06931ba549417b8af48e6b.svg","isPro":false,"fullname":"Linda Gonzalez","user":"NimbusLens","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"691d9d63284266ada1eb632a","name":"UCSantaBarbara","fullname":"University of California, Santa Barbara","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/AoPyKoRF9O8PHOoaJOi7h.png"},"query":{}}">
Papers
arxiv:2608.07458

CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG

Published on Aug 7
· Submitted by
Gyuwan Kim
on Aug 12
Authors:
,

Abstract

CoinRAG improves retrieval-augmented generation efficiency and accuracy by reusing fine-grained semantic nugget caches instead of full chunks.

Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise still remain in the coarse-grained chunks. This paper optimizes the Pareto frontier under low prefill latency constraints while maximizing accuracy by proposing CoinRAG (Contextualized Information Nugget KV Cache Reuse for Long-Context RAG). The name metaphorically reflects our core mechanism: much like assembling small tokens (or "coins") to accumulate a larger value, CoinRAG compositionally reuses offline-computed, fine-grained nugget caches to form a learned contextual representation efficiently in a more semantically relevant but compact manner. Specifically, instead of full-chunk encoding, CoinRAG identifies query-relevant semantic units within retrieved chunks through two-stage retrieval and seamlessly assembles their sliced KV representations with a chunk-level context. Extensive evaluations on LongBench multi-hop question answering tasks demonstrate that CoinRAG significantly reduces operational costs and outperforms the other baselines with a new Pareto frontier and an average 5.3% relative improvement in answer quality (F1) under a standard fast prefill latency budget.

Community

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.07458 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.07458 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.07458 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers