<a href=\"https://github.com/HJSang/LatentPress\" rel=\"nofollow\">https://github.com/HJSang/LatentPress</a></p>\n","updatedAt":"2026-09-04T01:57:24.806Z","author":{"_id":"646af200ca17a49700e94aa1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646af200ca17a49700e94aa1/7nqby5QEtLktCMrBCqwuo.jpeg","fullname":"Hejian Sang","name":"pb09204048","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":11,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6986408233642578},"editors":["pb09204048"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/646af200ca17a49700e94aa1/7nqby5QEtLktCMrBCqwuo.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.01507","authors":[{"_id":"6a98e8b0fea818274321fefd","name":"Zhengze Zhou","hidden":false},{"_id":"6a98e8b0fea818274321fefe","name":"Hejian Sang","hidden":false}],"publishedAt":"2026-09-01T00:00:00.000Z","submittedOnDailyAt":"2026-09-04T00:00:00.000Z","title":"LatentPress: Context Compression Beyond Text and Vision","submittedOnDailyBy":{"_id":"646af200ca17a49700e94aa1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646af200ca17a49700e94aa1/7nqby5QEtLktCMrBCqwuo.jpeg","isPro":false,"fullname":"Hejian Sang","user":"pb09204048","type":"user","name":"pb09204048"},"summary":"Compressed context is usually carried as human-readable text or as rendered images that must be decoded, even when its consumer is a language model. We introduce LatentPress, which writes conversational histories and long documents into a third representation: continuous memory tokens that a frozen decoder reads directly through its input-embedding interface, with no text reconstruction at inference. A small reader-matched writer compresses 4-16times while training only an adapter (4.2M-26.2M parameters, sim!0.1% of the decoder). On LongMemEval, LatentPress reaches 0.504 accuracy at 7.70times compression versus 0.490 for uncompressed evidence, outperforming text summaries (0.184) and OCR-based compression (0.426 to 0.312). On LongBench-QA, in-domain writers match or exceed raw-context reading at 4-8times compression, while 16times trails raw. Writing takes 43ms per conversation, roughly an order of magnitude faster than text summarization or OCR reconstruction, and reading is 5-9times faster than raw context or cached OCR. We validate the interface under two transfer settings, zero-shot from UltraChat to LongMemEval memory QA and from LongMemEval-derived QA to unseen LongBench document domains, establishing direct soft tokens as a practical machine-facing context interface beyond text and vision. The implementation of the experiments could be found at: https://github.com/xuyd16ai/context_softtoken_compress .","upvotes":56,"discussionId":"6a98e8b0fea818274321feff","projectPage":"https://github.com/HJSang/LatentPress","ai_summary":"LatentPress compresses conversational and document context into continuous memory tokens read directly by a frozen decoder, achieving high compression with faster inference and improved accuracy over text or OCR methods.","ai_keywords":["continuous memory tokens","frozen decoder","input-embedding interface","adapter","LongMemEval","LongBench-QA","soft tokens","context compression"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a8f701eedea20960e186dad","avatarUrl":"/avatars/ed114e3d5e8fdef9fb359db6456adf32.svg","isPro":false,"fullname":"Eri","user":"ssxqc","type":"user"},{"_id":"6a69ec931b577a27e51d8379","avatarUrl":"/avatars/d2226010b481746cf7609a641aea8b77.svg","isPro":false,"fullname":"David Anderson","user":"david-anderson-research","type":"user"},{"_id":"6a6a829fa698558ca76157ee","avatarUrl":"/avatars/98e9a42de130397e7f4efce0039cdacd.svg","isPro":false,"fullname":"Richard Williams","user":"richardwilliams","type":"user"},{"_id":"6a6a925739a7f0911330c546","avatarUrl":"/avatars/21568a0081650e63921f1f70775f8bdd.svg","isPro":false,"fullname":"Timothy Lee","user":"timothylee","type":"user"},{"_id":"646af200ca17a49700e94aa1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646af200ca17a49700e94aa1/7nqby5QEtLktCMrBCqwuo.jpeg","isPro":false,"fullname":"Hejian Sang","user":"pb09204048","type":"user"},{"_id":"6a6a92b662b8078b79bd4eb1","avatarUrl":"/avatars/47d17fa1bb148ae3fdc4e077954359ea.svg","isPro":false,"fullname":"Linda Taylor","user":"linda-taylor","type":"user"},{"_id":"6a6a9388b172d8c070b60494","avatarUrl":"/avatars/60f0e38dfc24d995b212b77987a24c38.svg","isPro":false,"fullname":"Steven Davis","user":"Cedar-Steven","type":"user"},{"_id":"6a6a9c88342f6961a3abcbad","avatarUrl":"/avatars/ffbf41907342cc03cad48d1f0ac01394.svg","isPro":false,"fullname":"George Martinez","user":"GeorgeMartinez","type":"user"},{"_id":"6a6aa11e9859d6ac81f379dd","avatarUrl":"/avatars/677da06f7762a169543cf98dfecfdbd6.svg","isPro":false,"fullname":"Andrew Rodriguez","user":"Velvet-Andrew","type":"user"},{"_id":"6a6aa186bbca071c718a0e36","avatarUrl":"/avatars/e384ce7c7b979118de08183475518018.svg","isPro":false,"fullname":"Richard Johnson","user":"zenithfield","type":"user"},{"_id":"6a6aa1f73550efadfe67147d","avatarUrl":"/avatars/783cd172a79ae72d02835468572bff74.svg","isPro":false,"fullname":"Anthony Williams","user":"Echo-Kai","type":"user"},{"_id":"6a6aa3c12101a78421dd2816","avatarUrl":"/avatars/5ffa1ef177241e57dc07849448b8a99e.svg","isPro":false,"fullname":"Anthony Williams","user":"Cedar-Remy","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.01507.md","query":{}}">
LatentPress: Context Compression Beyond Text and Vision
Abstract
LatentPress compresses conversational and document context into continuous memory tokens read directly by a frozen decoder, achieving high compression with faster inference and improved accuracy over text or OCR methods.
Compressed context is usually carried as human-readable text or as rendered images that must be decoded, even when its consumer is a language model. We introduce LatentPress, which writes conversational histories and long documents into a third representation: continuous memory tokens that a frozen decoder reads directly through its input-embedding interface, with no text reconstruction at inference. A small reader-matched writer compresses 4-16times while training only an adapter (4.2M-26.2M parameters, sim!0.1% of the decoder). On LongMemEval, LatentPress reaches 0.504 accuracy at 7.70times compression versus 0.490 for uncompressed evidence, outperforming text summaries (0.184) and OCR-based compression (0.426 to 0.312). On LongBench-QA, in-domain writers match or exceed raw-context reading at 4-8times compression, while 16times trails raw. Writing takes 43ms per conversation, roughly an order of magnitude faster than text summarization or OCR reconstruction, and reading is 5-9times faster than raw context or cached OCR. We validate the interface under two transfer settings, zero-shot from UltraChat to LongMemEval memory QA and from LongMemEval-derived QA to unseen LongBench document domains, establishing direct soft tokens as a practical machine-facing context interface beyond text and vision. The implementation of the experiments could be found at: https://github.com/xuyd16ai/context_softtoken_compress .
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.01507 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.01507 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.01507 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.