We introduce Memory Decoder at Scale, which disentangles long-term memory from reasoning and scales parametric memory to 6.9B parameters and 300B training tokens. Pythia-410M + Mem-6.9B surpasses Pythia-12B across 17 tasks with 39% fewer total parameters, revealing a promising \"small backbone, large memory\" paradigm.</p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/6609a53bd81d611249ef5266/zdlYdhNlZ3mGoxp5PeGUq.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/6609a53bd81d611249ef5266/zdlYdhNlZ3mGoxp5PeGUq.png\" alt=\"general_memory_avg_selected_path_scatter_expanded\"></a></p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/6609a53bd81d611249ef5266/T3mGOgCsQ6ow8a-uY7qm4.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/6609a53bd81d611249ef5266/T3mGOgCsQ6ow8a-uY7qm4.png\" alt=\"overview-final\"></a></p>\n","updatedAt":"2026-07-31T03:35:08.776Z","author":{"_id":"6609a53bd81d611249ef5266","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6609a53bd81d611249ef5266/h31hdQFl-jhRnO8R6Gr4C.png","fullname":"Rubin Wei","name":"Rubin-Wei","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5857882499694824},"editors":["Rubin-Wei"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6609a53bd81d611249ef5266/h31hdQFl-jhRnO8R6Gr4C.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.27919","authors":[{"_id":"6a6bfe197bd25d8874c0708c","user":{"_id":"6609a53bd81d611249ef5266","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6609a53bd81d611249ef5266/h31hdQFl-jhRnO8R6Gr4C.png","isPro":false,"fullname":"Rubin Wei","user":"Rubin-Wei","type":"user","name":"Rubin-Wei"},"name":"Rubin Wei","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.398Z","hidden":false},{"_id":"6a6bfe197bd25d8874c0708d","name":"Jiaqi Cao","hidden":false},{"_id":"6a6bfe197bd25d8874c0708e","name":"Jiarui Wang","hidden":false},{"_id":"6a6bfe197bd25d8874c0708f","name":"Junming Zhang","hidden":false},{"_id":"6a6bfe197bd25d8874c07090","name":"Qipeng Guo","hidden":false},{"_id":"6a6bfe197bd25d8874c07091","name":"Bowen Zhou","hidden":false},{"_id":"6a6bfe197bd25d8874c07092","name":"Zhouhan Lin","hidden":false}],"publishedAt":"2026-07-30T00:00:00.000Z","submittedOnDailyAt":"2026-07-31T00:00:00.000Z","title":"Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory","submittedOnDailyBy":{"_id":"6609a53bd81d611249ef5266","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6609a53bd81d611249ef5266/h31hdQFl-jhRnO8R6Gr4C.png","isPro":false,"fullname":"Rubin Wei","user":"Rubin-Wei","type":"user","name":"Rubin-Wei"},"summary":"Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it at a relatively small scale. In this work, we present Memory Decoder at Scale, scaling memory models up to 6.9B parameters and pretraining them on 300B tokens. At this data scale, the combined cost of indexing and search makes a standard Faiss pipeline infeasible. We address this bottleneck with a distributed pipeline for Faiss indexing and retrieval, together with sparse, batch-wise loading of kNN distributions. Across model scales, we find that allocating more parameters to memory yields a better parameter-performance tradeoff than scaling the base model alone. On 17 benchmarks, pairing a 6.9B general memory with Pythia-410M raises its average score from 29.86 to 37.34, surpassing Pythia-12B (37.24) with 39% fewer total parameters. For Qwen3 Base models ranging from 0.6B to 14B, 1.7B domain memories improve the average score across the three domains by more than 9 points at every scale. Overall, our results demonstrate that independently scaling pretrained memory offers a more parameter efficient path to improving language model performance.","upvotes":40,"discussionId":"6a6bfe1a7bd25d8874c07093","projectPage":"https://rubin-wei.github.io/memory-decoder-at-scale/","githubRepo":"https://github.com/LUMIA-Group/MemoryDecoder-at-Scale","githubRepoAddedBy":"user","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6609a53bd81d611249ef5266","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6609a53bd81d611249ef5266/h31hdQFl-jhRnO8R6Gr4C.png","isPro":false,"fullname":"Rubin Wei","user":"Rubin-Wei","type":"user"},{"_id":"6651f8441b1ce9f4a61a09ee","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/YM8m6_FLN_eF5PmqD7BfA.png","isPro":false,"fullname":"Zhuang Yumin","user":"Astricaelus","type":"user"},{"_id":"69e60f5ee5374a180de62a77","avatarUrl":"/avatars/c76d5f92719262ebf9c192f6541c4204.svg","isPro":false,"fullname":"Zhiqi Yang","user":"YSSYE","type":"user"},{"_id":"69dcd9ed7280ac222b59a41d","avatarUrl":"/avatars/383e5a489ec9120a648b98d0a3e50b3d.svg","isPro":false,"fullname":"Jingzhi Wang","user":"jzwang666","type":"user"},{"_id":"66d8512c54209e9101811e8e","avatarUrl":"/avatars/62dfd8e6261108f2508efe678d5a2a57.svg","isPro":false,"fullname":"M Saad Salman","user":"MSS444","type":"user"},{"_id":"69855c4acdc038b0a77f1514","avatarUrl":"/avatars/1ac66abc06d2803b0a2a344acd980ce7.svg","isPro":false,"fullname":"Hao Doou","user":"Hao126","type":"user"},{"_id":"66458107219ad12f47bc8fd4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66458107219ad12f47bc8fd4/8NqMRPB2Ko4GIOtL7ZzOj.jpeg","isPro":false,"fullname":"Yixuan Wang","user":"LuckyOrz","type":"user"},{"_id":"670a63544aba6241d593a81e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/670a63544aba6241d593a81e/4E_RifeWzh9rc1XaRP3Eq.png","isPro":false,"fullname":"ming","user":"hollyevil","type":"user"},{"_id":"68f48219a63f618c0137d354","avatarUrl":"/avatars/d48189da1e2ce350021288e2453d5cd6.svg","isPro":false,"fullname":"tuling","user":"merdft","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"674ec18bb094645555cc9e6a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ZCpui8sDqaBWfw3OMurLg.png","isPro":false,"fullname":"Rixin Rao","user":"kukojoy","type":"user"},{"_id":"66f2dee9e3e389facb05ee16","avatarUrl":"/avatars/451a16bf8ef5f3cf856521232eceaa4f.svg","isPro":false,"fullname":"(SII) Liu Jiawei","user":"Liu-Jiawei","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"query":{}}">
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
Abstract
Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it at a relatively small scale. In this work, we present Memory Decoder at Scale, scaling memory models up to 6.9B parameters and pretraining them on 300B tokens. At this data scale, the combined cost of indexing and search makes a standard Faiss pipeline infeasible. We address this bottleneck with a distributed pipeline for Faiss indexing and retrieval, together with sparse, batch-wise loading of kNN distributions. Across model scales, we find that allocating more parameters to memory yields a better parameter-performance tradeoff than scaling the base model alone. On 17 benchmarks, pairing a 6.9B general memory with Pythia-410M raises its average score from 29.86 to 37.34, surpassing Pythia-12B (37.24) with 39% fewer total parameters. For Qwen3 Base models ranging from 0.6B to 14B, 1.7B domain memories improve the average score across the three domains by more than 9 points at every scale. Overall, our results demonstrate that independently scaling pretrained memory offers a more parameter efficient path to improving language model performance.
Community
We introduce Memory Decoder at Scale, which disentangles long-term memory from reasoning and scales parametric memory to 6.9B parameters and 300B training tokens. Pythia-410M + Mem-6.9B surpasses Pythia-12B across 17 tasks with 39% fewer total parameters, revealing a promising "small backbone, large memory" paradigm.


Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.27919 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.27919 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.