A key question behind this work is: should associative memory treat every new observation with the same level of confidence?</p>\n<p>Existing Delta-rule models decide how strongly to update memory from the current token, but they do not explicitly track how certain the model already is about what has been stored. In Kalman Delta Networks (KDNs), we introduce uncertainty into recurrent associative memory through a state-space perspective, allowing each update to adapt to both accumulated evidence and observation reliability. </p>\n<p>This connects classical Kalman filtering with modern linear attention, while still leading to efficient, scan-compatible architectures for large-scale language modeling.</p>\n","updatedAt":"2026-09-09T03:31:50.694Z","author":{"_id":"65497582db4fc50d794c4d7e","avatarUrl":"/avatars/71a57f122a9fe6ee07e48a170b8860c1.svg","fullname":"Ngoc Bui","name":"ngocbh","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9199239611625671},"editors":["ngocbh"],"editorAvatarUrls":["/avatars/71a57f122a9fe6ee07e48a170b8860c1.svg"],"reactions":[{"reaction":"👀","users":["JunHill"],"count":1}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.07816","authors":[{"_id":"6aa0d106d0174964227bece9","name":"Ngoc Bui","hidden":false},{"_id":"6aa0d106d0174964227becea","name":"Tinglin Huang","hidden":false},{"_id":"6aa0d106d0174964227beceb","name":"Rex Ying","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/65497582db4fc50d794c4d7e/Xo_UjhCZtdFT2nZrPPiU9.png"],"publishedAt":"2026-09-07T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"Kalman Delta Networks: Uncertainty-aware Associative Memory","submittedOnDailyBy":{"_id":"65497582db4fc50d794c4d7e","avatarUrl":"/avatars/71a57f122a9fe6ee07e48a170b8860c1.svg","isPro":false,"fullname":"Ngoc Bui","user":"ngocbh","type":"user","name":"ngocbh"},"summary":"Linear attention is increasingly used in frontier language models for efficient long-context inference and constant-memory decoding. Its fixed-size recurrent memory, however, requires an online decision at each token: what to write and how strongly to overwrite existing associations before knowing which information future queries will require. Delta-rule models learn this strength from the current token embedding but do not track confidence in the memory estimate, preventing each write from adapting to accumulated evidence. To represent this uncertainty explicitly, we reformulate recurrent associative memory as a linear--Gaussian state-space model, for which the Kalman filter is the optimal recursive estimator, and introduce a new family of models, Kalman Delta Networks (KDNs). Within KDNs, the transition propagates both the memory state and its uncertainty, allowing the Kalman gain to weight each residual write by accumulated evidence and observation reliability. Under this formulation, Delta-style updates emerge as a special case that substitutes a token-wise isotropic surrogate for predictive covariance and omits covariance tracking. Exact tracking, however, entails a dense, state-dependent Riccati recursion that is poorly suited to GPU-parallel linear-attention scans. To address this issue, we introduce two scan-compatible KDN approximations. Diagonal KDN projects each one-step posterior onto the diagonal Gaussian family through online mean-field variational inference, whereas Isotropic KDN uses an isotropic approximation with a single uncertainty scalar per head. Their uncertainty recurrences are Mobius maps, enabling associative scans with logarithmic parallel depth. Across controlled pretraining at 750M and 1.3B parameters, KDN variants consistently improve perplexity and mean downstream accuracy over state-of-the-art linear-attention models.","upvotes":21,"discussionId":"6aa0d107d0174964227becec","githubRepo":"https://github.com/ngocbh/kalman-delta-networks","githubRepoAddedBy":"user","ai_summary":"Kalman Delta Networks reformulate linear attention as a linear-Gaussian state-space model with Kalman-filter updates to track memory uncertainty, yielding efficient scan-compatible approximations that improve language modeling performance.","ai_keywords":["linear attention","recurrent associative memory","linear-Gaussian state-space model","Kalman filter","Kalman Delta Networks","Kalman gain","Delta-rule","Riccati recursion","Diagonal KDN","Isotropic KDN","associative scans","Mobius maps"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"6756014dd83c390221a3815c","name":"YaleUniversity","fullname":"Yale University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6755ff00d3ff7f20aad244d2/xk0r-87S0AVfUA1XX8CG3.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"625f8694673e5862a8c05d5c","avatarUrl":"/avatars/4f22e22b23af8082a92659b72a639c55.svg","isPro":false,"fullname":"Nguyen Trung Hieu","user":"JunHill","type":"user"},{"_id":"67dabdbca6da274927c35687","avatarUrl":"/avatars/17f0b3c71acafea9dd73d249c12c99b3.svg","isPro":false,"fullname":"Dongxuan","user":"LucaXuan","type":"user"},{"_id":"68b2a4157f881fc640ba7d80","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/lMTgr3pe7pOHtMe7bVF7F.png","isPro":false,"fullname":"khtsly","user":"khtsly","type":"user"},{"_id":"6aa0d98dc5fce68381d337b5","avatarUrl":"/avatars/b12643b1a4351898d5f9defd2e16636f.svg","isPro":false,"fullname":"Linh Lê Quang","user":"linhlq","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6a6a9207ee246d8b226648d7","avatarUrl":"/avatars/549aafb4c48353b827bd5d83abb024ee.svg","isPro":false,"fullname":"Richard Smith","user":"richard-smith","type":"user"},{"_id":"6a6c86d4ef5c16124ece9275","avatarUrl":"/avatars/5b1345893c03e0465b6ecc551c05d3e0.svg","isPro":false,"fullname":"Susan Gonzalez","user":"susan-gonzalez","type":"user"},{"_id":"6a6de3b46d51d01100d6f251","avatarUrl":"/avatars/2441a21df4bfc94a281baded48b88c8b.svg","isPro":false,"fullname":"Barbara Perez","user":"barbara-perez","type":"user"},{"_id":"6a7d495da941bb0b6e4e9690","avatarUrl":"/avatars/4fbd955dd73b0a2f5e47e4c7242db1ac.svg","isPro":false,"fullname":"SableTrail","user":"SableTrail","type":"user"},{"_id":"6a9ae8c6c3a3d75ba18a5088","avatarUrl":"/avatars/d86a87f549787fc91c7d221f668b960b.svg","isPro":false,"fullname":"唐佳","user":"qiang699","type":"user"},{"_id":"6aa0dbc24bd91ffaa5df6da6","avatarUrl":"/avatars/d6a954da9e1d07d182fce20f777e0640.svg","isPro":false,"fullname":"杨桂英","user":"na06","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6756014dd83c390221a3815c","name":"YaleUniversity","fullname":"Yale University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6755ff00d3ff7f20aad244d2/xk0r-87S0AVfUA1XX8CG3.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.07816.md","query":{}}">
Kalman Delta Networks: Uncertainty-aware Associative Memory
Abstract
Kalman Delta Networks reformulate linear attention as a linear-Gaussian state-space model with Kalman-filter updates to track memory uncertainty, yielding efficient scan-compatible approximations that improve language modeling performance.
Linear attention is increasingly used in frontier language models for efficient long-context inference and constant-memory decoding. Its fixed-size recurrent memory, however, requires an online decision at each token: what to write and how strongly to overwrite existing associations before knowing which information future queries will require. Delta-rule models learn this strength from the current token embedding but do not track confidence in the memory estimate, preventing each write from adapting to accumulated evidence. To represent this uncertainty explicitly, we reformulate recurrent associative memory as a linear--Gaussian state-space model, for which the Kalman filter is the optimal recursive estimator, and introduce a new family of models, Kalman Delta Networks (KDNs). Within KDNs, the transition propagates both the memory state and its uncertainty, allowing the Kalman gain to weight each residual write by accumulated evidence and observation reliability. Under this formulation, Delta-style updates emerge as a special case that substitutes a token-wise isotropic surrogate for predictive covariance and omits covariance tracking. Exact tracking, however, entails a dense, state-dependent Riccati recursion that is poorly suited to GPU-parallel linear-attention scans. To address this issue, we introduce two scan-compatible KDN approximations. Diagonal KDN projects each one-step posterior onto the diagonal Gaussian family through online mean-field variational inference, whereas Isotropic KDN uses an isotropic approximation with a single uncertainty scalar per head. Their uncertainty recurrences are Mobius maps, enabling associative scans with logarithmic parallel depth. Across controlled pretraining at 750M and 1.3B parameters, KDN variants consistently improve perplexity and mean downstream accuracy over state-of-the-art linear-attention models.
Community
A key question behind this work is: should associative memory treat every new observation with the same level of confidence?
Existing Delta-rule models decide how strongly to update memory from the current token, but they do not explicitly track how certain the model already is about what has been stored. In Kalman Delta Networks (KDNs), we introduce uncertainty into recurrent associative memory through a state-space perspective, allowing each update to adapt to both accumulated evidence and observation reliability.
This connects classical Kalman filtering with modern linear attention, while still leading to efficient, scan-compatible architectures for large-scale language modeling.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.07816 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.07816 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.07816 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.