Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context modeling. However, existing context extension approaches typically apply continued pretraining directly without modifying these layers, overlooking the spectral properties of linear attention state dynamics. In this work, we study long-context extension of Gated DeltaNet (GDN) from a spectral perspective of transition matrix and identify two essential factors governing long-range information retrieval: (1) a sufficiently broad slow spectral band aligned with the target dependency length, and (2) the preservation of fast-decaying modes for state clearing and context switching. Based on this observation, we propose SpectralShift, a spectral reparameterization approach for long-context continual pretraining of GDNs. Specifically, SpectralShift reparameterizes the alpha projections initialization to reshape the decay spectrum by enhancing slow propagation capacity, and further introduces a learning-rate scaling for alpha projections to facilitate long-context training. Experiments show that SpectralShift consistently improves long-context capabilities over training, providing an effective and efficient solution for extending context windows of linear attention models.</p>\n","updatedAt":"2026-09-17T05:30:37.875Z","author":{"_id":"6317419f3eb2544b62389a79","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6317419f3eb2544b62389a79/8oU90du902ATtBPYCcLFK.jpeg","fullname":"Ivan Hu","name":"IvanHU","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":7,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8429237604141235},"editors":["IvanHU"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6317419f3eb2544b62389a79/8oU90du902ATtBPYCcLFK.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.14320","authors":[{"_id":"6aaa94b71d9cc4dec796222e","name":"Zian Liu","hidden":false},{"_id":"6aaa94b71d9cc4dec796222f","user":{"_id":"6317419f3eb2544b62389a79","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6317419f3eb2544b62389a79/8oU90du902ATtBPYCcLFK.jpeg","isPro":false,"fullname":"Ivan Hu","user":"IvanHU","type":"user","name":"IvanHU"},"name":"Yiwen Hu","status":"claimed_verified","statusLastChangedAt":"2026-09-17T09:10:05.275Z","hidden":false},{"_id":"6aaa94b71d9cc4dec7962230","name":"Zican Dong","hidden":false},{"_id":"6aaa94b71d9cc4dec7962231","name":"Tian Xie","hidden":false},{"_id":"6aaa94b71d9cc4dec7962232","name":"Wayne Xin Zhao","hidden":false},{"_id":"6aaa94b71d9cc4dec7962233","name":"Yucheng Ding","hidden":false},{"_id":"6aaa94b71d9cc4dec7962234","name":"Ran Tao","hidden":false},{"_id":"6aaa94b71d9cc4dec7962235","name":"Bryan Dai","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6317419f3eb2544b62389a79/bpUvmaP5ELHoVBe2ZB_E0.png"],"publishedAt":"2026-09-13T00:00:00.000Z","submittedOnDailyAt":"2026-09-17T00:00:00.000Z","title":"SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization","submittedOnDailyBy":{"_id":"6317419f3eb2544b62389a79","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6317419f3eb2544b62389a79/8oU90du902ATtBPYCcLFK.jpeg","isPro":false,"fullname":"Ivan Hu","user":"IvanHU","type":"user","name":"IvanHU"},"summary":"Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context modeling. However, existing context extension approaches typically apply continued pretraining directly without modifying these layers, overlooking the spectral properties of linear attention state dynamics. In this work, we study long-context extension of Gated DeltaNet (GDN) from a spectral perspective of transition matrix and identify two essential factors governing long-range information retrieval: (1) a sufficiently broad slow spectral band aligned with the target dependency length, and (2) the preservation of fast-decaying modes for state clearing and context switching. Based on this observation, we propose SpectralShift, a spectral reparameterization approach for long-context continual pretraining of GDNs. Specifically, SpectralShift reparameterizes the alpha projections initialization to reshape the decay spectrum by enhancing slow propagation capacity, and further introduces a learning-rate scaling for alpha projections to facilitate long-context training. Experiments show that SpectralShift consistently improves long-context capabilities over training, providing an effective and efficient solution for extending context windows of linear attention models. The code has been open-sourced at https://github.com/RUCAIBox/GDN-SpectralShift.","upvotes":21,"discussionId":"6aaa94b81d9cc4dec7962236","githubRepo":"https://github.com/RUCAIBox/GDN-SpectralShift","githubRepoAddedBy":"user","githubStars":1,"organization":{"_id":"6704ef33935b1a7c59795566","name":"RUC-AIBOX","fullname":"RUC-AIBOX","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/61b8405b516a20acdf3b85ff/Q3_mJHjNqZYfArFl1ZpAL.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6317419f3eb2544b62389a79","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6317419f3eb2544b62389a79/8oU90du902ATtBPYCcLFK.jpeg","isPro":false,"fullname":"Ivan Hu","user":"IvanHU","type":"user"},{"_id":"6a69ec703e202470681e0dad","avatarUrl":"/avatars/1812d9bb94b84c2d46903a3304e380b5.svg","isPro":false,"fullname":"Michael Lopez","user":"lopez-michael","type":"user"},{"_id":"6a6aa3723eb4da20611a1610","avatarUrl":"/avatars/5c290d2a94ef9b1a60b49f06f0c97fc9.svg","isPro":false,"fullname":"Timothy Harris","user":"Lunar-BloomW","type":"user"},{"_id":"6a6c7bee120299da51a2ac1e","avatarUrl":"/avatars/b2825420f51e1ab879aafb7724e77094.svg","isPro":false,"fullname":"Kenneth Clark","user":"CedarMap","type":"user"},{"_id":"6a6c9d4f1569d2cf12abdc48","avatarUrl":"/avatars/7c0a7fc35868bc1ac7446781853db117.svg","isPro":false,"fullname":"Edward Wilson","user":"novaLoom","type":"user"},{"_id":"6a6d57a4c1e23c5ff69d1331","avatarUrl":"/avatars/5b3f532a7739b93eb6a2e2c69ec7aa03.svg","isPro":false,"fullname":"William Thompson","user":"ZenithMind","type":"user"},{"_id":"6a6da8c0f372a51769692d07","avatarUrl":"/avatars/5b90452cb5968daa2d4822e2e9fc677a.svg","isPro":false,"fullname":"Robert Martinez","user":"IndigoPulse","type":"user"},{"_id":"6a6dc9e9097052be15666854","avatarUrl":"/avatars/a924e428da37f239c4bdf89900e1ba35.svg","isPro":false,"fullname":"Patricia Gonzalez","user":"Indigo-Patricia","type":"user"},{"_id":"6a6deee8c51edbf08f11aa2a","avatarUrl":"/avatars/e7c969d89201191959f8ce7541015afb.svg","isPro":false,"fullname":"Elizabeth Miller","user":"Elizabeth-Miller","type":"user"},{"_id":"6a7d48bf05b6bd357d677065","avatarUrl":"/avatars/477d5908e7bbb4bbf1b1b12bfec7446f.svg","isPro":false,"fullname":"QuietDawn","user":"QuietDawn","type":"user"},{"_id":"6a8115f038097d1508f45dc1","avatarUrl":"/avatars/99c90da410a200532f62a226c891c16f.svg","isPro":false,"fullname":"chenxi lin","user":"SableNico","type":"user"},{"_id":"6a9ae3710119f5dea2bbf2fb","avatarUrl":"/avatars/22b62a32d3c7c35761eabf05e8b43206.svg","isPro":false,"fullname":"서지현","user":"ti943","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6704ef33935b1a7c59795566","name":"RUC-AIBOX","fullname":"RUC-AIBOX","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/61b8405b516a20acdf3b85ff/Q3_mJHjNqZYfArFl1ZpAL.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.14320.md","query":{}}">
SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization
Abstract
Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context modeling. However, existing context extension approaches typically apply continued pretraining directly without modifying these layers, overlooking the spectral properties of linear attention state dynamics. In this work, we study long-context extension of Gated DeltaNet (GDN) from a spectral perspective of transition matrix and identify two essential factors governing long-range information retrieval: (1) a sufficiently broad slow spectral band aligned with the target dependency length, and (2) the preservation of fast-decaying modes for state clearing and context switching. Based on this observation, we propose SpectralShift, a spectral reparameterization approach for long-context continual pretraining of GDNs. Specifically, SpectralShift reparameterizes the alpha projections initialization to reshape the decay spectrum by enhancing slow propagation capacity, and further introduces a learning-rate scaling for alpha projections to facilitate long-context training. Experiments show that SpectralShift consistently improves long-context capabilities over training, providing an effective and efficient solution for extending context windows of linear attention models. The code has been open-sourced at https://github.com/RUCAIBox/GDN-SpectralShift.
Community
Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context modeling. However, existing context extension approaches typically apply continued pretraining directly without modifying these layers, overlooking the spectral properties of linear attention state dynamics. In this work, we study long-context extension of Gated DeltaNet (GDN) from a spectral perspective of transition matrix and identify two essential factors governing long-range information retrieval: (1) a sufficiently broad slow spectral band aligned with the target dependency length, and (2) the preservation of fast-decaying modes for state clearing and context switching. Based on this observation, we propose SpectralShift, a spectral reparameterization approach for long-context continual pretraining of GDNs. Specifically, SpectralShift reparameterizes the alpha projections initialization to reshape the decay spectrum by enhancing slow propagation capacity, and further introduces a learning-rate scaling for alpha projections to facilitate long-context training. Experiments show that SpectralShift consistently improves long-context capabilities over training, providing an effective and efficient solution for extending context windows of linear attention models.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.14320 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.14320 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.14320 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.