Hugging Face Daily Papers · · 4 min read

Morphing into Hybrid Attention Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Morphing into Hybrid Attention Models</p>\n","updatedAt":"2026-07-03T02:19:23.074Z","author":{"_id":"66ea643899af9ac3463639b1","avatarUrl":"/avatars/252d470e761a57834dee3dbc60dfefed.svg","fullname":"Disen Lan","name":"landisen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.848112165927887},"editors":["landisen"],"editorAvatarUrls":["/avatars/252d470e761a57834dee3dbc60dfefed.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2606.30562","authors":[{"_id":"6a471bb96ee372f6920de2c6","name":"Disen Lan","hidden":false},{"_id":"6a471bb96ee372f6920de2c7","name":"Jianbin Zheng","hidden":false},{"_id":"6a471bb96ee372f6920de2c8","name":"Yuxi Ren","hidden":false},{"_id":"6a471bb96ee372f6920de2c9","name":"Xin Xia","hidden":false},{"_id":"6a471bb96ee372f6920de2ca","name":"Xuanda Wang","hidden":false},{"_id":"6a471bb96ee372f6920de2cb","name":"Xuefeng Xiao","hidden":false},{"_id":"6a471bb96ee372f6920de2cc","name":"Xipeng Qiu","hidden":false},{"_id":"6a471bb96ee372f6920de2cd","name":"Yu Cheng","hidden":false}],"publishedAt":"2026-06-29T00:00:00.000Z","submittedOnDailyAt":"2026-07-03T00:00:00.000Z","title":"Morphing into Hybrid Attention Models","submittedOnDailyBy":{"_id":"66ea643899af9ac3463639b1","avatarUrl":"/avatars/252d470e761a57834dee3dbc60dfefed.svg","isPro":false,"fullname":"Disen Lan","user":"landisen","type":"user","name":"landisen"},"summary":"Hybrid attention models improve long-context efficiency by retaining only a subset of full-attention layers and replacing the remaining layers with linear attention. However, the effectiveness of Transformer-to-hybrid conversion critically depends on which layers preserve full attention. Existing hybrid layer selection methods typically rely on heuristic strategies such as fixed placement patterns or layerwise scoring, implicitly treating layer importance as isolated and overlooking the interdependent layer effect under a global hybrid configuration. In this work, we formulate hybrid layer selection as a budget-constrained subset optimization problem. We further propose FlashMorph (Fast LAyer Selection for Hybrid MORPHing), an effective, efficient and scalable layer selection method for Transformer-to-hybrid conversion. FlashMorph first constructs a morphable model by equipping each full-attention layer with a converted linear-attention branch. It then freezes all model weights and jointly optimizes layerwise gates on synthetic long-context retrieval data, with a linearization regularization that encourages the model to rely on linear attention for efficiency. The learned gates are discretized under a preset full-attention budget to instantiate the hybrid architecture, followed by standard logits distillation and long-context finetuning. Extensive experiments show that FlashMorph discovers more effective hybrid configurations, preserves strong long-context recall and general benchmark performance while substantially reducing layer selection cost compared with existing layer selection methods, demonstrating its effectiveness, efficiency, and scalability.","upvotes":26,"discussionId":"6a471bb96ee372f6920de2ce","githubRepo":"https://github.com/LanDisen/FlashMorph","githubRepoAddedBy":"user","ai_summary":"FlashMorph is an efficient layer selection method that formulates hybrid layer selection as a budget-constrained optimization problem, using morphable models and linearization regularization to improve long-context efficiency in Transformers.","ai_keywords":["hybrid attention models","full-attention layers","linear attention","Transformer-to-hybrid conversion","subset optimization problem","morphable model","layerwise gates","synthetic long-context retrieval data","linearization regularization","logits distillation","long-context finetuning"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":4,"organization":{"_id":"67d1140985ea0644e2f14b99","name":"ByteDance-Seed","fullname":"ByteDance Seed","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6535c9e88bde2fae19b6fb25/flkDUqd_YEuFsjeNET3r-.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66ea643899af9ac3463639b1","avatarUrl":"/avatars/252d470e761a57834dee3dbc60dfefed.svg","isPro":false,"fullname":"Disen Lan","user":"landisen","type":"user"},{"_id":"6672f7c82376de4b7ab9fbd5","avatarUrl":"/avatars/60fe64dbe4c814fd8b1df7ce2bebc951.svg","isPro":false,"fullname":"Guo","user":"Rongjin03","type":"user"},{"_id":"656d8d4b1f8d9b618de91369","avatarUrl":"/avatars/884dba9e56936241034b179d11a513b9.svg","isPro":false,"fullname":"Xiangdong Zhang","user":"aHapBean","type":"user"},{"_id":"6363a1fa123a5d5cd4a800e2","avatarUrl":"/avatars/a0961ca5463aae05de0b1574c0064fae.svg","isPro":false,"fullname":"gbz","user":"greeky","type":"user"},{"_id":"64ba47b129d10d4185c46af1","avatarUrl":"/avatars/84a776d283b01f0558a28a5625115f83.svg","isPro":false,"fullname":"Zhilin Wang","user":"linzw","type":"user"},{"_id":"6570450a78d7aca0c361a177","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6570450a78d7aca0c361a177/MX7jHhTQwLs-BvYIu5rqb.jpeg","isPro":false,"fullname":"Harold Chen","user":"Harold328","type":"user"},{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"},{"_id":"640d6e06d9fcfbf4a56d89f2","avatarUrl":"/avatars/c9d50eaae109483178c152ac77387ef5.svg","isPro":false,"fullname":"Nathan Zhou","user":"Nathan01012","type":"user"},{"_id":"62f98cfa9fd0218c293b6044","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62f98cfa9fd0218c293b6044/hPNZLXOMC5simVHGx9qXI.jpeg","isPro":false,"fullname":"Hanz","user":"hanzceo","type":"user"},{"_id":"67679b5cfeac1e9f62571cf9","avatarUrl":"/avatars/4b0a0348dd0bf871aa40f8ff37703efa.svg","isPro":false,"fullname":"Zhuowen Liang","user":"SetonLiang2","type":"user"},{"_id":"66a8b2c349d0b1014615d4fa","avatarUrl":"/avatars/4acadc5097bf8a1b33e4d60bc7c821de.svg","isPro":false,"fullname":"|||||","user":"Happygameee","type":"user"},{"_id":"646cd947da8e99940b6e55cf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646cd947da8e99940b6e55cf/9c0P0WppFqNW9pdo8LgOS.jpeg","isPro":false,"fullname":"Shengyuan Ding","user":"ChrisDing1105","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"67d1140985ea0644e2f14b99","name":"ByteDance-Seed","fullname":"ByteDance Seed","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6535c9e88bde2fae19b6fb25/flkDUqd_YEuFsjeNET3r-.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2606/2606.30562.md","query":{}}">
Papers
arxiv:2606.30562

Morphing into Hybrid Attention Models

Published on Jun 29
· Submitted by
Disen Lan
on Jul 3
Authors:
,
,
,
,
,
,
,

Abstract

FlashMorph is an efficient layer selection method that formulates hybrid layer selection as a budget-constrained optimization problem, using morphable models and linearization regularization to improve long-context efficiency in Transformers.

Hybrid attention models improve long-context efficiency by retaining only a subset of full-attention layers and replacing the remaining layers with linear attention. However, the effectiveness of Transformer-to-hybrid conversion critically depends on which layers preserve full attention. Existing hybrid layer selection methods typically rely on heuristic strategies such as fixed placement patterns or layerwise scoring, implicitly treating layer importance as isolated and overlooking the interdependent layer effect under a global hybrid configuration. In this work, we formulate hybrid layer selection as a budget-constrained subset optimization problem. We further propose FlashMorph (Fast LAyer Selection for Hybrid MORPHing), an effective, efficient and scalable layer selection method for Transformer-to-hybrid conversion. FlashMorph first constructs a morphable model by equipping each full-attention layer with a converted linear-attention branch. It then freezes all model weights and jointly optimizes layerwise gates on synthetic long-context retrieval data, with a linearization regularization that encourages the model to rely on linear attention for efficiency. The learned gates are discretized under a preset full-attention budget to instantiate the hybrid architecture, followed by standard logits distillation and long-context finetuning. Extensive experiments show that FlashMorph discovers more effective hybrid configurations, preserves strong long-context recall and general benchmark performance while substantially reducing layer selection cost compared with existing layer selection methods, demonstrating its effectiveness, efficiency, and scalability.

Community

Paper submitter about 8 hours ago

Morphing into Hybrid Attention Models

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2606.30562
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2606.30562 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2606.30562 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2606.30562 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers