Hugging Face Daily Papers · · 3 min read

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

🚀Project Page: <a href=\"https://microsoft.github.io/Mage\" rel=\"nofollow\">https://microsoft.github.io/Mage</a><br>🔥 Code: <a href=\"https://github.com/microsoft/Mage\" rel=\"nofollow\">https://github.com/microsoft/Mage</a></p>\n","updatedAt":"2026-07-29T02:58:20.615Z","author":{"_id":"6527b7280ae663e384eb8499","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6527b7280ae663e384eb8499/73yF3eu2cUx7jVZrhXnXx.jpeg","fullname":"Senqiao Yang","name":"Senqiao","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":15,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6198073029518127},"editors":["Senqiao"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6527b7280ae663e384eb8499/73yF3eu2cUx7jVZrhXnXx.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.24904","authors":[{"_id":"6a6968859d3a1231d492b83f","name":"Senqiao Yang","hidden":false},{"_id":"6a6968859d3a1231d492b840","name":"Kaichen Zhang","hidden":false},{"_id":"6a6968859d3a1231d492b841","name":"Zhaoyang Jia","hidden":false},{"_id":"6a6968859d3a1231d492b842","name":"Jinghao Guo","hidden":false},{"_id":"6a6968859d3a1231d492b843","name":"Yifei Shen","hidden":false},{"_id":"6a6968859d3a1231d492b844","name":"Xinjie Zhang","hidden":false},{"_id":"6a6968859d3a1231d492b845","name":"Xiaoyi Zhang","hidden":false},{"_id":"6a6968859d3a1231d492b846","name":"Haoqing Wang","hidden":false},{"_id":"6a6968859d3a1231d492b847","name":"Xiao Li","hidden":false},{"_id":"6a6968859d3a1231d492b848","name":"Peng Zhang","hidden":false},{"_id":"6a6968859d3a1231d492b849","user":{"_id":"6478679d7b370854241b2ad8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6478679d7b370854241b2ad8/dBczWYYdfEt9tQcnVGhQk.jpeg","isPro":false,"fullname":"xiangan","user":"xiangan","type":"user","name":"xiangan"},"name":"Xiang An","status":"claimed_verified","statusLastChangedAt":"2026-07-29T08:45:04.464Z","hidden":false},{"_id":"6a6968859d3a1231d492b84a","name":"Yin Xie","hidden":false},{"_id":"6a6968859d3a1231d492b84b","name":"Zhening Liu","hidden":false},{"_id":"6a6968859d3a1231d492b84c","name":"Xun Guo","hidden":false},{"_id":"6a6968859d3a1231d492b84d","name":"Jiahao Li","hidden":false},{"_id":"6a6968859d3a1231d492b84e","name":"Shicheng Zheng","hidden":false},{"_id":"6a6968859d3a1231d492b84f","name":"Jinglu Wang","hidden":false},{"_id":"6a6968859d3a1231d492b850","name":"Zongyu Guo","hidden":false},{"_id":"6a6968859d3a1231d492b851","name":"Wenxuan Xie","hidden":false},{"_id":"6a6968859d3a1231d492b852","name":"Zihan Zheng","hidden":false},{"_id":"6a6968859d3a1231d492b853","name":"Yuxuan Luo","hidden":false},{"_id":"6a6968859d3a1231d492b854","name":"Bin Li","hidden":false},{"_id":"6a6968859d3a1231d492b855","name":"Yan Lu","hidden":false}],"publishedAt":"2026-07-27T00:00:00.000Z","submittedOnDailyAt":"2026-07-29T00:00:00.000Z","title":"Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model","submittedOnDailyBy":{"_id":"6527b7280ae663e384eb8499","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6527b7280ae663e384eb8499/73yF3eu2cUx7jVZrhXnXx.jpeg","isPro":false,"fullname":"Senqiao Yang","user":"Senqiao","type":"user","name":"Senqiao"},"summary":"Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction. At its core, our custom tokenizer, Mage-ViT, replaces uniform frame sampling by selectively encoding dynamic, entropy-rich regions using motion vectors and residual energy across sparse anchor (I) and predicted (P) frames. Operating at a 16 x 16 patch level, this reduces visual token consumption by over 75% while preserving spatiotemporal context. Trained from scratch on approximately 560M unlabeled images and 100M unlabeled video frames, Mage-ViT matches or outperforms flagship encoders trained on billions of image-text pairs. We establish AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes. Furthermore, through a bio-inspired dual-system architecture - a lightweight System 1 event gate and a causal System 2 decoder - Mage-VL enables proactive streaming perception. Extensive evaluations show that Mage-VL-4B matches Qwen3-VL-4B on static tasks while achieving strong gains in video understanding and 2D/3D spatial reasoning, with up to a 3.5x wall-clock inference speedup, and comprehensively surpasses the 15B Phi-4-reasoning-vision baseline. Beyond model artifacts, we deliver seven key empirical findings covering pre-training data efficiency, variable-resolution scaling, codec system acceleration, VideoQA SFT redundancy, motion-spatial synergy, AI4AI data pipelines, and Zero-Vision SFT for multimodal RL.","upvotes":16,"discussionId":"6a6968859d3a1231d492b856","projectPage":"https://microsoft.github.io/Mage","githubRepo":"https://github.com/microsoft/Mage","githubRepoAddedBy":"user","githubStars":776,"organization":{"_id":"5e6485f787403103f9f1055e","name":"microsoft","fullname":"Microsoft","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583646260758-5e64858c87403103f9f1055d.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6527b7280ae663e384eb8499","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6527b7280ae663e384eb8499/73yF3eu2cUx7jVZrhXnXx.jpeg","isPro":false,"fullname":"Senqiao Yang","user":"Senqiao","type":"user"},{"_id":"64338d1c4521083b9d2d21da","avatarUrl":"/avatars/54b809021d794f1c4b762fbc5d0c7c90.svg","isPro":false,"fullname":"Xinjie","user":"Xinjie-Q","type":"user"},{"_id":"64bb77e786e7fb5b8a317a43","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64bb77e786e7fb5b8a317a43/J0jOrlZJ9gazdYaeSH2Bo.png","isPro":false,"fullname":"kcz","user":"kcz358","type":"user"},{"_id":"649aa367c6cf3cc95bc1b7f6","avatarUrl":"/avatars/4bf5446c261eab08fc06caebf4c5779a.svg","isPro":false,"fullname":"Yifei Shen","user":"yshenaw","type":"user"},{"_id":"6478679d7b370854241b2ad8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6478679d7b370854241b2ad8/dBczWYYdfEt9tQcnVGhQk.jpeg","isPro":false,"fullname":"xiangan","user":"xiangan","type":"user"},{"_id":"68f59ae49315a06ad9a01464","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68f59ae49315a06ad9a01464/welpgtUr6TY9Qa7nPBigg.jpeg","isPro":false,"fullname":"Sean Yu","user":"yushaohan","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"65c4eb7cd1dcbd30d86febec","avatarUrl":"/avatars/001c8f02e8ce794b2c21883628b2da72.svg","isPro":false,"fullname":"free-bit","user":"free-bit","type":"user"},{"_id":"6970c006897d3834ff6bb385","avatarUrl":"/avatars/c627588cac0ae0c6f689800862290219.svg","isPro":false,"fullname":"gg dsf","user":"ggg93949943","type":"user"},{"_id":"667aec90c06f44945546fc58","avatarUrl":"/avatars/c87ec4f74f6216c99685ace0b9e9080f.svg","isPro":false,"fullname":"Wen","user":"caes0r","type":"user"},{"_id":"688c72c011ef3399b561dee7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/688c72c011ef3399b561dee7/puhgnTOAfZYetsC46hqGm.jpeg","isPro":false,"fullname":"BoxueYang","user":"Boxue","type":"user"},{"_id":"646e1ef5075bbcc48ddf21e8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646e1ef5075bbcc48ddf21e8/g-nFu-plmEdTpnAJh_pUx.png","isPro":false,"fullname":"Pu Fanyi","user":"pufanyi","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"5e6485f787403103f9f1055e","name":"microsoft","fullname":"Microsoft","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583646260758-5e64858c87403103f9f1055d.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.24904.md","query":{}}">
Papers
arxiv:2607.24904

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

Published on Jul 27
· Submitted by
Senqiao Yang
on Jul 29
Authors:
,

Abstract

Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction. At its core, our custom tokenizer, Mage-ViT, replaces uniform frame sampling by selectively encoding dynamic, entropy-rich regions using motion vectors and residual energy across sparse anchor (I) and predicted (P) frames. Operating at a 16 x 16 patch level, this reduces visual token consumption by over 75% while preserving spatiotemporal context. Trained from scratch on approximately 560M unlabeled images and 100M unlabeled video frames, Mage-ViT matches or outperforms flagship encoders trained on billions of image-text pairs. We establish AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes. Furthermore, through a bio-inspired dual-system architecture - a lightweight System 1 event gate and a causal System 2 decoder - Mage-VL enables proactive streaming perception. Extensive evaluations show that Mage-VL-4B matches Qwen3-VL-4B on static tasks while achieving strong gains in video understanding and 2D/3D spatial reasoning, with up to a 3.5x wall-clock inference speedup, and comprehensively surpasses the 15B Phi-4-reasoning-vision baseline. Beyond model artifacts, we deliver seven key empirical findings covering pre-training data efficiency, variable-resolution scaling, codec system acceleration, VideoQA SFT redundancy, motion-spatial synergy, AI4AI data pipelines, and Zero-Vision SFT for multimodal RL.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.24904
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.24904 in a dataset README.md to link it from this page.

Spaces citing this paper

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers