Hugging Face Daily Papers · · 3 min read

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Explorations into multimodal pretraining.</p>\n","updatedAt":"2026-08-06T02:05:21.387Z","author":{"_id":"636e6ee287545ca5a136b4c3","avatarUrl":"/avatars/208d32b1202e2da210146027212dbdd3.svg","fullname":"Junlin Han","name":"Junlinh","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":9,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8534981608390808},"editors":["Junlinh"],"editorAvatarUrls":["/avatars/208d32b1202e2da210146027212dbdd3.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.05000","authors":[{"_id":"6a73ebe0c5e410d076869a16","name":"Junlin Han","hidden":false},{"_id":"6a73ebe0c5e410d076869a17","name":"Shengbang Tong","hidden":false},{"_id":"6a73ebe0c5e410d076869a18","name":"David Fan","hidden":false},{"_id":"6a73ebe0c5e410d076869a19","name":"Minghao Chen","hidden":false},{"_id":"6a73ebe0c5e410d076869a1a","name":"Philip Torr","hidden":false},{"_id":"6a73ebe0c5e410d076869a1b","name":"Filippos Kokkinos","hidden":false},{"_id":"6a73ebe0c5e410d076869a1c","name":"Mike Lewis","hidden":false}],"publishedAt":"2026-08-05T00:00:00.000Z","submittedOnDailyAt":"2026-08-06T00:00:00.000Z","title":"Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes","submittedOnDailyBy":{"_id":"636e6ee287545ca5a136b4c3","avatarUrl":"/avatars/208d32b1202e2da210146027212dbdd3.svg","isPro":false,"fullname":"Junlin Han","user":"Junlinh","type":"user","name":"Junlinh"},"summary":"Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain underexplored. We provide empirical clarity through a systematic exploration of multimodal pretraining. Our controlled experiments on both synthetic and large-scale real-world datasets yield four key insights into the physics of multimodal pretraining: (i) Knowledge Flow: We disentangle how language, visual understanding, and visual generation transfer knowledge across modalities, revealing distinct patterns of influence and asymmetry; (ii) Synergy vs. Competition: We show that data \"complexity\" largely determines whether modalities are synergistic, identify architectural choices that promote synergy: such as shared attention and normalization with modality-specific feed-forward layers, and find that these behaviors generalize across different visual tokenizer designs; (iii) Early Unification: Unifying modalities from the very early stages and training them jointly is shown to be more effective than late alignment or sequential training. This process uncovers a vision laziness phenomenon, where delayed integration leads models to rely on language priors; (iv) Recipes: We derive efficient pretraining recipes that achieve strong generative performance using only 5% of the compute budget. These core findings are subsequently validated at scale by training multiple 13.5B MoE models on 2T tokens. We hope this study provides a principled foundation for understanding and scaling multimodal pretraining.","upvotes":28,"discussionId":"6a73ebe0c5e410d076869a1d","projectPage":"https://junlinhan.github.io/projects/physics_of_mm_pretrain/","organization":{"_id":"5e63d8713071d5be688861b8","name":"facebook","fullname":"AI at Meta","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1592839207516-noauth.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"636e6ee287545ca5a136b4c3","avatarUrl":"/avatars/208d32b1202e2da210146027212dbdd3.svg","isPro":false,"fullname":"Junlin Han","user":"Junlinh","type":"user"},{"_id":"6873264209494f4808b720a6","avatarUrl":"/avatars/695b5bfba06165b3dbdfeda2d743a588.svg","isPro":false,"fullname":"Zhang","user":"Shangzhan","type":"user"},{"_id":"62f0ecd2700bdc19558360de","avatarUrl":"/avatars/5325b4b763f30c41f30e3aec0d2b59fa.svg","isPro":false,"fullname":"Junyi Zhang","user":"Junyi42","type":"user"},{"_id":"670880950e79a8b46f7ff9dd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/670880950e79a8b46f7ff9dd/hA1TLhwlQblkFsq8wLrkB.jpeg","isPro":false,"fullname":"Juanxi Tian","user":"Juanxi","type":"user"},{"_id":"6508768dd48ed98e63864b5b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/qEXRx-OSNyOEhOJPfvrG4.jpeg","isPro":false,"fullname":"Xinhao Liu","user":"Gaaaavin","type":"user"},{"_id":"64a5d8219f3b568c202b3137","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64a5d8219f3b568c202b3137/TI20Z1lWHzZpLsMkayDbU.png","isPro":false,"fullname":"Di Chang","user":"Boese0601","type":"user"},{"_id":"63172831c92fd6fee3181f50","avatarUrl":"/avatars/0f57068a138cb181e9451bfc1ed3d1c0.svg","isPro":false,"fullname":"Xichen Pan","user":"xcpan","type":"user"},{"_id":"67ef93cc8cc3d1b9946eefad","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/M_9NtJsM4Mk3icC1VXjWw.jpeg","isPro":false,"fullname":"HouJiadong","user":"iChubai","type":"user"},{"_id":"652964f762a4885189484775","avatarUrl":"/avatars/8c7e7db5a7d0aae1c4d5e843131304b6.svg","isPro":false,"fullname":"Jianhao Yuan","user":"JianhaoDYDY","type":"user"},{"_id":"627ccf058b4e56cfc2716425","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1652346592327-noauth.jpeg","isPro":true,"fullname":"Shusheng Yang","user":"ShushengYang","type":"user"},{"_id":"65119b03d6e446efb3a5b4ce","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65119b03d6e446efb3a5b4ce/axKQipd3gPsoO0I3FNL0O.jpeg","isPro":true,"fullname":"Juexiao Zhang","user":"juexzz","type":"user"},{"_id":"679517e61703797636585760","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/Dmd8ws39tRZLwUH7wwk8i.png","isPro":false,"fullname":"Zeren Jiang","user":"jzr99","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"5e63d8713071d5be688861b8","name":"facebook","fullname":"AI at Meta","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1592839207516-noauth.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.05000.md","query":{}}">
Papers
arxiv:2608.05000

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

Published on Aug 5
· Submitted by
Junlin Han
on Aug 6
#3 Paper of the day
Authors:
,

Abstract

Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain underexplored. We provide empirical clarity through a systematic exploration of multimodal pretraining. Our controlled experiments on both synthetic and large-scale real-world datasets yield four key insights into the physics of multimodal pretraining: (i) Knowledge Flow: We disentangle how language, visual understanding, and visual generation transfer knowledge across modalities, revealing distinct patterns of influence and asymmetry; (ii) Synergy vs. Competition: We show that data "complexity" largely determines whether modalities are synergistic, identify architectural choices that promote synergy: such as shared attention and normalization with modality-specific feed-forward layers, and find that these behaviors generalize across different visual tokenizer designs; (iii) Early Unification: Unifying modalities from the very early stages and training them jointly is shown to be more effective than late alignment or sequential training. This process uncovers a vision laziness phenomenon, where delayed integration leads models to rely on language priors; (iv) Recipes: We derive efficient pretraining recipes that achieve strong generative performance using only 5% of the compute budget. These core findings are subsequently validated at scale by training multiple 13.5B MoE models on 2T tokens. We hope this study provides a principled foundation for understanding and scaling multimodal pretraining.

Community

Paper submitter about 8 hours ago

Explorations into multimodal pretraining.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.05000
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.05000 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.05000 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.05000 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers