Hugging Face Daily Papers · · 7 min read

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Image tokenizers should be designed and evaluated as visual languages — not just as image compressors.</p>\n<p>📄 Paper: <a href=\"https://arxiv.org/abs/2609.09143\" rel=\"nofollow\">https://arxiv.org/abs/2609.09143</a><br>💻 Code: <a href=\"https://github.com/amazon-far/Tokenizer_UMM\" rel=\"nofollow\">https://github.com/amazon-far/Tokenizer_UMM</a><br>🌐 Website: <a href=\"https://lst627.github.io/tokenizers_as_visual_languages\" rel=\"nofollow\">https://lst627.github.io/tokenizers_as_visual_languages</a></p>\n","updatedAt":"2026-09-11T21:37:39.837Z","author":{"_id":"6555cece9dc61e22c5255fb4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/lnOCZmwJAElEI3EiN5up2.png","fullname":"lst","name":"lst627","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7991886138916016},"editors":["lst627"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/lnOCZmwJAElEI3EiN5up2.png"],"reactions":[],"isReport":false}},{"id":"6aa4a8e043a0c5ca08ed8430","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":379,"isUserFollowing":false},"createdAt":"2026-09-12T01:20:32.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [VoT: Vision-of-Thought for Unified Multimodal Representation Alignment](https://huggingface.co/papers/2609.07815) (2026)\n* [UniSpace: Unified Visual Representation and Scalable Multimodal Modeling](https://huggingface.co/papers/2608.08676) (2026)\n* [Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation](https://huggingface.co/papers/2607.25527) (2026)\n* [Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction](https://huggingface.co/papers/2608.12209) (2026)\n* [MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations](https://huggingface.co/papers/2608.25575) (2026)\n* [MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment](https://huggingface.co/papers/2608.11167) (2026)\n* [Where Does Generative Difficulty Reside? An Empirical Study of Target Representations](https://huggingface.co/papers/2608.00626) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2609.07815\">VoT: Vision-of-Thought for Unified Multimodal Representation Alignment</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.08676\">UniSpace: Unified Visual Representation and Scalable Multimodal Modeling</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.25527\">Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.12209\">Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.25575\">MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.11167\">MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.00626\">Where Does Generative Difficulty Reside? An Empirical Study of Target Representations</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-09-12T01:20:32.215Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":379,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.719621479511261},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.09143","authors":[{"_id":"6aa473377ba345d44ad143e5","name":"Siting Li","hidden":false},{"_id":"6aa473377ba345d44ad143e6","name":"Zhengyang Wang","hidden":false},{"_id":"6aa473377ba345d44ad143e7","name":"Simon Shaolei Du","hidden":false},{"_id":"6aa473377ba345d44ad143e8","name":"Xi Chen","hidden":false},{"_id":"6aa473377ba345d44ad143e9","name":"Yang Liu","hidden":false}],"publishedAt":"2026-09-08T00:00:00.000Z","submittedOnDailyAt":"2026-09-11T00:00:00.000Z","title":"Studying Image Tokenizers as Visual Languages in Unified Multimodal Models","submittedOnDailyBy":{"_id":"6555cece9dc61e22c5255fb4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/lnOCZmwJAElEI3EiN5up2.png","isPro":false,"fullname":"lst","user":"lst627","type":"user","name":"lst627"},"summary":"Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These evaluations do not fully capture how visual tokens behave when modeled jointly with text. We build a controlled pure-autoregressive testbed and track task-specific validation losses during multimodal continual pretraining across text, image, text-to-image (T2I), and image-to-text (I2T) prediction. We examine how these losses scale and relate to downstream performance, then use them to study multimodal learnability---how well image and text tokens are jointly modeled---and tokenizer design. We find that (1) losses should be analyzed by task, since they exhibit distinct scaling behavior and rank tokenizers differently. (2) The loss--performance relationship depends on the predicted token space: for a fixed tokenizer, T2I and I2T losses correlate with generation quality, but across tokenizers, the T2I loss--performance relationship shifts with the image-token space, whereas I2T loss, computed over a shared text vocabulary, provides a more consistent signal. I2T loss also correlates with both generation and visual understanding performance after supervised finetuning. Using losses as a lens, we show that (3) better reconstruction does not necessarily yield lower task-specific losses or stronger downstream performance, and that (4) image tokenizer choice can affect text modeling under joint optimization. As case studies, we revisit three tokenizer design axes---the discriminator, semantic supervision, and vocabulary size---to examine their effects on joint modeling and downstream performance. Together, our testbed offers a complementary perspective on image tokenizers as visual languages, highlighting their interplay with text in joint multimodal training.","upvotes":17,"discussionId":"6aa473387ba345d44ad143ea","projectPage":"https://lst627.github.io/tokenizers_as_visual_languages/","githubRepo":"https://github.com/amazon-far/Tokenizer_UMM","githubRepoAddedBy":"user","ai_summary":"Using a controlled autoregressive testbed, the study analyzes task-specific validation losses during multimodal pretraining to evaluate how image tokenizer design affects joint text-image modeling and downstream performance.","ai_keywords":["image tokenizers","multimodal continual pretraining","text-to-image","image-to-text","autoregressive","cross-modal scaling","tokenizer design","reconstruction","discriminator","semantic supervision","vocabulary size"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"68d5924a76c6b4bfe8b4ab60","name":"Amazon-FAR","fullname":"Amazon Frontier AI & Robotics","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68251d059667f9f347003874/YPxVjC8rCyexwo42cznvE.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6555cece9dc61e22c5255fb4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/lnOCZmwJAElEI3EiN5up2.png","isPro":false,"fullname":"lst","user":"lst627","type":"user"},{"_id":"689cfe7aec3a63420d1acdda","avatarUrl":"/avatars/708a8611dc99c19da5d4a5615e1067a5.svg","isPro":false,"fullname":"xinqi wang","user":"ElliotKaxddWang","type":"user"},{"_id":"6374aa8f33bb19fd1cc1856b","avatarUrl":"/avatars/3758bd877ab85542bcca61e11c5d2828.svg","isPro":false,"fullname":"Tangqi Fang","user":"chnftq","type":"user"},{"_id":"653efc4c3ea696d4637b92c8","avatarUrl":"/avatars/8a8f3468dfd55133ce04f17220b84aa5.svg","isPro":true,"fullname":"pc","user":"neocxi","type":"user"},{"_id":"6a1469c343be728859606ebe","avatarUrl":"/avatars/90c45a3ea8667e7abf08fdd701767ca4.svg","isPro":false,"fullname":"郑雨田","user":"masoncampbell","type":"user"},{"_id":"6a6aa38fbbca071c718a29de","avatarUrl":"/avatars/93253032d7014e832fa66c1575e1fb9e.svg","isPro":false,"fullname":"James Johnson","user":"Kestrel-Evan","type":"user"},{"_id":"6a6c8a1109f0af0927dee1c0","avatarUrl":"/avatars/f6e88ed87230495e6a3bda2d6cbd8aef.svg","isPro":false,"fullname":"James Anderson","user":"Harbor-James","type":"user"},{"_id":"6a6dca4ec1e23c5ff6a36aba","avatarUrl":"/avatars/71ed0a47e7c665ce9820aaa414cd2ff9.svg","isPro":false,"fullname":"Jennifer Lopez","user":"cobaltRemy","type":"user"},{"_id":"6a701ed9418ef8ed5ce8852b","avatarUrl":"/avatars/2bfcc94945d474ed6e2bf99e9e8f8095.svg","isPro":false,"fullname":"Thomas White","user":"Nova-Thomas","type":"user"},{"_id":"6a81161b845a39b1ec065bed","avatarUrl":"/avatars/1b1d6cb2b3e37bc7dbbda0b815255995.svg","isPro":false,"fullname":"Patricia Lee","user":"Harbor-Leo","type":"user"},{"_id":"6a9aee5b754f1ace116aeda4","avatarUrl":"/avatars/2c99b357dde5055347a31152703d2f06.svg","isPro":false,"fullname":"Jordan Hunter","user":"meridianwing","type":"user"},{"_id":"6a9b5535a36b6b0317c87099","avatarUrl":"/avatars/6039749bdce1e422ff890d3a8d5fefd4.svg","isPro":false,"fullname":"胡玉梅","user":"jun51","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68d5924a76c6b4bfe8b4ab60","name":"Amazon-FAR","fullname":"Amazon Frontier AI & Robotics","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68251d059667f9f347003874/YPxVjC8rCyexwo42cznvE.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.09143.md","query":{}}">
Papers
arxiv:2609.09143

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

Published on Sep 8
· Submitted by
lst
on Sep 11
Authors:
,

Abstract

Using a controlled autoregressive testbed, the study analyzes task-specific validation losses during multimodal pretraining to evaluate how image tokenizer design affects joint text-image modeling and downstream performance.

Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These evaluations do not fully capture how visual tokens behave when modeled jointly with text. We build a controlled pure-autoregressive testbed and track task-specific validation losses during multimodal continual pretraining across text, image, text-to-image (T2I), and image-to-text (I2T) prediction. We examine how these losses scale and relate to downstream performance, then use them to study multimodal learnability---how well image and text tokens are jointly modeled---and tokenizer design. We find that (1) losses should be analyzed by task, since they exhibit distinct scaling behavior and rank tokenizers differently. (2) The loss--performance relationship depends on the predicted token space: for a fixed tokenizer, T2I and I2T losses correlate with generation quality, but across tokenizers, the T2I loss--performance relationship shifts with the image-token space, whereas I2T loss, computed over a shared text vocabulary, provides a more consistent signal. I2T loss also correlates with both generation and visual understanding performance after supervised finetuning. Using losses as a lens, we show that (3) better reconstruction does not necessarily yield lower task-specific losses or stronger downstream performance, and that (4) image tokenizer choice can affect text modeling under joint optimization. As case studies, we revisit three tokenizer design axes---the discriminator, semantic supervision, and vocabulary size---to examine their effects on joint modeling and downstream performance. Together, our testbed offers a complementary perspective on image tokenizers as visual languages, highlighting their interplay with text in joint multimodal training.

Community

Paper submitter about 18 hours ago

Image tokenizers should be designed and evaluated as visual languages — not just as image compressors.

📄 Paper: https://arxiv.org/abs/2609.09143
💻 Code: https://github.com/amazon-far/Tokenizer_UMM
🌐 Website: https://lst627.github.io/tokenizers_as_visual_languages

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.09143
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.09143 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.09143 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.09143 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers