👀 @uminaty and I are very proud to share NeoMME: a family of 260M and 800M Multimodal-Native Multilingual Encoders trained from scratch!</p>\n<p><em>NeoMME</em> uses one bidirectional Transformer for text tokens and raw image patches, without a pretrained vision tower, text encoder, or text decoder. To demonstrate its downstream capabilities, we fine-tuned for visual document retrieval. On ViDoRe v3, <em>NeoMME</em>-Retriever models are competitive for their size and deliver high encoding throughput on high-resolution images.</p>\n","updatedAt":"2026-09-03T10:27:48.860Z","author":{"_id":"6264f9655f6f2e14d6ac981c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6264f9655f6f2e14d6ac981c/bJ5lPjGaOfatQne--KQ10.jpeg","fullname":"Tony Wu","name":"tonywu71","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":44,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/677d3f355f847864bb644112/OQyAJ33sssiTDIQEQ7oH_.png","fullname":"H company","name":"Hcompany","type":"org","isHf":false,"plan":"team"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8438438177108765},"editors":["tonywu71"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6264f9655f6f2e14d6ac981c/bJ5lPjGaOfatQne--KQ10.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.01657","authors":[{"_id":"6a992e18fea818274322009f","name":"Aurélien Lac","hidden":false},{"_id":"6a992e18fea81827432200a0","name":"Tony Wu","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6264f9655f6f2e14d6ac981c/P897oMaYKjqSTRNhJKRvE.webp","https://cdn-uploads.huggingface.co/production/uploads/6264f9655f6f2e14d6ac981c/MliSwa5XWfDohXgKW5LE1.webp","https://cdn-uploads.huggingface.co/production/uploads/6264f9655f6f2e14d6ac981c/24bKqYVF9gYdmGWBIFA1g.webp","https://cdn-uploads.huggingface.co/production/uploads/6264f9655f6f2e14d6ac981c/Y3rfgfSLMGeyR_gXVbWRT.webp"],"publishedAt":"2026-08-31T00:00:00.000Z","submittedOnDailyAt":"2026-09-03T00:00:00.000Z","title":"NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference","submittedOnDailyBy":{"_id":"6264f9655f6f2e14d6ac981c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6264f9655f6f2e14d6ac981c/bJ5lPjGaOfatQne--KQ10.jpeg","isPro":false,"fullname":"Tony Wu","user":"tonywu71","type":"user","name":"tonywu71"},"summary":"Multimodal models often build on architectures designed for generative vision-language modeling, typically combining separately pretrained vision encoders with causal language models. Visual document retrievers such as ColPali repurpose these models as encoders, carrying over the parameter and compute overhead of a VLM for a non-generative task.\n We introduce NeoMME, a family of 260M and 800M-parameter Multimodal and Multilingual bidirectional Encoders that process multilingual text and raw image patches in a single bidirectional Transformer encoder. Both models are pretrained from scratch with a masked discrete-diffusion text objective, conditioned on visible image patches for multimodal examples. Both support a 16,384-token context, enough to encode up to two standard 4K UHD images.\n To demonstrate its downstream capabilities, we fine-tune NeoMME with jointly trained dense and late-interaction heads. On the ViDoRe v3 benchmark, the resulting NeoMME-Retriever 260M outperforms all evaluated models strictly below 800M parameters with 0.523 nDCG@10, while NeoMME-Retriever 800M reaches 0.556. At a matched 2048x2048 image input size on an NVIDIA L40S, NeoMME-260M encodes pages with about 2x the throughput of ColModernVBERT. Hierarchical token pooling and asymmetric quantization compress late-interaction multimodal document embeddings by 255x while preserving over 95% of baseline nDCG@10. We contribute NeoMME to Hugging Face Transformers and release the pretrained backbone and retrieval-compatible checkpoints under Apache 2.0 at https://hf.co/collections/Hcompany/neomme.","upvotes":18,"discussionId":"6a992e18fea81827432200a1","projectPage":"https://huggingface.co/docs/transformers/main/en/model_doc/neomme","ai_summary":"NeoMME introduces small bidirectional multimodal encoders pretrained with masked discrete diffusion that achieve strong visual document retrieval and high compression of late-interaction embeddings.","ai_keywords":["masked discrete-diffusion","bidirectional Transformer encoder","multimodal multilingual encoder","late-interaction heads","hierarchical token pooling","asymmetric quantization","late-interaction multimodal embeddings","ViDoRe v3","nDCG@10"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"66faa91224f27d02bd96ef70","name":"Hcompany","fullname":"H company","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/677d3f355f847864bb644112/OQyAJ33sssiTDIQEQ7oH_.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69b1cc79bfc0ece1116d8a67","avatarUrl":"/avatars/aa17e2cea733f87cc4bc09896b9bb9e4.svg","isPro":false,"fullname":"Gary De'Snake","user":"GaryDSnake","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"5fda30110e761cf3183cfd5c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1654000939422-5fda30110e761cf3183cfd5c.png","isPro":false,"fullname":"Connor Shorten","user":"CShorten","type":"user"},{"_id":"67e184d667350657a7d68aac","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67e184d667350657a7d68aac/I4e1DvZbdv7zQUqndrRgX.png","isPro":false,"fullname":"Paulo Moura","user":"paulomouraj","type":"user"},{"_id":"68c435c869d3aa0e590b0361","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68c435c869d3aa0e590b0361/RZcEZherMz2CeJo5kuWHh.png","isPro":false,"fullname":"Antoine Bonnet","user":"ABonnetH","type":"user"},{"_id":"6a5a4e5cbebeac8471deb09d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/gCLGNIpWsU7LJC0U9fEZc.png","isPro":false,"fullname":"Enzo Damion","user":"enzo-damion-h","type":"user"},{"_id":"68c7e6b768632ba6d83f6d75","avatarUrl":"/avatars/cb119b58fabbb25ca80c7cc975581a9e.svg","isPro":false,"fullname":"Mathieu Diaz","user":"Diazmathh","type":"user"},{"_id":"6a22cf2952f1c6b6cebbc781","avatarUrl":"/avatars/e15d64d814c3ee65ffc7fe39d82969ea.svg","isPro":false,"fullname":"Vincent Coyette","user":"vincentcoyette","type":"user"},{"_id":"69cc3064a245f2c5f71322db","avatarUrl":"/avatars/c9920b2a634b67720b5f9e72895a2e0c.svg","isPro":false,"fullname":"Emrick Sinitambirivoutin","user":"emricksini-h","type":"user"},{"_id":"6a997d2d8061cc5fd1fafbab","avatarUrl":"/avatars/800cd399e5592fe781f1b5e236b16eb1.svg","isPro":false,"fullname":"Pauline Jouitteau","user":"PJ-H","type":"user"},{"_id":"6981cda5964078d4a588f94c","avatarUrl":"/avatars/8460159705ac1d164708fe37d537843f.svg","isPro":false,"fullname":"Hussenot","user":"GeoffroyRecruit","type":"user"},{"_id":"6377b63b24d97f9f7ec70064","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6377b63b24d97f9f7ec70064/vBl68PAwwp353MOeOahMH.jpeg","isPro":false,"fullname":"Antonio Loison","user":"antonioloison","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"66faa91224f27d02bd96ef70","name":"Hcompany","fullname":"H company","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/677d3f355f847864bb644112/OQyAJ33sssiTDIQEQ7oH_.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.01657.md","query":{}}">
NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference
Abstract
NeoMME introduces small bidirectional multimodal encoders pretrained with masked discrete diffusion that achieve strong visual document retrieval and high compression of late-interaction embeddings.
Multimodal models often build on architectures designed for generative vision-language modeling, typically combining separately pretrained vision encoders with causal language models. Visual document retrievers such as ColPali repurpose these models as encoders, carrying over the parameter and compute overhead of a VLM for a non-generative task.
We introduce NeoMME, a family of 260M and 800M-parameter Multimodal and Multilingual bidirectional Encoders that process multilingual text and raw image patches in a single bidirectional Transformer encoder. Both models are pretrained from scratch with a masked discrete-diffusion text objective, conditioned on visible image patches for multimodal examples. Both support a 16,384-token context, enough to encode up to two standard 4K UHD images.
To demonstrate its downstream capabilities, we fine-tune NeoMME with jointly trained dense and late-interaction heads. On the ViDoRe v3 benchmark, the resulting NeoMME-Retriever 260M outperforms all evaluated models strictly below 800M parameters with 0.523 nDCG@10, while NeoMME-Retriever 800M reaches 0.556. At a matched 2048x2048 image input size on an NVIDIA L40S, NeoMME-260M encodes pages with about 2x the throughput of ColModernVBERT. Hierarchical token pooling and asymmetric quantization compress late-interaction multimodal document embeddings by 255x while preserving over 95% of baseline nDCG@10. We contribute NeoMME to Hugging Face Transformers and release the pretrained backbone and retrieval-compatible checkpoints under Apache 2.0 at https://hf.co/collections/Hcompany/neomme.
Community
👀 @uminaty and I are very proud to share NeoMME: a family of 260M and 800M Multimodal-Native Multilingual Encoders trained from scratch!
NeoMME uses one bidirectional Transformer for text tokens and raw image patches, without a pretrained vision tower, text encoder, or text decoder. To demonstrate its downstream capabilities, we fine-tuned for visual document retrieval. On ViDoRe v3, NeoMME-Retriever models are competitive for their size and deliver high encoding throughput on high-resolution images.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.01657 in a dataset README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.