Hugging Face Daily Papers · · 3 min read

UEmbed: Unified Sparse and Dense Multimodal Embeddings

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

UEmbed offers a new paradigm: it unifies dense and sparse embeddings in one model, while further extending sparse retrieval to unify text and multimodal inputs.</p>\n","updatedAt":"2026-08-04T03:31:06.003Z","author":{"_id":"64dc29d9b5d625e0e9a6ecb9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/QxGBsnk1cNsBEPqSx4ae-.jpeg","fullname":"Tingyu Song","name":"songtingyu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7412709593772888},"editors":["songtingyu"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/QxGBsnk1cNsBEPqSx4ae-.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.02583","authors":[{"_id":"6a7158ffec5082b9f872ce37","name":"Tingyu Song","hidden":false},{"_id":"6a7158ffec5082b9f872ce38","name":"Mingxin Li","hidden":false},{"_id":"6a7158ffec5082b9f872ce39","name":"Yanzhao Zhang","hidden":false},{"_id":"6a7158ffec5082b9f872ce3a","name":"Dingkun Long","hidden":false},{"_id":"6a7158ffec5082b9f872ce3b","name":"Pengjun Xie","hidden":false},{"_id":"6a7158ffec5082b9f872ce3c","name":"Zhijie Nie","hidden":false},{"_id":"6a7158ffec5082b9f872ce3d","name":"Yilun Zhao","hidden":false},{"_id":"6a7158ffec5082b9f872ce3e","name":"Shu Wu","hidden":false}],"publishedAt":"2026-08-03T00:00:00.000Z","submittedOnDailyAt":"2026-08-04T00:00:00.000Z","title":"UEmbed: Unified Sparse and Dense Multimodal Embeddings","submittedOnDailyBy":{"_id":"64dc29d9b5d625e0e9a6ecb9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/QxGBsnk1cNsBEPqSx4ae-.jpeg","isPro":false,"fullname":"Tingyu Song","user":"songtingyu","type":"user","name":"songtingyu"},"summary":"Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work has introduced Learned Sparse Retrieval (LSR) to push beyond exact lexical matching toward richer semantics. Yet LSR has so far remained tied to encoder-style bidirectional architectures, and its extension to multimodal settings still relies heavily on auxiliary cross-modal modules. To address these limitations, we introduce UEmbed (Unified Embedding), a decoder-only multimodal embedding model that produces both sparse lexical and dense representations in one causal forward pass. UEmbed appends N learnable special tokens to the input and partitions the vocabulary into N disjoint subsets. Each token's causal hidden state predicts sparse weights over its assigned subset, and the N subsets are concatenated into the full sparse vector. Trained on public data, we release UEmbed at 2B, 4B, and 9B scales. UEmbed-9B reaches 71.8 (dense) and 71.0 (sparse) on MMEB-v2, outperforming multimodal embedding models trained on publicly available data (e.g., RzenEmbed). On BEIR, UEmbed also remains competitive with strong dense and sparse baselines. Furthermore, we demonstrate the practical utility of UEmbed across three dimensions: effectiveness, efficiency, and agentic applications. Overall, UEmbed offers a new paradigm: it unifies dense and sparse embeddings in one model, while further extending sparse retrieval to unify text and multimodal inputs.","upvotes":17,"discussionId":"6a7158ffec5082b9f872ce3f","projectPage":"https://alibaba-nlp.github.io/UEmbed/","githubRepo":"https://github.com/Alibaba-NLP/UEmbed","githubRepoAddedBy":"user","githubStars":2,"organization":{"_id":"661f98de142a51d630dbbcc4","name":"Alibaba-NLP","fullname":"Alibaba-NLP","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63fc4c00a3c067e62899d32b/dfd_EcIfylvu3sdc2WMqX.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"62f662bcc58915315c4eccea","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62f662bcc58915315c4eccea/zOAQLONfMP88zr70sxHK-.jpeg","isPro":true,"fullname":"Yilun Zhao","user":"yilunzhao","type":"user"},{"_id":"64dc29d9b5d625e0e9a6ecb9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/QxGBsnk1cNsBEPqSx4ae-.jpeg","isPro":false,"fullname":"Tingyu Song","user":"songtingyu","type":"user"},{"_id":"6434c530a5aed21dd119a393","avatarUrl":"/avatars/a5aa3d8dc8b3b987e6de39280a4c0765.svg","isPro":false,"fullname":"Bro","user":"H34lthy","type":"user"},{"_id":"68084d54aca60e6178b3afb5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68084d54aca60e6178b3afb5/TshN3Ka3VRFD_I3WJ6Vys.jpeg","isPro":false,"fullname":"Lin Fu","user":"minuzero","type":"user"},{"_id":"65dfeee3d16fb170031df293","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65dfeee3d16fb170031df293/2VbNuqcpN3XrWB18NfzRQ.jpeg","isPro":false,"fullname":"gan","user":"guo9","type":"user"},{"_id":"683c642b02c1a474a867964e","avatarUrl":"/avatars/63e44a9cf788ee7b3ad236407700ceca.svg","isPro":false,"fullname":"Jinbiao Wei","user":"mikeweii","type":"user"},{"_id":"6900b7a8ecfa14b7b16368fb","avatarUrl":"/avatars/183456c0fa0d5e21c27813ce48d3bad2.svg","isPro":false,"fullname":"Sentinel","user":"Sentinel7","type":"user"},{"_id":"66e258bdc70c02b46dfed6e3","avatarUrl":"/avatars/ccc2d604616c018f45a268a610472cac.svg","isPro":false,"fullname":"Yuzheng Cai","user":"Ucreate","type":"user"},{"_id":"64b16ffa5c1ffb0870f81074","avatarUrl":"/avatars/97a27ea102fe4818a909f1b5d94623ef.svg","isPro":false,"fullname":"Jiawei","user":"jzhoubu","type":"user"},{"_id":"6a6c0b0eee5cdeaf4af82f83","avatarUrl":"/avatars/a382a02b896a9a4b3cfcf8630f6bdff8.svg","isPro":false,"fullname":"Jun Li","user":"lijun2005","type":"user"},{"_id":"63c1699e40a26dd2db32400d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63c1699e40a26dd2db32400d/3N0-Zp8igv8-52mXAdiiq.jpeg","isPro":false,"fullname":"Chroma","user":"Chroma111","type":"user"},{"_id":"616adb8578833ce5997e441a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/616adb8578833ce5997e441a/Oy5YkTaCMq-FEc0k3FTMd.jpeg","isPro":false,"fullname":"Dingkun Long","user":"thenlper","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"661f98de142a51d630dbbcc4","name":"Alibaba-NLP","fullname":"Alibaba-NLP","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63fc4c00a3c067e62899d32b/dfd_EcIfylvu3sdc2WMqX.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.02583.md","query":{}}">
Papers
arxiv:2608.02583

UEmbed: Unified Sparse and Dense Multimodal Embeddings

Published on Aug 3
· Submitted by
Tingyu Song
on Aug 4
Authors:
,

Abstract

Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work has introduced Learned Sparse Retrieval (LSR) to push beyond exact lexical matching toward richer semantics. Yet LSR has so far remained tied to encoder-style bidirectional architectures, and its extension to multimodal settings still relies heavily on auxiliary cross-modal modules. To address these limitations, we introduce UEmbed (Unified Embedding), a decoder-only multimodal embedding model that produces both sparse lexical and dense representations in one causal forward pass. UEmbed appends N learnable special tokens to the input and partitions the vocabulary into N disjoint subsets. Each token's causal hidden state predicts sparse weights over its assigned subset, and the N subsets are concatenated into the full sparse vector. Trained on public data, we release UEmbed at 2B, 4B, and 9B scales. UEmbed-9B reaches 71.8 (dense) and 71.0 (sparse) on MMEB-v2, outperforming multimodal embedding models trained on publicly available data (e.g., RzenEmbed). On BEIR, UEmbed also remains competitive with strong dense and sparse baselines. Furthermore, we demonstrate the practical utility of UEmbed across three dimensions: effectiveness, efficiency, and agentic applications. Overall, UEmbed offers a new paradigm: it unifies dense and sparse embeddings in one model, while further extending sparse retrieval to unify text and multimodal inputs.

Community

Paper submitter about 5 hours ago

UEmbed offers a new paradigm: it unifies dense and sparse embeddings in one model, while further extending sparse retrieval to unify text and multimodal inputs.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.02583
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.02583 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.02583 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers