Hugging Face Daily Papers · · 5 min read

Generative Late-Interaction Embeddings For Visual Document Retrieval

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Visual document retrieval typically stores around a thousand vectors per page. GLIE learns a compact code that retrieves pages and supports on-demand regeneration of their embeddings. A small refiner and decoder are trained while the document encoder stays frozen. Search uses the stored codes; in the reported evaluation, only the top 20 candidates are expanded and reranked with MaxSim.</p>\n","updatedAt":"2026-09-11T06:05:14.035Z","author":{"_id":"6331c242e092098b57bd8e58","avatarUrl":"/avatars/8bcaf3cb3482a002ded96d3206b04947.svg","fullname":"Mohamed Eltahir","name":"mohammad2012191","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8377684354782104},"editors":["mohammad2012191"],"editorAvatarUrls":["/avatars/8bcaf3cb3482a002ded96d3206b04947.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.11808","authors":[{"_id":"6aa399db47a406da7901e831","name":"Mohamed Eltahir","hidden":false},{"_id":"6aa399db47a406da7901e832","name":"Talal Aloushan","hidden":false},{"_id":"6aa399db47a406da7901e833","name":"Rose Khairoalsendi","hidden":false},{"_id":"6aa399db47a406da7901e834","name":"Jana Shata","hidden":false},{"_id":"6aa399db47a406da7901e835","name":"Mohammed Alhassan","hidden":false},{"_id":"6aa399db47a406da7901e836","name":"Leen Alrehaili","hidden":false},{"_id":"6aa399db47a406da7901e837","name":"Tanveer Hussain","hidden":false},{"_id":"6aa399db47a406da7901e838","name":"Naeemullah Khan","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6331c242e092098b57bd8e58/xxMXWNZHmbrtioq5ajrQk.png","https://cdn-uploads.huggingface.co/production/uploads/6331c242e092098b57bd8e58/BTMD42iBWGobxVrnGEC0Q.png","https://cdn-uploads.huggingface.co/production/uploads/6331c242e092098b57bd8e58/QyVO6eD9ZU2wfHrQOR2Ut.png","https://cdn-uploads.huggingface.co/production/uploads/6331c242e092098b57bd8e58/VfkjdPO-43vXEQ1wUhh_D.png"],"publishedAt":"2026-09-10T00:00:00.000Z","submittedOnDailyAt":"2026-09-11T00:00:00.000Z","title":"Generative Late-Interaction Embeddings For Visual Document Retrieval","submittedOnDailyBy":{"_id":"6331c242e092098b57bd8e58","avatarUrl":"/avatars/8bcaf3cb3482a002ded96d3206b04947.svg","isPro":false,"fullname":"Mohamed Eltahir","user":"mohammad2012191","type":"user","name":"mohammad2012191"},"summary":"Late-interaction retrieval is the state-of-the-art for visual document search, but it pays for its accuracy in storage. Existing compression methods retain a subset or local average of the N~1,000 vectors per page. Under aggressive storage budgets, however, these methods degrade sharply, and alternatives require retraining the encoder. Investigating this degradation across three encoders, we found two consistent properties: the vectors lie exactly on the unit sphere and concentrate near a manifold of intrinsic dimension five to six. This geometry yields two insights. First, standard k-means centroids fall inside the sphere, causing systematic underestimation of MaxSim scores. Normalizing them to the surface is a free correction worth up to +0.093 nDCG@5 over raw centroids. Second, because the page manifold has few degrees of freedom, the full set of vectors can be regenerated from only a few. To this end, we introduce Generative Late-Interaction Embeddings (GLIE): k << N vectors per page learned from the normalized centroids to serve as both a lightweight index and a basis for regenerating the page's full embedding set. At query time, search runs exclusively on these k vectors, and a decoder expands only the top candidates back to all N vectors for exact rescoring. At four vectors per page on ViDoRe v1, GLIE retains nearly 80% of the uncompressed system's nDCG@5, against 70% for the best prior post-hoc method. These results use a 415K-parameter network fitted in under three GPU-minutes on just a thousand training pages. At a matched training budget, fine-tuning the encoder does not reach even the training-free stage of GLIE, and the full system beats it at every budget. These patterns hold across a second encoder and ViDoRe v2. By reconstructing evidence on demand rather than sampling it, GLIE opens a new axis for storage-efficient retrieval, with the decoder as its main design surface.","upvotes":2,"discussionId":"6aa399dc47a406da7901e839","projectPage":"https://mohammad2012191.github.io/GLIE/","githubRepo":"https://github.com/mohammad2012191/GLIE","githubRepoAddedBy":"user","ai_summary":"Generative Late-Interaction Embeddings compress visual document retrieval vectors by learning a small basis set that regenerates full embeddings on demand, improving accuracy under tight storage limits without retraining the encoder.","ai_keywords":["late-interaction retrieval","MaxSim","k-means","unit sphere","manifold","intrinsic dimension","Generative Late-Interaction Embeddings","GLIE","decoder","nDCG@5","ViDoRe"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"642bf38ba208ae9adcebe075","name":"KAUST","fullname":"King Abdullah University of Science and Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6315fb0b29411a6864b05b35/6egitr9tcwgl5i6ikCbkY.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6331c242e092098b57bd8e58","avatarUrl":"/avatars/8bcaf3cb3482a002ded96d3206b04947.svg","isPro":false,"fullname":"Mohamed Eltahir","user":"mohammad2012191","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"642bf38ba208ae9adcebe075","name":"KAUST","fullname":"King Abdullah University of Science and Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6315fb0b29411a6864b05b35/6egitr9tcwgl5i6ikCbkY.jpeg"},"query":{}}">
Papers
arxiv:2609.11808

Generative Late-Interaction Embeddings For Visual Document Retrieval

Published on Sep 10
· Submitted by
Mohamed Eltahir
on Sep 11
Authors:
,

Abstract

Generative Late-Interaction Embeddings compress visual document retrieval vectors by learning a small basis set that regenerates full embeddings on demand, improving accuracy under tight storage limits without retraining the encoder.

Late-interaction retrieval is the state-of-the-art for visual document search, but it pays for its accuracy in storage. Existing compression methods retain a subset or local average of the N~1,000 vectors per page. Under aggressive storage budgets, however, these methods degrade sharply, and alternatives require retraining the encoder. Investigating this degradation across three encoders, we found two consistent properties: the vectors lie exactly on the unit sphere and concentrate near a manifold of intrinsic dimension five to six. This geometry yields two insights. First, standard k-means centroids fall inside the sphere, causing systematic underestimation of MaxSim scores. Normalizing them to the surface is a free correction worth up to +0.093 nDCG@5 over raw centroids. Second, because the page manifold has few degrees of freedom, the full set of vectors can be regenerated from only a few. To this end, we introduce Generative Late-Interaction Embeddings (GLIE): k << N vectors per page learned from the normalized centroids to serve as both a lightweight index and a basis for regenerating the page's full embedding set. At query time, search runs exclusively on these k vectors, and a decoder expands only the top candidates back to all N vectors for exact rescoring. At four vectors per page on ViDoRe v1, GLIE retains nearly 80% of the uncompressed system's nDCG@5, against 70% for the best prior post-hoc method. These results use a 415K-parameter network fitted in under three GPU-minutes on just a thousand training pages. At a matched training budget, fine-tuning the encoder does not reach even the training-free stage of GLIE, and the full system beats it at every budget. These patterns hold across a second encoder and ViDoRe v2. By reconstructing evidence on demand rather than sampling it, GLIE opens a new axis for storage-efficient retrieval, with the decoder as its main design surface.

Community

Visual document retrieval typically stores around a thousand vectors per page. GLIE learns a compact code that retrieves pages and supports on-demand regeneration of their embeddings. A small refiner and decoder are trained while the document encoder stays frozen. Search uses the stored codes; in the reported evaluation, only the top 20 candidates are expanded and reranked with MaxSim.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.11808 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.11808 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.11808 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers