Hugging Face Daily Papers · · 3 min read

Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Accepted for publication at ECCV 2026</p>\n","updatedAt":"2026-07-02T14:45:49.872Z","author":{"_id":"63595b637d959cab6329a127","avatarUrl":"/avatars/1fd6057391a5d78d7973b6b36518469e.svg","fullname":"Jongoh Jeong","name":"jeong2","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.897340714931488},"editors":["jeong2"],"editorAvatarUrls":["/avatars/1fd6057391a5d78d7973b6b36518469e.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2606.29464","authors":[{"_id":"6a467984350f1bd0175de0b1","name":"Jongoh Jeong","hidden":false},{"_id":"6a467984350f1bd0175de0b2","name":"Sun-Kyung Lee","hidden":false},{"_id":"6a467984350f1bd0175de0b3","name":"Kuk-Jin Yoon","hidden":false}],"publishedAt":"2026-06-28T00:00:00.000Z","submittedOnDailyAt":"2026-07-02T00:00:00.000Z","title":"Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation","submittedOnDailyBy":{"_id":"63595b637d959cab6329a127","avatarUrl":"/avatars/1fd6057391a5d78d7973b6b36518469e.svg","isPro":false,"fullname":"Jongoh Jeong","user":"jeong2","type":"user","name":"jeong2"},"summary":"Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of synthetic pairs that can efficiently train contrastive vision-language models under strict data and compute budgets. Most existing methods match expert trajectories or cross-modal statistics, yet still enforce full-dimensional alignment in a Euclidean embedding space. This is often overly restrictive due to rank-deficient image--text correlation, with shared semantics concentrated in a low-dimensional range and remaining variation spread across a weakly correlated residual subspace. LoRS relaxes alignment at the similarity level by low-rank factorization, but does not explicitly control dominant alignment capacity and structure in the representation space. We thus propose a rank-aware hyperbolic alignment (RAHA) that combines hierarchical geometry with explicit alignment-capacity control. RAHA lifts multimodal representations to hyperbolic space and optimizes distilled pairs with asymmetric objectives that enforce geodesic alignment in the shared range while regularizing the residual subspace to preserve modality-private diversity and improve transfer robustness. Experiments on benchmarks show that RAHA demonstrates competitive cross-modal retrieval and improved transfer indicators under fixed budgets.","upvotes":3,"discussionId":"6a467984350f1bd0175de0b4","projectPage":"https://andyj1.github.io/raha","githubRepo":"https://github.com/andyj1/raha","githubRepoAddedBy":"user","ai_summary":"Vision-language dataset distillation method using rank-aware hyperbolic alignment to optimize synthetic image-text pairs for efficient contrastive model training while preserving modality-specific diversity.","ai_keywords":["vision-language dataset distillation","contrastive vision-language models","data distillation","low-rank factorization","hyperbolic space","geodesic alignment","multimodal representations","cross-modal retrieval","transfer robustness"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63595b637d959cab6329a127","avatarUrl":"/avatars/1fd6057391a5d78d7973b6b36518469e.svg","isPro":false,"fullname":"Jongoh Jeong","user":"jeong2","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"66935bdc5489e4f73c76bc7b","avatarUrl":"/avatars/129d1e86bbaf764b507501f4feb177db.svg","isPro":false,"fullname":"Abidoye Aanuoluwapo","user":"Aanuoluwapo65","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2606/2606.29464.md","query":{}}">
Papers
arxiv:2606.29464

Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation

Published on Jun 28
· Submitted by
Jongoh Jeong
on Jul 2
Authors:
,
,

Abstract

Vision-language dataset distillation method using rank-aware hyperbolic alignment to optimize synthetic image-text pairs for efficient contrastive model training while preserving modality-specific diversity.

Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of synthetic pairs that can efficiently train contrastive vision-language models under strict data and compute budgets. Most existing methods match expert trajectories or cross-modal statistics, yet still enforce full-dimensional alignment in a Euclidean embedding space. This is often overly restrictive due to rank-deficient image--text correlation, with shared semantics concentrated in a low-dimensional range and remaining variation spread across a weakly correlated residual subspace. LoRS relaxes alignment at the similarity level by low-rank factorization, but does not explicitly control dominant alignment capacity and structure in the representation space. We thus propose a rank-aware hyperbolic alignment (RAHA) that combines hierarchical geometry with explicit alignment-capacity control. RAHA lifts multimodal representations to hyperbolic space and optimizes distilled pairs with asymmetric objectives that enforce geodesic alignment in the shared range while regularizing the residual subspace to preserve modality-private diversity and improve transfer robustness. Experiments on benchmarks show that RAHA demonstrates competitive cross-modal retrieval and improved transfer indicators under fixed budgets.

Community

Paper submitter about 10 hours ago

Accepted for publication at ECCV 2026

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2606.29464
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2606.29464 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2606.29464 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2606.29464 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers