Hugging Face Daily Papers · · 5 min read

ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<video src=\"https://cdn-uploads.huggingface.co/production/uploads/6828d5f2682a9738b94ef72f/XI0u8hcOC4mwP3WKHW5bM.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\n","updatedAt":"2026-09-08T10:42:51.848Z","author":{"_id":"6828d5f2682a9738b94ef72f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/G8WyQrC5sq6iD2Q8mN1kG.png","fullname":"Chema Garabito","name":"chgara","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6828d5f2682a9738b94ef72f/X6_QZmboVD5aFL4DdRy83.png","fullname":"speridlabs","name":"speridlabs","type":"org","isHf":false,"plan":"team"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6131085157394409},"editors":["chgara"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/G8WyQrC5sq6iD2Q8mN1kG.png"],"reactions":[],"isReport":false}},{"id":"6a9fe96fb08948eac55df2b0","author":{"_id":"6828d5f2682a9738b94ef72f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/G8WyQrC5sq6iD2Q8mN1kG.png","fullname":"Chema Garabito","name":"chgara","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6828d5f2682a9738b94ef72f/X6_QZmboVD5aFL4DdRy83.png","fullname":"speridlabs","name":"speridlabs","type":"org","isHf":false,"plan":"team"}},"createdAt":"2026-09-08T10:54:39.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"🚀 We’re releasing **ENEAS**, a text-promptable method for robust instance tracking and open-concept semantic discovery.\n\nENEAS targets several failure modes we encountered with foundation segmentation models such as SAM 3: target disappearance/re-entry, extreme close-ups, and semantic doppelgängers such as statues, paintings and reflections.\n\n🎥 Demo above\n🌐 Project: https://speridlabs.com/research/eneas\n💻 Code: https://github.com/speridlabs/eneas\n🤗 Demo: https://huggingface.co/spaces/speridlabs/eneas","html":"<p>🚀 We’re releasing <strong>ENEAS</strong>, a text-promptable method for robust instance tracking and open-concept semantic discovery.</p>\n<p>ENEAS targets several failure modes we encountered with foundation segmentation models such as SAM 3: target disappearance/re-entry, extreme close-ups, and semantic doppelgängers such as statues, paintings and reflections.</p>\n<p>🎥 Demo above<br>🌐 Project: <a href=\"https://speridlabs.com/research/eneas\" rel=\"nofollow\">https://speridlabs.com/research/eneas</a><br>💻 Code: <a href=\"https://github.com/speridlabs/eneas\" rel=\"nofollow\">https://github.com/speridlabs/eneas</a><br>🤗 Demo: <a href=\"https://huggingface.co/spaces/speridlabs/eneas\">https://huggingface.co/spaces/speridlabs/eneas</a></p>\n","updatedAt":"2026-09-08T10:54:39.859Z","author":{"_id":"6828d5f2682a9738b94ef72f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/G8WyQrC5sq6iD2Q8mN1kG.png","fullname":"Chema Garabito","name":"chgara","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6828d5f2682a9738b94ef72f/X6_QZmboVD5aFL4DdRy83.png","fullname":"speridlabs","name":"speridlabs","type":"org","isHf":false,"plan":"team"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8491736650466919},"editors":["chgara"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/G8WyQrC5sq6iD2Q8mN1kG.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.03756","authors":[{"_id":"6a9ef34a6c8e10537d563a4b","name":"Javier del Pino","hidden":false},{"_id":"6a9ef34a6c8e10537d563a4c","name":"Salvador Rodríguez","hidden":false},{"_id":"6a9ef34a6c8e10537d563a4d","name":"Alejandro Garabito","hidden":false},{"_id":"6a9ef34a6c8e10537d563a4e","name":"Javier Álvarez","hidden":false},{"_id":"6a9ef34a6c8e10537d563a4f","user":{"_id":"6828d5f2682a9738b94ef72f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/G8WyQrC5sq6iD2Q8mN1kG.png","isPro":false,"fullname":"Chema Garabito","user":"chgara","type":"user","name":"chgara"},"name":"Chema Garabito","status":"claimed_verified","statusLastChangedAt":"2026-09-08T00:45:04.313Z","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6828d5f2682a9738b94ef72f/fsYpwS-A89nawf3jwbzYv.mp4"],"publishedAt":"2026-09-03T00:00:00.000Z","submittedOnDailyAt":"2026-09-08T00:00:00.000Z","title":"ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation","submittedOnDailyBy":{"_id":"6828d5f2682a9738b94ef72f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/G8WyQrC5sq6iD2Q8mN1kG.png","isPro":false,"fullname":"Chema Garabito","user":"chgara","type":"user","name":"chgara"},"summary":"We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absence when an object leaves the field of view, segment local textures instead of the complete object during extreme close-ups, and prioritize visual features over ontological reality, so that visually similar artifacts such as statues, paintings, or reflections are segmented as target entities.\n ENEAS works two ways from a single method: precise tracking and high-quality segmentation of a unique instance, and open-concept discovery of every instance a text query names, resolved by a semantic verification layer. For tracking, we extend the geometrically robust SeC architecture, previously limited to point interactions, with a text-prompting adapter and leverage its temporal memory, so that the target is held through disappearance without drifting to distractors and kept whole even when it fills the entire view. For discovery, the verification layer combines high-speed visual embedding matching with conditional VLM refinement, invoking semantic reasoning only for ambiguous candidates, which filters out the ontological errors that visual-only models cannot distinguish while keeping latency low. Designed with 3D reconstruction in mind, where a single misclassified distractor corrupts the asset, ENEAS unlocks high-quality semantic tracking and segmentation of video, of broad libraries, and of collections of temporally or spatially unordered data, together with the discrimination to tell true instances from their doppelgangers: things that look alike but are not the same. The code and models are available at https://github.com/speridlabs/eneas","upvotes":10,"discussionId":"6a9ef34b6c8e10537d563a50","projectPage":"https://speridlabs.com/research/eneas","githubRepo":"https://github.com/speridlabs/eneas","githubRepoAddedBy":"user","ai_summary":"ENEAS unifies text-prompted instance tracking and open-concept semantic discovery via temporal memory extension and a verification layer that filters visual distractors.","ai_keywords":["text-promptable segmentation","SAM","SeC architecture","text-prompting adapter","temporal memory","semantic verification layer","VLM refinement","visual embedding matching","3D reconstruction"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":41,"organization":{"_id":"6936db1f3eabe013bfe5cf9b","name":"speridlabs","fullname":"speridlabs","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6828d5f2682a9738b94ef72f/X6_QZmboVD5aFL4DdRy83.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6828d5f2682a9738b94ef72f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/G8WyQrC5sq6iD2Q8mN1kG.png","isPro":false,"fullname":"Chema Garabito","user":"chgara","type":"user"},{"_id":"62e0fd5e02a6c13e467411f6","avatarUrl":"/avatars/64bcd6c6fbe6e7c363ffda41f4a378cd.svg","isPro":false,"fullname":"Javier","user":"javipd99","type":"user"},{"_id":"63076cf2cd148dbc5e4ccce7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63076cf2cd148dbc5e4ccce7/kpVPt1TXnpqWNqtMbtRSJ.jpeg","isPro":false,"fullname":"Alejandro Garabito","user":"alexgara","type":"user"},{"_id":"66e6bab86e6ce3af72764f96","avatarUrl":"/avatars/3c6b2bb431c206bf02d94de44f8bdfed.svg","isPro":false,"fullname":"Salvador Rodriguez","user":"salvy9978","type":"user"},{"_id":"64be25334561d0aca2b209d9","avatarUrl":"/avatars/8c2eabe465958acda78b62fc1c5e3a6e.svg","isPro":false,"fullname":"Javier Álvarez","user":"jalvarezz13","type":"user"},{"_id":"6a8594cd200f083f782d9f65","avatarUrl":"/avatars/bacdf9a32672678e55f0b58cb8c4817d.svg","isPro":false,"fullname":"alejandro","user":"acanadaSpl","type":"user"},{"_id":"6a54a66d5cae84d3c8fe30e2","avatarUrl":"/avatars/47a8d9258a951c2e3c80cf5da41538ef.svg","isPro":false,"fullname":"Hanqiu Li Cai","user":"lisperidlabs","type":"user"},{"_id":"6a821520aeecc8b39689ff2b","avatarUrl":"/avatars/2849365947491db3d2c41d50372720ca.svg","isPro":false,"fullname":"Joko M. Sari","user":"jokosari","type":"user"},{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"6936db1f3eabe013bfe5cf9b","name":"speridlabs","fullname":"speridlabs","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6828d5f2682a9738b94ef72f/X6_QZmboVD5aFL4DdRy83.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.03756.md","query":{}}">
Papers
arxiv:2609.03756

ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation

Published on Sep 3
· Submitted by
Chema Garabito
on Sep 8
#3 Paper of the day
Authors:
,

Abstract

ENEAS unifies text-prompted instance tracking and open-concept semantic discovery via temporal memory extension and a verification layer that filters visual distractors.

We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absence when an object leaves the field of view, segment local textures instead of the complete object during extreme close-ups, and prioritize visual features over ontological reality, so that visually similar artifacts such as statues, paintings, or reflections are segmented as target entities. ENEAS works two ways from a single method: precise tracking and high-quality segmentation of a unique instance, and open-concept discovery of every instance a text query names, resolved by a semantic verification layer. For tracking, we extend the geometrically robust SeC architecture, previously limited to point interactions, with a text-prompting adapter and leverage its temporal memory, so that the target is held through disappearance without drifting to distractors and kept whole even when it fills the entire view. For discovery, the verification layer combines high-speed visual embedding matching with conditional VLM refinement, invoking semantic reasoning only for ambiguous candidates, which filters out the ontological errors that visual-only models cannot distinguish while keeping latency low. Designed with 3D reconstruction in mind, where a single misclassified distractor corrupts the asset, ENEAS unlocks high-quality semantic tracking and segmentation of video, of broad libraries, and of collections of temporally or spatially unordered data, together with the discrimination to tell true instances from their doppelgangers: things that look alike but are not the same. The code and models are available at https://github.com/speridlabs/eneas

Community

Paper author Paper submitter about 3 hours ago

Paper author Paper submitter about 3 hours ago

🚀 We’re releasing ENEAS, a text-promptable method for robust instance tracking and open-concept semantic discovery.

ENEAS targets several failure modes we encountered with foundation segmentation models such as SAM 3: target disappearance/re-entry, extreme close-ups, and semantic doppelgängers such as statues, paintings and reflections.

🎥 Demo above
🌐 Project: https://speridlabs.com/research/eneas
💻 Code: https://github.com/speridlabs/eneas
🤗 Demo: https://huggingface.co/spaces/speridlabs/eneas

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.03756
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.03756 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.03756 in a dataset README.md to link it from this page.

Spaces citing this paper

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers