<video src=\"https://cdn-uploads.huggingface.co/production/uploads/6828d5f2682a9738b94ef72f/XI0u8hcOC4mwP3WKHW5bM.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\n","updatedAt":"2026-09-08T10:42:51.848Z","author":{"_id":"6828d5f2682a9738b94ef72f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/G8WyQrC5sq6iD2Q8mN1kG.png","fullname":"Chema Garabito","name":"chgara","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6828d5f2682a9738b94ef72f/X6_QZmboVD5aFL4DdRy83.png","fullname":"speridlabs","name":"speridlabs","type":"org","isHf":false,"plan":"team"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6131085157394409},"editors":["chgara"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/G8WyQrC5sq6iD2Q8mN1kG.png"],"reactions":[],"isReport":false}},{"id":"6a9fe96fb08948eac55df2b0","author":{"_id":"6828d5f2682a9738b94ef72f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/G8WyQrC5sq6iD2Q8mN1kG.png","fullname":"Chema Garabito","name":"chgara","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6828d5f2682a9738b94ef72f/X6_QZmboVD5aFL4DdRy83.png","fullname":"speridlabs","name":"speridlabs","type":"org","isHf":false,"plan":"team"}},"createdAt":"2026-09-08T10:54:39.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"🚀 We’re releasing **ENEAS**, a text-promptable method for robust instance tracking and open-concept semantic discovery.\n\nENEAS targets several failure modes we encountered with foundation segmentation models such as SAM 3: target disappearance/re-entry, extreme close-ups, and semantic doppelgängers such as statues, paintings and reflections.\n\n🎥 Demo above\n🌐 Project: https://speridlabs.com/research/eneas\n💻 Code: https://github.com/speridlabs/eneas\n🤗 Demo: https://huggingface.co/spaces/speridlabs/eneas","html":"<p>🚀 We’re releasing <strong>ENEAS</strong>, a text-promptable method for robust instance tracking and open-concept semantic discovery.</p>\n<p>ENEAS targets several failure modes we encountered with foundation segmentation models such as SAM 3: target disappearance/re-entry, extreme close-ups, and semantic doppelgängers such as statues, paintings and reflections.</p>\n<p>🎥 Demo above<br>🌐 Project: <a href=\"https://speridlabs.com/research/eneas\" rel=\"nofollow\">https://speridlabs.com/research/eneas</a><br>💻 Code: <a href=\"https://github.com/speridlabs/eneas\" rel=\"nofollow\">https://github.com/speridlabs/eneas</a><br>🤗 Demo: <a href=\"https://huggingface.co/spaces/speridlabs/eneas\">https://huggingface.co/spaces/speridlabs/eneas</a></p>\n","updatedAt":"2026-09-08T10:54:39.859Z","author":{"_id":"6828d5f2682a9738b94ef72f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/G8WyQrC5sq6iD2Q8mN1kG.png","fullname":"Chema Garabito","name":"chgara","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6828d5f2682a9738b94ef72f/X6_QZmboVD5aFL4DdRy83.png","fullname":"speridlabs","name":"speridlabs","type":"org","isHf":false,"plan":"team"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8491736650466919},"editors":["chgara"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/G8WyQrC5sq6iD2Q8mN1kG.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.03756","authors":[{"_id":"6a9ef34a6c8e10537d563a4b","name":"Javier del Pino","hidden":false},{"_id":"6a9ef34a6c8e10537d563a4c","name":"Salvador Rodríguez","hidden":false},{"_id":"6a9ef34a6c8e10537d563a4d","name":"Alejandro Garabito","hidden":false},{"_id":"6a9ef34a6c8e10537d563a4e","name":"Javier Álvarez","hidden":false},{"_id":"6a9ef34a6c8e10537d563a4f","user":{"_id":"6828d5f2682a9738b94ef72f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/G8WyQrC5sq6iD2Q8mN1kG.png","isPro":false,"fullname":"Chema Garabito","user":"chgara","type":"user","name":"chgara"},"name":"Chema Garabito","status":"claimed_verified","statusLastChangedAt":"2026-09-08T00:45:04.313Z","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6828d5f2682a9738b94ef72f/fsYpwS-A89nawf3jwbzYv.mp4"],"publishedAt":"2026-09-03T00:00:00.000Z","submittedOnDailyAt":"2026-09-08T00:00:00.000Z","title":"ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation","submittedOnDailyBy":{"_id":"6828d5f2682a9738b94ef72f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/G8WyQrC5sq6iD2Q8mN1kG.png","isPro":false,"fullname":"Chema Garabito","user":"chgara","type":"user","name":"chgara"},"summary":"We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absence when an object leaves the field of view, segment local textures instead of the complete object during extreme close-ups, and prioritize visual features over ontological reality, so that visually similar artifacts such as statues, paintings, or reflections are segmented as target entities.\n ENEAS works two ways from a single method: precise tracking and high-quality segmentation of a unique instance, and open-concept discovery of every instance a text query names, resolved by a semantic verification layer. For tracking, we extend the geometrically robust SeC architecture, previously limited to point interactions, with a text-prompting adapter and leverage its temporal memory, so that the target is held through disappearance without drifting to distractors and kept whole even when it fills the entire view. For discovery, the verification layer combines high-speed visual embedding matching with conditional VLM refinement, invoking semantic reasoning only for ambiguous candidates, which filters out the ontological errors that visual-only models cannot distinguish while keeping latency low. Designed with 3D reconstruction in mind, where a single misclassified distractor corrupts the asset, ENEAS unlocks high-quality semantic tracking and segmentation of video, of broad libraries, and of collections of temporally or spatially unordered data, together with the discrimination to tell true instances from their doppelgangers: things that look alike but are not the same. The code and models are available at https://github.com/speridlabs/eneas","upvotes":10,"discussionId":"6a9ef34b6c8e10537d563a50","projectPage":"https://speridlabs.com/research/eneas","githubRepo":"https://github.com/speridlabs/eneas","githubRepoAddedBy":"user","ai_summary":"ENEAS unifies text-prompted instance tracking and open-concept semantic discovery via temporal memory extension and a verification layer that filters visual distractors.","ai_keywords":["text-promptable segmentation","SAM","SeC architecture","text-prompting adapter","temporal memory","semantic verification layer","VLM refinement","visual embedding matching","3D reconstruction"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":41,"organization":{"_id":"6936db1f3eabe013bfe5cf9b","name":"speridlabs","fullname":"speridlabs","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6828d5f2682a9738b94ef72f/X6_QZmboVD5aFL4DdRy83.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6828d5f2682a9738b94ef72f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/G8WyQrC5sq6iD2Q8mN1kG.png","isPro":false,"fullname":"Chema Garabito","user":"chgara","type":"user"},{"_id":"62e0fd5e02a6c13e467411f6","avatarUrl":"/avatars/64bcd6c6fbe6e7c363ffda41f4a378cd.svg","isPro":false,"fullname":"Javier","user":"javipd99","type":"user"},{"_id":"63076cf2cd148dbc5e4ccce7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63076cf2cd148dbc5e4ccce7/kpVPt1TXnpqWNqtMbtRSJ.jpeg","isPro":false,"fullname":"Alejandro Garabito","user":"alexgara","type":"user"},{"_id":"66e6bab86e6ce3af72764f96","avatarUrl":"/avatars/3c6b2bb431c206bf02d94de44f8bdfed.svg","isPro":false,"fullname":"Salvador Rodriguez","user":"salvy9978","type":"user"},{"_id":"64be25334561d0aca2b209d9","avatarUrl":"/avatars/8c2eabe465958acda78b62fc1c5e3a6e.svg","isPro":false,"fullname":"Javier Álvarez","user":"jalvarezz13","type":"user"},{"_id":"6a8594cd200f083f782d9f65","avatarUrl":"/avatars/bacdf9a32672678e55f0b58cb8c4817d.svg","isPro":false,"fullname":"alejandro","user":"acanadaSpl","type":"user"},{"_id":"6a54a66d5cae84d3c8fe30e2","avatarUrl":"/avatars/47a8d9258a951c2e3c80cf5da41538ef.svg","isPro":false,"fullname":"Hanqiu Li Cai","user":"lisperidlabs","type":"user"},{"_id":"6a821520aeecc8b39689ff2b","avatarUrl":"/avatars/2849365947491db3d2c41d50372720ca.svg","isPro":false,"fullname":"Joko M. Sari","user":"jokosari","type":"user"},{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"6936db1f3eabe013bfe5cf9b","name":"speridlabs","fullname":"speridlabs","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6828d5f2682a9738b94ef72f/X6_QZmboVD5aFL4DdRy83.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.03756.md","query":{}}">
ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation
Abstract
ENEAS unifies text-prompted instance tracking and open-concept semantic discovery via temporal memory extension and a verification layer that filters visual distractors.
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absence when an object leaves the field of view, segment local textures instead of the complete object during extreme close-ups, and prioritize visual features over ontological reality, so that visually similar artifacts such as statues, paintings, or reflections are segmented as target entities.
ENEAS works two ways from a single method: precise tracking and high-quality segmentation of a unique instance, and open-concept discovery of every instance a text query names, resolved by a semantic verification layer. For tracking, we extend the geometrically robust SeC architecture, previously limited to point interactions, with a text-prompting adapter and leverage its temporal memory, so that the target is held through disappearance without drifting to distractors and kept whole even when it fills the entire view. For discovery, the verification layer combines high-speed visual embedding matching with conditional VLM refinement, invoking semantic reasoning only for ambiguous candidates, which filters out the ontological errors that visual-only models cannot distinguish while keeping latency low. Designed with 3D reconstruction in mind, where a single misclassified distractor corrupts the asset, ENEAS unlocks high-quality semantic tracking and segmentation of video, of broad libraries, and of collections of temporally or spatially unordered data, together with the discrimination to tell true instances from their doppelgangers: things that look alike but are not the same. The code and models are available at https://github.com/speridlabs/eneas
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.03756 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.03756 in a dataset README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.