We introduce Splash (ECCV 2026), a mask-isolated tactile alignment learning framework for MLLMs.<br>Splash partitions the pretrained parameter space into a frozen critical subspace that safeguards general vision-language knowledge and a dormant subspace updated for tactile alignment, enabling non-destructive modality expansion without catastrophic forgetting. Splash achieves state-of-the-art performance on visuo-tactile benchmarks (SSVTP, TVL, TacQuad) with no additional inference overhead, while preserving original general-purpose capabilities.</p>\n","updatedAt":"2026-07-09T13:45:41.275Z","author":{"_id":"69105d44cf26541ba945d8ef","avatarUrl":"/avatars/c39ef81bb63b6434b322ad5b35330066.svg","fullname":"김민지","name":"xxinzzi","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7230032086372375},"editors":["xxinzzi"],"editorAvatarUrls":["/avatars/c39ef81bb63b6434b322ad5b35330066.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.00302","authors":[{"_id":"6a4e732048d70828b718dce3","name":"Yoonhyung Park","hidden":false},{"_id":"6a4e732048d70828b718dce4","user":{"_id":"69105d44cf26541ba945d8ef","avatarUrl":"/avatars/c39ef81bb63b6434b322ad5b35330066.svg","isPro":false,"fullname":"김민지","user":"xxinzzi","type":"user","name":"xxinzzi"},"name":"Minji Kim","status":"claimed_verified","statusLastChangedAt":"2026-07-09T11:53:58.787Z","hidden":false},{"_id":"6a4e732048d70828b718dce5","name":"Sungwon Moon","hidden":false},{"_id":"6a4e732048d70828b718dce6","name":"Jiyoung Lee","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/69105d44cf26541ba945d8ef/tB0Qtqfuq1VmGp8ff5FOB.png"],"publishedAt":"2026-07-01T00:00:00.000Z","submittedOnDailyAt":"2026-07-09T00:00:00.000Z","title":"Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs","submittedOnDailyBy":{"_id":"69105d44cf26541ba945d8ef","avatarUrl":"/avatars/c39ef81bb63b6434b322ad5b35330066.svg","isPro":false,"fullname":"김민지","user":"xxinzzi","type":"user","name":"xxinzzi"},"summary":"Touch supplies the physical grounding needed to perceive intrinsic material properties, such as friction and compliance, that vision alone often cannot resolve. Recent efforts for equipping multimodal LLMs with this tactile sense, however, expose a zero-sum trade-off: the limited parameter budget of compact models forces a choice between acquiring the new sensory modality and preserving the established vision-language reasoning. We present Splash, a mask-isolated tactile alignment learning framework for MLLMs. Splash quantifies the significance of each pretrained parameter, and partitions the parameter space into a dormant and critical subspace. While the frozen critical subspace acts as a stable anchor to safeguard general visual knowledge, Splash updates the isolated dormant subspace to internalize tactile alignment towards LLMs. This selective, non-destructive expansion effectively prevents catastrophic forgetting and ensures non-destructive modality expansion. Extensive experiments show that Splash effectively achieves tactile reasoning without additional inference overhead in the LLM part, demonstrating state-of-the-art performance on visuo-tactile benchmarks, including SSVTP, TVL, and TacQuad, while preserving its original general-purpose capabilities.","upvotes":0,"discussionId":"6a4e732048d70828b718dce7","projectPage":"https://ewha-mmai.github.io/splash/","githubRepo":"https://github.com/ewha-mmai/splash","githubRepoAddedBy":"user","ai_summary":"Splash is a mask-isolated tactile alignment learning framework that enables multimodal LLMs to acquire tactile sensing capabilities without sacrificing vision-language reasoning through selective parameter updating that prevents catastrophic forgetting.","ai_keywords":["multimodal LLMs","tactile sense","parameter-efficient fine-tuning","catastrophic forgetting","mask-isolated tactile alignment learning","pretrained parameters","dormant subspace","critical subspace","visual knowledge preservation","visuo-tactile benchmarks","SSVTP","TVL","TacQuad"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":0,"organization":{"_id":"65f920f81e0c65c13a4254bd","name":"Ewha","fullname":"Ewha Womans University","avatar":"https://www.gravatar.com/avatar/0a41f78b7f450ca4538bcc6197e91e81?d=retro&size=100"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[],"acceptLanguages":["en"],"organization":{"_id":"65f920f81e0c65c13a4254bd","name":"Ewha","fullname":"Ewha Womans University","avatar":"https://www.gravatar.com/avatar/0a41f78b7f450ca4538bcc6197e91e81?d=retro&size=100"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.00302.md","query":{}}">
Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs
Published on Jul 1
· Submitted by 김민지 on Jul 9 Abstract
Splash is a mask-isolated tactile alignment learning framework that enables multimodal LLMs to acquire tactile sensing capabilities without sacrificing vision-language reasoning through selective parameter updating that prevents catastrophic forgetting.
Touch supplies the physical grounding needed to perceive intrinsic material properties, such as friction and compliance, that vision alone often cannot resolve. Recent efforts for equipping multimodal LLMs with this tactile sense, however, expose a zero-sum trade-off: the limited parameter budget of compact models forces a choice between acquiring the new sensory modality and preserving the established vision-language reasoning. We present Splash, a mask-isolated tactile alignment learning framework for MLLMs. Splash quantifies the significance of each pretrained parameter, and partitions the parameter space into a dormant and critical subspace. While the frozen critical subspace acts as a stable anchor to safeguard general visual knowledge, Splash updates the isolated dormant subspace to internalize tactile alignment towards LLMs. This selective, non-destructive expansion effectively prevents catastrophic forgetting and ensures non-destructive modality expansion. Extensive experiments show that Splash effectively achieves tactile reasoning without additional inference overhead in the LLM part, demonstrating state-of-the-art performance on visuo-tactile benchmarks, including SSVTP, TVL, and TacQuad, while preserving its original general-purpose capabilities.
Community
We introduce Splash (ECCV 2026), a mask-isolated tactile alignment learning framework for MLLMs.
Splash partitions the pretrained parameter space into a frozen critical subspace that safeguards general vision-language knowledge and a dormant subspace updated for tactile alignment, enabling non-destructive modality expansion without catastrophic forgetting. Splash achieves state-of-the-art performance on visuo-tactile benchmarks (SSVTP, TVL, TacQuad) with no additional inference overhead, while preserving original general-purpose capabilities.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.00302 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.00302 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.00302 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.