Hugging Face Daily Papers · · 5 min read

N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Hi all, author here 👋</p>\n<p>We're excited to share N0-VTLA, a vision–tactile–language–action (VTLA) foundation model built for two things current VLA backbones struggle with: fine-grained contact-rich manipulation with real tactile feedback control, and offline policy improvement from data you've already collected during deployment.</p>\n<p>The recipe has three parts:</p>\n<ul>\n<li>Visuo-tactile pre-training on NeoData, our large-scale visuo-tactile robot dataset. To our knowledge this makes N0-VTLA the first VTLA model pre-trained on tactile data at scale.</li>\n<li>Staged tactile-pathway integration in post-training, a predictive tactile pathway that distills the contact priors learned at scale into fine motion adjustments for downstream tactile-centric tasks.</li>\n<li>[ALTER], an advantage-conditioned offline RL method that turns relative progress and trajectory-event comparisons into binary advantage labels, so a fixed deployment corpus can keep improving the policy.</li>\n</ul>\n<p>Results: N0-VTLA wins all nine real-robot NeoReal tasks, and reaches 63.8% mean success on our 20-task simulation suite vs. 44.0% for the strongest baseline. With [ALTER], policies hit 75–95% success on three long-horizon real-robot tasks, including deformable object manipulation.</p>\n<p>Happy to answer questions here — feedback very welcome!</p>\n","updatedAt":"2026-08-03T01:47:18.437Z","author":{"_id":"660d17d6c9be0dcd31a30b3d","avatarUrl":"/avatars/3743fe9b695c488ebe33f0d8fd607a8a.svg","fullname":"Zhou Heng","name":"henggg","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8648383617401123},"editors":["henggg"],"editorAvatarUrls":["/avatars/3743fe9b695c488ebe33f0d8fd607a8a.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.23782","authors":[{"_id":"6a6fef6abbe824e6bcc4659b","name":"NeoteAI Team","hidden":false},{"_id":"6a6fef6abbe824e6bcc4659c","name":"Fudan TEAI Team","hidden":false}],"publishedAt":"2026-07-26T00:00:00.000Z","submittedOnDailyAt":"2026-08-03T00:00:00.000Z","title":"N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens","submittedOnDailyBy":{"_id":"660d17d6c9be0dcd31a30b3d","avatarUrl":"/avatars/3743fe9b695c488ebe33f0d8fd607a8a.svg","isPro":false,"fullname":"Zhou Heng","user":"henggg","type":"user","name":"henggg"},"summary":"We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current vision-based backbones, we propose a training recipe for tactile integration consisting of visuo-tactile pre-training, staged tactile-pathway integration, and advantage-conditioned offline policy improvement. During pre-training, the policy learns broad contact priors from NeoData, our large-scale visuo-tactile robot dataset; to our knowledge, N_0-VTLA is the first VTLA model pretrained on tactile data at scale. During post-training, we augment the policy with a predictive tactile pathway that distills the contact patterns learned at scale into the fine motion adjustments required by downstream tactile-centric manipulation. For offline policy improvement, we introduce ALTER, an advantage-conditioned offline reinforcement learning method that converts relative progress and trajectory-event comparisons into binary advantage labels for policy training on a fixed deployment corpus, further improving task-specific learning on contact-rich skills such as deformable object manipulation. Across contact-rich benchmarks, N_0-VTLA outperforms strong baselines by wide margins: it wins all nine real-robot NeoReal tasks and reaches 63.8% mean success on a twenty-task simulation suite, against 44.0% for the strongest baseline. N_0-VTLA policies trained with ALTER reach 75-95% success on three long-horizon real-robot tasks. These results lay a foundation for versatile tactile-driven manipulation policies.","upvotes":32,"discussionId":"6a6fef6abbe824e6bcc4659d","projectPage":"https://research.neoteai.com/n0-vtla/","githubRepo":"https://github.com/neoteai/N0-VTLA","githubRepoAddedBy":"user","githubStars":31,"organization":{"_id":"6a28f192fff7a3f4f2589b29","name":"NeoteAIEmbodied","fullname":"NeoteAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a27afdfd205b09ba5dce236/UsQK_CJO_GwMdMcVp8g2f.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"660d17d6c9be0dcd31a30b3d","avatarUrl":"/avatars/3743fe9b695c488ebe33f0d8fd607a8a.svg","isPro":false,"fullname":"Zhou Heng","user":"henggg","type":"user"},{"_id":"64858110b28e374571bbc550","avatarUrl":"/avatars/62c169a3149d5a7942c960753df3a673.svg","isPro":false,"fullname":"Mi Boyu","user":"miboyu5","type":"user"},{"_id":"67dd16495b97d6983787693c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/4KptoRtQEguHAmAjsD0va.png","isPro":false,"fullname":"martelzhang","user":"martelzhang","type":"user"},{"_id":"64cdf8230fbfb00b91225087","avatarUrl":"/avatars/5379727f55782e146302ed20c7661932.svg","isPro":false,"fullname":"Xiufeng Song","user":"sparklexfantasy","type":"user"},{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"},{"_id":"658a6c1399ed106ac8c822b1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/658a6c1399ed106ac8c822b1/Wk2KXCcK39rUvXx6mpmGD.jpeg","isPro":false,"fullname":"yiranqin","user":"IranQin","type":"user"},{"_id":"6463554dd2044cd1d7c6e0bf","avatarUrl":"/avatars/d7653623117268c545a7063fec69664b.svg","isPro":false,"fullname":"Bingzheng Wei","user":"Bingzheng","type":"user"},{"_id":"691843fe670e3c0c9f7d4c56","avatarUrl":"/avatars/25bda4a3d7364b4f9bec83370bfda31f.svg","isPro":false,"fullname":"Liu Qi","user":"apulupai","type":"user"},{"_id":"6943ed2de4a7bdfd0a56b7d7","avatarUrl":"/avatars/f4e03233132a63c193543bab738dcbb9.svg","isPro":false,"fullname":"Yuchen Fan","user":"yuchenfan49","type":"user"},{"_id":"6a69eb62c5e36f5d1b3c77e4","avatarUrl":"/avatars/8481802389ca8628ac440d9b926ca9ac.svg","isPro":false,"fullname":"William Lopez","user":"johnny-2859013","type":"user"},{"_id":"6a6a93d2bbca071c7189619a","avatarUrl":"/avatars/2d133bd4af457635809bf72d2c8ca3b4.svg","isPro":false,"fullname":"Mary Hernandez","user":"patrick-7343686","type":"user"},{"_id":"6a6aa1c6557526be22f38a09","avatarUrl":"/avatars/8677177172aaef08cbc5b35680dce73a.svg","isPro":false,"fullname":"Steven Hernandez","user":"jessica-0290891","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"6a28f192fff7a3f4f2589b29","name":"NeoteAIEmbodied","fullname":"NeoteAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a27afdfd205b09ba5dce236/UsQK_CJO_GwMdMcVp8g2f.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.23782.md","query":{}}">
Papers
arxiv:2607.23782

N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

Published on Jul 26
· Submitted by
Zhou Heng
on Aug 3
#2 Paper of the day
Authors:
,

Abstract

We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current vision-based backbones, we propose a training recipe for tactile integration consisting of visuo-tactile pre-training, staged tactile-pathway integration, and advantage-conditioned offline policy improvement. During pre-training, the policy learns broad contact priors from NeoData, our large-scale visuo-tactile robot dataset; to our knowledge, N_0-VTLA is the first VTLA model pretrained on tactile data at scale. During post-training, we augment the policy with a predictive tactile pathway that distills the contact patterns learned at scale into the fine motion adjustments required by downstream tactile-centric manipulation. For offline policy improvement, we introduce ALTER, an advantage-conditioned offline reinforcement learning method that converts relative progress and trajectory-event comparisons into binary advantage labels for policy training on a fixed deployment corpus, further improving task-specific learning on contact-rich skills such as deformable object manipulation. Across contact-rich benchmarks, N_0-VTLA outperforms strong baselines by wide margins: it wins all nine real-robot NeoReal tasks and reaches 63.8% mean success on a twenty-task simulation suite, against 44.0% for the strongest baseline. N_0-VTLA policies trained with ALTER reach 75-95% success on three long-horizon real-robot tasks. These results lay a foundation for versatile tactile-driven manipulation policies.

Community

Paper submitter about 6 hours ago

Hi all, author here 👋

We're excited to share N0-VTLA, a vision–tactile–language–action (VTLA) foundation model built for two things current VLA backbones struggle with: fine-grained contact-rich manipulation with real tactile feedback control, and offline policy improvement from data you've already collected during deployment.

The recipe has three parts:

  • Visuo-tactile pre-training on NeoData, our large-scale visuo-tactile robot dataset. To our knowledge this makes N0-VTLA the first VTLA model pre-trained on tactile data at scale.
  • Staged tactile-pathway integration in post-training, a predictive tactile pathway that distills the contact priors learned at scale into fine motion adjustments for downstream tactile-centric tasks.
  • [ALTER], an advantage-conditioned offline RL method that turns relative progress and trajectory-event comparisons into binary advantage labels, so a fixed deployment corpus can keep improving the policy.

Results: N0-VTLA wins all nine real-robot NeoReal tasks, and reaches 63.8% mean success on our 20-task simulation suite vs. 44.0% for the strongest baseline. With [ALTER], policies hit 75–95% success on three long-horizon real-robot tasks, including deformable object manipulation.

Happy to answer questions here — feedback very welcome!

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.23782
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.23782 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.23782 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.23782 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers