Hugging Face Daily Papers · · 4 min read

RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Model weights and code are available at<br>Github: <a href=\"https://github.com/alibaba-damo-academy/RynnValue\" rel=\"nofollow\">https://github.com/alibaba-damo-academy/RynnValue</a><br>HuggingFace: <a href=\"https://huggingface.co/collections/Alibaba-DAMO-Academy/rynnvalue\">https://huggingface.co/collections/Alibaba-DAMO-Academy/rynnvalue</a><br>Modelscope: <a href=\"https://www.modelscope.cn/collections/DAMO_Academy/RynnValue\" rel=\"nofollow\">https://www.modelscope.cn/collections/DAMO_Academy/RynnValue</a></p>\n","updatedAt":"2026-08-11T04:29:29.733Z","author":{"_id":"65fd82762bf2cd20ddaa193f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/yBYbWp_mT7UusYdkqtAvw.png","fullname":"Siteng Huang","name":"huangsiteng","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":12,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6808e64de5dd22427c006e10/9J3vdB62CdeTOd_YrGh9w.jpeg","fullname":"DAMO Academy","name":"Alibaba-DAMO-Academy","type":"org","isHf":false}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7525149583816528},"editors":["huangsiteng"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/yBYbWp_mT7UusYdkqtAvw.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.09853","authors":[{"_id":"6a7aa442019ce76dc7b3aae1","name":"Dongchi Huang","hidden":false},{"_id":"6a7aa442019ce76dc7b3aae2","name":"Hongyin Zhang","hidden":false},{"_id":"6a7aa442019ce76dc7b3aae3","name":"Bohan Hou","hidden":false},{"_id":"6a7aa442019ce76dc7b3aae4","name":"Siteng Huang","hidden":false},{"_id":"6a7aa442019ce76dc7b3aae5","name":"Zhian Su","hidden":false},{"_id":"6a7aa442019ce76dc7b3aae6","name":"Hang Guo","hidden":false},{"_id":"6a7aa442019ce76dc7b3aae7","name":"Tong Lu","hidden":false},{"_id":"6a7aa442019ce76dc7b3aae8","name":"Zhaofeng Xu","hidden":false},{"_id":"6a7aa442019ce76dc7b3aae9","name":"Jiahao Tang","hidden":false},{"_id":"6a7aa442019ce76dc7b3aaea","name":"Jianfei Yang","hidden":false},{"_id":"6a7aa442019ce76dc7b3aaeb","name":"Donglin Wang","hidden":false},{"_id":"6a7aa442019ce76dc7b3aaec","name":"Peixi Peng","hidden":false},{"_id":"6a7aa442019ce76dc7b3aaed","name":"Mingxiu Chen","hidden":false},{"_id":"6a7aa442019ce76dc7b3aaee","name":"Deli Zhao","hidden":false},{"_id":"6a7aa442019ce76dc7b3aaef","name":"Xin Li","hidden":false}],"publishedAt":"2026-08-10T00:00:00.000Z","submittedOnDailyAt":"2026-08-11T00:00:00.000Z","title":"RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance","submittedOnDailyBy":{"_id":"65fd82762bf2cd20ddaa193f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/yBYbWp_mT7UusYdkqtAvw.png","isPro":false,"fullname":"Siteng Huang","user":"huangsiteng","type":"user","name":"huangsiteng"},"summary":"General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing approaches tie supervision to task-internal anchors such as preferences or normalized progress, none of which transfer cleanly across embodiments and data sources. We introduce RynnValue, an open-source value foundation model for robotic manipulation that replaces these anchors with temporal distance, the directed cost-to-go from an observation to the language-specified goal. Because temporal-distance labels can be derived directly from timestamps, RynnValue scales to over 7,000 hours and roughly 3M instruction-conditioned clips without preference or progress annotations. To make temporal-value learning reliable at scale, we combine random temporal sampling, temporal-order shuffling, and value-isolation attention, suppressing shortcuts that would leave predictions insensitive to failures and regressions. Trained without preference labels, RynnValue attains an average Kendall's tau_a of 0.675 on RBM-EVAL-OOD, surpassing the fully preference-supervised state of the art (0.655) and more than doubling a progress-only counterpart (0.292), while generalizing zero-shot to unseen tasks, embodiments, and viewpoints. Converted into dense rewards via potential-based shaping, it raises real-world policy success from 52.5% to 72.5% online and from 63.8% to 82.5% offline. These results establish temporal distance as a scalable supervision target and practical reward interface for generalist robot policies.","upvotes":6,"discussionId":"6a7aa443019ce76dc7b3aaf0","projectPage":"https://alibaba-damo-academy.github.io/RynnValue.github.io/","githubRepo":"https://github.com/alibaba-damo-academy/RynnValue","githubRepoAddedBy":"user","ai_summary":"RynnValue is a scalable open-source value foundation model for robot manipulation that uses temporal distance instead of preferences or progress to learn generalizable value predictions and improve real-world policy success.","ai_keywords":["reward models","value foundation model","temporal distance","cost-to-go","value-isolation attention","potential-based shaping","robot manipulation","generalist robot policies"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":11,"organization":{"_id":"6808e7522a4d69d5111da55f","name":"Alibaba-DAMO-Academy","fullname":"DAMO Academy","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6808e64de5dd22427c006e10/9J3vdB62CdeTOd_YrGh9w.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65fd82762bf2cd20ddaa193f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/yBYbWp_mT7UusYdkqtAvw.png","isPro":false,"fullname":"Siteng Huang","user":"huangsiteng","type":"user"},{"_id":"63c1699e40a26dd2db32400d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63c1699e40a26dd2db32400d/3N0-Zp8igv8-52mXAdiiq.jpeg","isPro":false,"fullname":"Chroma","user":"Chroma111","type":"user"},{"_id":"63913b120cf6b11c487ca31d","avatarUrl":"/avatars/aec44edd5470dd6e767e0a25efd6fb5d.svg","isPro":false,"fullname":"Xin Li","user":"lixin4ever","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"66974212a9e7257fc37798dc","avatarUrl":"/avatars/4063ac7e4a39f1a761374136983b7305.svg","isPro":false,"fullname":"Bohan Hou","user":"hbh123","type":"user"},{"_id":"6270324ebecab9e2dcf245de","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6270324ebecab9e2dcf245de/cMbtWSasyNlYc9hvsEEzt.jpeg","isPro":false,"fullname":"Kye Gomez","user":"kye","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6808e7522a4d69d5111da55f","name":"Alibaba-DAMO-Academy","fullname":"DAMO Academy","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6808e64de5dd22427c006e10/9J3vdB62CdeTOd_YrGh9w.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.09853.md","query":{}}">
Papers
arxiv:2608.09853

RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

Published on Aug 10
· Submitted by
Siteng Huang
on Aug 11
Authors:
,

Abstract

RynnValue is a scalable open-source value foundation model for robot manipulation that uses temporal distance instead of preferences or progress to learn generalizable value predictions and improve real-world policy success.

General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing approaches tie supervision to task-internal anchors such as preferences or normalized progress, none of which transfer cleanly across embodiments and data sources. We introduce RynnValue, an open-source value foundation model for robotic manipulation that replaces these anchors with temporal distance, the directed cost-to-go from an observation to the language-specified goal. Because temporal-distance labels can be derived directly from timestamps, RynnValue scales to over 7,000 hours and roughly 3M instruction-conditioned clips without preference or progress annotations. To make temporal-value learning reliable at scale, we combine random temporal sampling, temporal-order shuffling, and value-isolation attention, suppressing shortcuts that would leave predictions insensitive to failures and regressions. Trained without preference labels, RynnValue attains an average Kendall's tau_a of 0.675 on RBM-EVAL-OOD, surpassing the fully preference-supervised state of the art (0.655) and more than doubling a progress-only counterpart (0.292), while generalizing zero-shot to unseen tasks, embodiments, and viewpoints. Converted into dense rewards via potential-based shaping, it raises real-world policy success from 52.5% to 72.5% online and from 63.8% to 82.5% offline. These results establish temporal distance as a scalable supervision target and practical reward interface for generalist robot policies.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.09853
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.09853 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.09853 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers