Hugging Face Daily Papers · · 4 min read

RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional. As a result, these models can be difficult to use for planning or policy extraction. To address these limitations, we propose RoboTALES, a single-stage framework that learns task-aligned simulated futures and uses them to train robot policies. Our approach introduces two key innovations: (1) a hierarchical LLM-based planner that breaks complex tasks into a sequence of subgoals to guide the model's imagination; and (2) a VLM-based critic that evaluates these ``imagined'' futures and uses reward-based feedback to keep the model's internal representations focused on the goal. By anchoring the video generator in abstract reasoning, we produce temporally consistent rollouts and more coherent actions. We evaluate RoboTALES on diverse manipulation tasks from RoboCasa and LIBERO10, and show that our method consistently outperforms existing methods, especially in long-horizon tasks</p>\n","updatedAt":"2026-07-09T22:55:13.230Z","author":{"_id":"64eddf17e630ae6d575f6231","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64eddf17e630ae6d575f6231/W8oofVqStLEAK1Xi4fQTh.png","fullname":"Hanan Gani","name":"hanangani","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8932214379310608},"editors":["hanangani"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64eddf17e630ae6d575f6231/W8oofVqStLEAK1Xi4fQTh.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.06018","authors":[{"_id":"6a5021e975fd3d966bd45ce1","name":"Hanan Gani","hidden":false},{"_id":"6a5021e975fd3d966bd45ce2","name":"Tejal Kulkarni","hidden":false},{"_id":"6a5021e975fd3d966bd45ce3","name":"Madhoolika Chodavarapu","hidden":false},{"_id":"6a5021e975fd3d966bd45ce4","name":"Nicklas Hansen","hidden":false},{"_id":"6a5021e975fd3d966bd45ce5","name":"Manmohan Chandraker","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/64eddf17e630ae6d575f6231/5FPjC4FbMvyH2DzcOYXrv.jpeg","https://cdn-uploads.huggingface.co/production/uploads/64eddf17e630ae6d575f6231/lEFJ_onJMuNTN-c6v_eUH.jpeg"],"publishedAt":"2026-07-07T00:00:00.000Z","submittedOnDailyAt":"2026-07-09T00:00:00.000Z","title":"RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures","submittedOnDailyBy":{"_id":"64eddf17e630ae6d575f6231","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64eddf17e630ae6d575f6231/W8oofVqStLEAK1Xi4fQTh.png","isPro":false,"fullname":"Hanan Gani","user":"hanangani","type":"user","name":"hanangani"},"summary":"Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional. As a result, these models can be difficult to use for planning or policy extraction. To address these limitations, we propose RoboTALES, a single-stage framework that learns task-aligned simulated futures and uses them to train robot policies. Our approach introduces two key innovations: (1) a hierarchical LLM-based planner that breaks complex tasks into a sequence of subgoals to guide the model's imagination; and (2) a VLM-based critic that evaluates these ``imagined'' futures and uses reward-based feedback to keep the model's internal representations focused on the goal. By anchoring the video generator in abstract reasoning, we produce temporally consistent rollouts and more coherent actions. We evaluate RoboTALES on diverse manipulation tasks from RoboCasa and LIBERO10, and show that our method consistently outperforms existing methods, especially in long-horizon tasks. Our code and models are publicly available at https://github.com/hananshafi/RoboTALES.","upvotes":3,"discussionId":"6a5021ea75fd3d966bd45ce6","projectPage":"https://hananshafi.github.io/RoboTALES","githubRepo":"https://github.com/hananshafi/RoboTALES","githubRepoAddedBy":"user","ai_summary":"RoboTALES introduces a two-stage framework that combines LLM-based planning and VLM-based criticism to improve task-aligned video generation and robotic policy training.","ai_keywords":["video generative models","visuomotor control","task-aligned simulated futures","hierarchical LLM-based planner","VLM-based critic","robot policies","temporal consistency","long-horizon tasks"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":2,"organization":{"_id":"6659672c197d3500e0e02a34","name":"UniversityofCaliforniaSanDiego","fullname":"University of California San Diego","avatar":"https://www.gravatar.com/avatar/a1e8d1ce6033a6208abcf8f3da33fc64?d=retro&size=100"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"64eddf17e630ae6d575f6231","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64eddf17e630ae6d575f6231/W8oofVqStLEAK1Xi4fQTh.png","isPro":false,"fullname":"Hanan Gani","user":"hanangani","type":"user"},{"_id":"6a1b190271023578a3c45cf6","avatarUrl":"/avatars/8d50de1fa3677a058bc5906febc465b3.svg","isPro":false,"fullname":"Tejal Kulkarni","user":"tejalkulkarni","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6659672c197d3500e0e02a34","name":"UniversityofCaliforniaSanDiego","fullname":"University of California San Diego","avatar":"https://www.gravatar.com/avatar/a1e8d1ce6033a6208abcf8f3da33fc64?d=retro&size=100"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.06018.md","query":{}}">
Papers
arxiv:2607.06018

RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures

Published on Jul 7
· Submitted by
Hanan Gani
on Jul 9
Authors:
,

Abstract

RoboTALES introduces a two-stage framework that combines LLM-based planning and VLM-based criticism to improve task-aligned video generation and robotic policy training.

Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional. As a result, these models can be difficult to use for planning or policy extraction. To address these limitations, we propose RoboTALES, a single-stage framework that learns task-aligned simulated futures and uses them to train robot policies. Our approach introduces two key innovations: (1) a hierarchical LLM-based planner that breaks complex tasks into a sequence of subgoals to guide the model's imagination; and (2) a VLM-based critic that evaluates these ``imagined'' futures and uses reward-based feedback to keep the model's internal representations focused on the goal. By anchoring the video generator in abstract reasoning, we produce temporally consistent rollouts and more coherent actions. We evaluate RoboTALES on diverse manipulation tasks from RoboCasa and LIBERO10, and show that our method consistently outperforms existing methods, especially in long-horizon tasks. Our code and models are publicly available at https://github.com/hananshafi/RoboTALES.

Community

Paper submitter about 2 hours ago

Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional. As a result, these models can be difficult to use for planning or policy extraction. To address these limitations, we propose RoboTALES, a single-stage framework that learns task-aligned simulated futures and uses them to train robot policies. Our approach introduces two key innovations: (1) a hierarchical LLM-based planner that breaks complex tasks into a sequence of subgoals to guide the model's imagination; and (2) a VLM-based critic that evaluates these ``imagined'' futures and uses reward-based feedback to keep the model's internal representations focused on the goal. By anchoring the video generator in abstract reasoning, we produce temporally consistent rollouts and more coherent actions. We evaluate RoboTALES on diverse manipulation tasks from RoboCasa and LIBERO10, and show that our method consistently outperforms existing methods, especially in long-horizon tasks

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.06018
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.06018 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.06018 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers