Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional. As a result, these models can be difficult to use for planning or policy extraction. To address these limitations, we propose RoboTALES, a single-stage framework that learns task-aligned simulated futures and uses them to train robot policies. Our approach introduces two key innovations: (1) a hierarchical LLM-based planner that breaks complex tasks into a sequence of subgoals to guide the model's imagination; and (2) a VLM-based critic that evaluates these ``imagined'' futures and uses reward-based feedback to keep the model's internal representations focused on the goal. By anchoring the video generator in abstract reasoning, we produce temporally consistent rollouts and more coherent actions. We evaluate RoboTALES on diverse manipulation tasks from RoboCasa and LIBERO10, and show that our method consistently outperforms existing methods, especially in long-horizon tasks</p>\n","updatedAt":"2026-07-09T22:55:13.230Z","author":{"_id":"64eddf17e630ae6d575f6231","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64eddf17e630ae6d575f6231/W8oofVqStLEAK1Xi4fQTh.png","fullname":"Hanan Gani","name":"hanangani","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8932214379310608},"editors":["hanangani"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64eddf17e630ae6d575f6231/W8oofVqStLEAK1Xi4fQTh.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.06018","authors":[{"_id":"6a5021e975fd3d966bd45ce1","name":"Hanan Gani","hidden":false},{"_id":"6a5021e975fd3d966bd45ce2","name":"Tejal Kulkarni","hidden":false},{"_id":"6a5021e975fd3d966bd45ce3","name":"Madhoolika Chodavarapu","hidden":false},{"_id":"6a5021e975fd3d966bd45ce4","name":"Nicklas Hansen","hidden":false},{"_id":"6a5021e975fd3d966bd45ce5","name":"Manmohan Chandraker","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/64eddf17e630ae6d575f6231/5FPjC4FbMvyH2DzcOYXrv.jpeg","https://cdn-uploads.huggingface.co/production/uploads/64eddf17e630ae6d575f6231/lEFJ_onJMuNTN-c6v_eUH.jpeg"],"publishedAt":"2026-07-07T00:00:00.000Z","submittedOnDailyAt":"2026-07-09T00:00:00.000Z","title":"RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures","submittedOnDailyBy":{"_id":"64eddf17e630ae6d575f6231","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64eddf17e630ae6d575f6231/W8oofVqStLEAK1Xi4fQTh.png","isPro":false,"fullname":"Hanan Gani","user":"hanangani","type":"user","name":"hanangani"},"summary":"Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional. As a result, these models can be difficult to use for planning or policy extraction. To address these limitations, we propose RoboTALES, a single-stage framework that learns task-aligned simulated futures and uses them to train robot policies. Our approach introduces two key innovations: (1) a hierarchical LLM-based planner that breaks complex tasks into a sequence of subgoals to guide the model's imagination; and (2) a VLM-based critic that evaluates these ``imagined'' futures and uses reward-based feedback to keep the model's internal representations focused on the goal. By anchoring the video generator in abstract reasoning, we produce temporally consistent rollouts and more coherent actions. We evaluate RoboTALES on diverse manipulation tasks from RoboCasa and LIBERO10, and show that our method consistently outperforms existing methods, especially in long-horizon tasks. Our code and models are publicly available at https://github.com/hananshafi/RoboTALES.","upvotes":3,"discussionId":"6a5021ea75fd3d966bd45ce6","projectPage":"https://hananshafi.github.io/RoboTALES","githubRepo":"https://github.com/hananshafi/RoboTALES","githubRepoAddedBy":"user","ai_summary":"RoboTALES introduces a two-stage framework that combines LLM-based planning and VLM-based criticism to improve task-aligned video generation and robotic policy training.","ai_keywords":["video generative models","visuomotor control","task-aligned simulated futures","hierarchical LLM-based planner","VLM-based critic","robot policies","temporal consistency","long-horizon tasks"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":2,"organization":{"_id":"6659672c197d3500e0e02a34","name":"UniversityofCaliforniaSanDiego","fullname":"University of California San Diego","avatar":"https://www.gravatar.com/avatar/a1e8d1ce6033a6208abcf8f3da33fc64?d=retro&size=100"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"64eddf17e630ae6d575f6231","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64eddf17e630ae6d575f6231/W8oofVqStLEAK1Xi4fQTh.png","isPro":false,"fullname":"Hanan Gani","user":"hanangani","type":"user"},{"_id":"6a1b190271023578a3c45cf6","avatarUrl":"/avatars/8d50de1fa3677a058bc5906febc465b3.svg","isPro":false,"fullname":"Tejal Kulkarni","user":"tejalkulkarni","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6659672c197d3500e0e02a34","name":"UniversityofCaliforniaSanDiego","fullname":"University of California San Diego","avatar":"https://www.gravatar.com/avatar/a1e8d1ce6033a6208abcf8f3da33fc64?d=retro&size=100"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.06018.md","query":{}}">
RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures
Abstract
RoboTALES introduces a two-stage framework that combines LLM-based planning and VLM-based criticism to improve task-aligned video generation and robotic policy training.
Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional. As a result, these models can be difficult to use for planning or policy extraction. To address these limitations, we propose RoboTALES, a single-stage framework that learns task-aligned simulated futures and uses them to train robot policies. Our approach introduces two key innovations: (1) a hierarchical LLM-based planner that breaks complex tasks into a sequence of subgoals to guide the model's imagination; and (2) a VLM-based critic that evaluates these ``imagined'' futures and uses reward-based feedback to keep the model's internal representations focused on the goal. By anchoring the video generator in abstract reasoning, we produce temporally consistent rollouts and more coherent actions. We evaluate RoboTALES on diverse manipulation tasks from RoboCasa and LIBERO10, and show that our method consistently outperforms existing methods, especially in long-horizon tasks. Our code and models are publicly available at https://github.com/hananshafi/RoboTALES.
Community
Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional. As a result, these models can be difficult to use for planning or policy extraction. To address these limitations, we propose RoboTALES, a single-stage framework that learns task-aligned simulated futures and uses them to train robot policies. Our approach introduces two key innovations: (1) a hierarchical LLM-based planner that breaks complex tasks into a sequence of subgoals to guide the model's imagination; and (2) a VLM-based critic that evaluates these ``imagined'' futures and uses reward-based feedback to keep the model's internal representations focused on the goal. By anchoring the video generator in abstract reasoning, we produce temporally consistent rollouts and more coherent actions. We evaluate RoboTALES on diverse manipulation tasks from RoboCasa and LIBERO10, and show that our method consistently outperforms existing methods, especially in long-horizon tasks
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.06018 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.06018 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.