<strong>👋 Authors here — happy to answer questions about WorldCycle!</strong></p>\n<p><strong>TL;DR:</strong> Video world models look great frame-by-frame, but ask one to <em>walk forward and back</em> and it never returns to where it started. We turn that failure into free supervision. <strong>WorldCycle</strong> uses reversible action cycles as a self-verifiable RL signal — no ground-truth trajectories needed — and cuts long-horizon state-returning drift by up to <strong>44%</strong>, while boosting composite-action accuracy nearly <strong>4×</strong>.</p>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"the-problem-nobody-was-measuring-🔍\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#the-problem-nobody-was-measuring-🔍\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\tThe problem nobody was measuring 🔍\n\t</span>\n</h2>\n<p>Post-training for video world models mostly optimizes short-horizon visual quality or per-step action alignment. But the real bottleneck for <em>long</em> horizons is that <strong>there's no supervision signal</strong>: for an arbitrary action sequence, no ground-truth future state exists to measure accumulated error. So models drift, and we can't even tell.</p>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"the-key-insight-💡\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#the-key-insight-💡\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\tThe key insight 💡\n\t</span>\n</h2>\n<p>Physics gives us one exact, annotation-free reference for free: <strong>a reversible action cycle must return to its starting state.</strong> Move the camera forward then back → you should see the exact same view. We show that strong baselines (WorldPlay, WorldCompass) fail this minimal check in two ways:</p>\n<ul>\n<li><strong>Spatial closure failure</strong> — an inverse action sequence doesn't recover the starting view.</li>\n<li><strong>Temporal consistency failure</strong> — the <em>same</em> action produces different displacements at different rollout depths.</li>\n</ul>\n<p>Both are invisible to any short-horizon reward, and they compound with horizon length.</p>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"method-🛠️\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#method-🛠️\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\tMethod 🛠️\n\t</span>\n</h2>\n<p>We turn closed action programs into <em>dense</em> supervision:</p>\n<ul>\n<li><strong>Spatial closure reward</strong> — build mirrored frame pairs at every intermediate depth of a cycle and compare them, so each chunk becomes an independently verifiable closure check (no sparse endpoint-only signal).</li>\n<li><strong>Temporal state-consistency reward</strong> — repeat cycles and compare co-indexed frames across them, penalizing drift of identical actions over time.</li>\n</ul>\n<p>Together these force the model to learn actions as <strong>consistent state operators</strong> rather than memorized temporal patterns — and the same objective extends to <strong>composite actions</strong> (e.g. move-while-turning) with <em>no</em> ground-truth video.</p>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"cyclebench-📊\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#cyclebench-📊\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\tCycleBench 📊\n\t</span>\n</h2>\n<p>Since no existing benchmark diagnoses state-returning consistency, we release <strong>CycleBench</strong>: reversible, repeated, and composite cycles across short-/mid-/long-horizon settings — evaluating video world models as <em>simulators</em>, not just short-horizon generators.</p>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"results\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#results\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\tResults\n\t</span>\n</h2>\n<ul>\n<li>Up to <strong>44%</strong> reduction in state-returning drift.</li>\n<li>Nearly <strong>4×</strong> improvement on composite-action accuracy over the baseline (which itself shows a 5× accuracy collapse on composite vs. simple actions).</li>\n</ul>\n<hr>\n","updatedAt":"2026-08-06T02:37:51.571Z","author":{"_id":"67d792701e998e70c607abfd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67d792701e998e70c607abfd/_iXHF0A5uTnSEGFJ9sRYu.jpeg","fullname":"Fury James","name":"MarcusGu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8491027355194092},"editors":["MarcusGu"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/67d792701e998e70c607abfd/_iXHF0A5uTnSEGFJ9sRYu.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.04964","authors":[{"_id":"6a73f1a2c5e410d076869a95","name":"Bohai Gu","hidden":false},{"_id":"6a73f1a2c5e410d076869a96","name":"Yueyang Yuan","hidden":false},{"_id":"6a73f1a2c5e410d076869a97","name":"Taiyi Wu","hidden":false},{"_id":"6a73f1a2c5e410d076869a98","name":"Dazhao Du","hidden":false},{"_id":"6a73f1a2c5e410d076869a99","name":"Jian Liu","hidden":false},{"_id":"6a73f1a2c5e410d076869a9a","name":"Xiaoyi Pang","hidden":false},{"_id":"6a73f1a2c5e410d076869a9b","name":"Jie Zhang","hidden":false},{"_id":"6a73f1a2c5e410d076869a9c","name":"Xiaocheng Lu","hidden":false},{"_id":"6a73f1a2c5e410d076869a9d","name":"Haobin Zhong","hidden":false},{"_id":"6a73f1a2c5e410d076869a9e","name":"Xiaotong Zhao","hidden":false},{"_id":"6a73f1a2c5e410d076869a9f","name":"Alan Zhao","hidden":false},{"_id":"6a73f1a2c5e410d076869aa0","name":"Song Guo","hidden":false}],"publishedAt":"2026-08-05T00:00:00.000Z","submittedOnDailyAt":"2026-08-06T00:00:00.000Z","title":"WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models","submittedOnDailyBy":{"_id":"67d792701e998e70c607abfd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67d792701e998e70c607abfd/_iXHF0A5uTnSEGFJ9sRYu.jpeg","isPro":false,"fullname":"Fury James","user":"MarcusGu","type":"user","name":"MarcusGu"},"summary":"Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compounding errors. Post-training methods such as reinforcement learning (RL) can improve these models, but they hit a verification bottleneck: for arbitrary action sequences, no ground-truth future state exists to measure long-term drift. Our key insight is that reversible action cycles make this verification possible: a sequence composed with its inverse must analytically return to the initial state, yielding annotation-free supervision on long-horizon correctness. Building on this, we introduce WorldCycle, a self-verifiable RL framework that constructs closed action cycles and their repeated executions from ordinary action sequences, and optimizes two complementary rewards: a spatial closure reward enforcing symmetry between mirrored forward and reverse segments, and a temporal consistency reward aligning states across repeated cycle executions. These rewards force the model to learn actions as consistent state operators rather than memorized temporal patterns, and extend naturally to out-of-distribution composite action cycles that the base model handles poorly. We further release CycleBench, a diagnostic benchmark for state-returning ability under complex action structures. WorldCycle reduces state returning drift by up to 44% and lifts composite-action accuracy nearly 4x over the base model, providing a vital foundation for physically grounded world models.","upvotes":7,"discussionId":"6a73f1a2c5e410d076869aa1","projectPage":"https://nevsnev.github.io/Worldcycle/","organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67d792701e998e70c607abfd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67d792701e998e70c607abfd/_iXHF0A5uTnSEGFJ9sRYu.jpeg","isPro":false,"fullname":"Fury James","user":"MarcusGu","type":"user"},{"_id":"674d9e399304bf7ffaa5bdb4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/674d9e399304bf7ffaa5bdb4/RIhRlWlJiG1TjZl_kB1CW.jpeg","isPro":false,"fullname":"yueyangyuan","user":"yueyangyuan","type":"user"},{"_id":"63c1699e40a26dd2db32400d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63c1699e40a26dd2db32400d/3N0-Zp8igv8-52mXAdiiq.jpeg","isPro":false,"fullname":"Chroma","user":"Chroma111","type":"user"},{"_id":"6a6a92b662b8078b79bd4eb1","avatarUrl":"/avatars/47d17fa1bb148ae3fdc4e077954359ea.svg","isPro":false,"fullname":"Linda Taylor","user":"linda-taylor","type":"user"},{"_id":"6a6aa6de726441725a30a027","avatarUrl":"/avatars/2e777075d9d0d6dcdaaf0e97860f551e.svg","isPro":false,"fullname":"Sarah Clark","user":"granitecore","type":"user"},{"_id":"6a6aa77d342f6961a3ac5feb","avatarUrl":"/avatars/5b391a7dd212e505964da69ce2f00d24.svg","isPro":false,"fullname":"David Lopez","user":"Cobalt-Glade","type":"user"},{"_id":"6a6de993067f0e2726f81651","avatarUrl":"/avatars/f5162244a4a7aad4706867f5fded818c.svg","isPro":false,"fullname":"Jennifer Taylor","user":"david-6479060","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.04964.md","query":{}}">
WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models
Abstract
Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compounding errors. Post-training methods such as reinforcement learning (RL) can improve these models, but they hit a verification bottleneck: for arbitrary action sequences, no ground-truth future state exists to measure long-term drift. Our key insight is that reversible action cycles make this verification possible: a sequence composed with its inverse must analytically return to the initial state, yielding annotation-free supervision on long-horizon correctness. Building on this, we introduce WorldCycle, a self-verifiable RL framework that constructs closed action cycles and their repeated executions from ordinary action sequences, and optimizes two complementary rewards: a spatial closure reward enforcing symmetry between mirrored forward and reverse segments, and a temporal consistency reward aligning states across repeated cycle executions. These rewards force the model to learn actions as consistent state operators rather than memorized temporal patterns, and extend naturally to out-of-distribution composite action cycles that the base model handles poorly. We further release CycleBench, a diagnostic benchmark for state-returning ability under complex action structures. WorldCycle reduces state returning drift by up to 44% and lifts composite-action accuracy nearly 4x over the base model, providing a vital foundation for physically grounded world models.
Community
👋 Authors here — happy to answer questions about WorldCycle!
TL;DR: Video world models look great frame-by-frame, but ask one to walk forward and back and it never returns to where it started. We turn that failure into free supervision. WorldCycle uses reversible action cycles as a self-verifiable RL signal — no ground-truth trajectories needed — and cuts long-horizon state-returning drift by up to 44%, while boosting composite-action accuracy nearly 4×.
The problem nobody was measuring 🔍
Post-training for video world models mostly optimizes short-horizon visual quality or per-step action alignment. But the real bottleneck for long horizons is that there's no supervision signal: for an arbitrary action sequence, no ground-truth future state exists to measure accumulated error. So models drift, and we can't even tell.
The key insight 💡
Physics gives us one exact, annotation-free reference for free: a reversible action cycle must return to its starting state. Move the camera forward then back → you should see the exact same view. We show that strong baselines (WorldPlay, WorldCompass) fail this minimal check in two ways:
- Spatial closure failure — an inverse action sequence doesn't recover the starting view.
- Temporal consistency failure — the same action produces different displacements at different rollout depths.
Both are invisible to any short-horizon reward, and they compound with horizon length.
Method 🛠️
We turn closed action programs into dense supervision:
- Spatial closure reward — build mirrored frame pairs at every intermediate depth of a cycle and compare them, so each chunk becomes an independently verifiable closure check (no sparse endpoint-only signal).
- Temporal state-consistency reward — repeat cycles and compare co-indexed frames across them, penalizing drift of identical actions over time.
Together these force the model to learn actions as consistent state operators rather than memorized temporal patterns — and the same objective extends to composite actions (e.g. move-while-turning) with no ground-truth video.
CycleBench 📊
Since no existing benchmark diagnoses state-returning consistency, we release CycleBench: reversible, repeated, and composite cycles across short-/mid-/long-horizon settings — evaluating video world models as simulators, not just short-horizon generators.
Results
- Up to 44% reduction in state-returning drift.
- Nearly 4× improvement on composite-action accuracy over the baseline (which itself shows a 5× accuracy collapse on composite vs. simple actions).
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.04964 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.04964 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.04964 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.