ActionSplice enables in-flight action editing for chunk-autoregressive video world models. At interruption step r, a lightweight corrector performs Counterfactual State Transport, mapping the active backbone-native representation toward the same-step state induced by the revised action while keeping the world model and sampler frozen and avoiding replay of completed evaluations. CST-R retargets the entire active chunk; CST-T preserves a temporal prefix and edits only the suffix. Across minWM–Wan Action2V and HY-WM1.5, CST-R reduces rollback-relative LPIPS by 61.5% and 75.9% relative to direct condition swapping. CST-T reduces suffix LPIPS by 56.1% and 77.5%, with 2.73× and 1.69× pixel-ready speedups over waiting.</p>\n","updatedAt":"2026-09-14T13:37:37.385Z","author":{"_id":"65d60e4d15f94930d75f2e41","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/LISLywrxIHyfet7ds9ogq.png","fullname":"pardis","name":"PardisTaghavi","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":0,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.842275083065033},"editors":["PardisTaghavi"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/LISLywrxIHyfet7ds9ogq.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.08230","authors":[{"_id":"6aa1b057a2aeb74440b1dcbc","name":"Pardis Taghavi","hidden":false},{"_id":"6aa1b057a2aeb74440b1dcbd","name":"Tingyu Guo","hidden":false},{"_id":"6aa1b057a2aeb74440b1dcbe","name":"Jonas Lossner","hidden":false},{"_id":"6aa1b057a2aeb74440b1dcbf","name":"Gaurav Pandey","hidden":false},{"_id":"6aa1b057a2aeb74440b1dcc0","name":"Reza Langari","hidden":false}],"publishedAt":"2026-09-08T00:00:00.000Z","submittedOnDailyAt":"2026-09-14T00:00:00.000Z","title":"ActionSplice: In-Flight Action Editing for Interactive World Models","submittedOnDailyBy":{"_id":"65d60e4d15f94930d75f2e41","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/LISLywrxIHyfet7ds9ogq.png","isPro":false,"fullname":"pardis","user":"PardisTaghavi","type":"user","name":"PardisTaghavi"},"summary":"Chunk-autoregressive video world models typically condition each generated chunk on one action. An action received during sampling must therefore wait for the next chunk, condition future solver evaluations on a state produced under the previous action, or trigger rollback that repeats completed evaluations. We introduce ActionSplice, an inference framework that formulates this problem as Counterfactual State Transport (CST). A lightweight corrector transports the interrupted backbone-native representation toward the matched state induced by the revised action at the same solver step. The world model and sampler remain frozen, and sampling resumes without replaying completed evaluations. The retargeting variant CST*{R} updates the entire active chunk, while the temporal-splicing variant CST*{T} preserves a temporal prefix and updates only the suffix. Across minWM-Wan Action2V and HY-WM1.5, CST*{R} reduces rollback-relative LPIPS by 61.5% and 75.9% relative to direct condition swapping. CST*{T} reduces suffix LPIPS by 56.1% and 77.5%, respectively, while providing 2.73times and 1.69times pixel-ready speedups over waiting. Under the HY-WorldPlay protocol, CST_{R} obtains a PSNR of 25.66 dB, an SSIM of 0.6902, and an LPIPS of 0.1337 against the original rollout.","upvotes":0,"discussionId":"6aa1b058a2aeb74440b1dcc1","projectPage":"https://pardistaghavi.github.io/actionsplice-website/","githubRepo":"https://github.com/PardisTaghavi/ActionSplice","githubRepoAddedBy":"user","ai_summary":"ActionSplice introduces counterfactual state transport to splice revised actions into chunk-autoregressive video world models without replaying completed evaluations, improving fidelity and speed.","ai_keywords":["chunk-autoregressive video world models","Counterfactual State Transport","ActionSplice","corrector","backbone-native representation","CST_R","CST_T","LPIPS","PSNR","SSIM"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":3,"organization":{"_id":"693049768605dfa68334b46d","name":"TexasAMUniversity","fullname":"Texas A&M University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/uv9z1cu15X7vyo70DW0tH.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[],"acceptLanguages":["en"],"organization":{"_id":"693049768605dfa68334b46d","name":"TexasAMUniversity","fullname":"Texas A&M University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/uv9z1cu15X7vyo70DW0tH.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.08230.md","query":{}}">
ActionSplice: In-Flight Action Editing for Interactive World Models
Published on Sep 8
· Submitted by pardis on Sep 14 Abstract
ActionSplice introduces counterfactual state transport to splice revised actions into chunk-autoregressive video world models without replaying completed evaluations, improving fidelity and speed.
Chunk-autoregressive video world models typically condition each generated chunk on one action. An action received during sampling must therefore wait for the next chunk, condition future solver evaluations on a state produced under the previous action, or trigger rollback that repeats completed evaluations. We introduce ActionSplice, an inference framework that formulates this problem as Counterfactual State Transport (CST). A lightweight corrector transports the interrupted backbone-native representation toward the matched state induced by the revised action at the same solver step. The world model and sampler remain frozen, and sampling resumes without replaying completed evaluations. The retargeting variant CST*{R} updates the entire active chunk, while the temporal-splicing variant CST*{T} preserves a temporal prefix and updates only the suffix. Across minWM-Wan Action2V and HY-WM1.5, CST*{R} reduces rollback-relative LPIPS by 61.5% and 75.9% relative to direct condition swapping. CST*{T} reduces suffix LPIPS by 56.1% and 77.5%, respectively, while providing 2.73times and 1.69times pixel-ready speedups over waiting. Under the HY-WorldPlay protocol, CST_{R} obtains a PSNR of 25.66 dB, an SSIM of 0.6902, and an LPIPS of 0.1337 against the original rollout.
Community
ActionSplice enables in-flight action editing for chunk-autoregressive video world models. At interruption step r, a lightweight corrector performs Counterfactual State Transport, mapping the active backbone-native representation toward the same-step state induced by the revised action while keeping the world model and sampler frozen and avoiding replay of completed evaluations. CST-R retargets the entire active chunk; CST-T preserves a temporal prefix and edits only the suffix. Across minWM–Wan Action2V and HY-WM1.5, CST-R reduces rollback-relative LPIPS by 61.5% and 75.9% relative to direct condition swapping. CST-T reduces suffix LPIPS by 56.1% and 77.5%, with 2.73× and 1.69× pixel-ready speedups over waiting.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.08230 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.08230 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.08230 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.