🚀 <strong>StatePlay: Beyond pixel-level realism toward mechanics-consistent Game World Models!</strong> Instead of modeling gameplay through visual observations alone, we explicitly predict internal game states and use them to guide frame generation, ensuring consistency with the underlying game mechanics.</p>\n<p>✨ <strong>Highlights:</strong></p>\n<p><strong>State-Aware Generation:</strong> Jointly predicts visual content and precise game states, including health points, skill meters, and timers.</p>\n<p><strong>Mechanics-Consistent:</strong> Couples predicted states with frame generation to enforce state-dependent rules such as valid skill activation, health reduction, and game termination.</p>\n<p><strong>Specialized Architecture:</strong> A Mixture-of-Transformers-style design preserves modality-specific visual and state representations while enabling effective cross-modal interaction.</p>\n<p><strong>Strong Performance:</strong> Achieves a normalized state-prediction error below 0.06 and improves mechanics fidelity by <strong>18.6%</strong> over models without explicit state modeling.</p>\n<p>👇 <strong>Dive in:</strong></p>\n<p>📄 <strong>Paper:</strong> <a href=\"https://arxiv.org/abs/2607.26754\" rel=\"nofollow\">https://arxiv.org/abs/2607.26754</a><br>🏠 <strong>Project:</strong> <a href=\"https://jimntu.github.io/stateplay_page/\" rel=\"nofollow\">https://jimntu.github.io/stateplay_page/</a><br>💻 <strong>Code:</strong> <a href=\"https://github.com/Jimntu/StatePlay\" rel=\"nofollow\">https://github.com/Jimntu/StatePlay</a><br>🤗 <strong>Model:</strong> <a href=\"https://huggingface.co/onepiece1999/StatePlay\">https://huggingface.co/onepiece1999/StatePlay</a><br>🗄️ <strong>Dataset:</strong> <a href=\"https://huggingface.co/datasets/onepiece1999/StatePlay-Dataset\">https://huggingface.co/datasets/onepiece1999/StatePlay-Dataset</a></p>\n","updatedAt":"2026-07-30T09:12:01.727Z","author":{"_id":"677e31bb114aeff62d62bc11","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/JA85qWXtPO_mnK7_Bu2HJ.png","fullname":"ZIJUN LIN","name":"onepiece1999","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.7279802560806274},"editors":["onepiece1999"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/JA85qWXtPO_mnK7_Bu2HJ.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.26754","authors":[{"_id":"6a6af60c4463a8a84bdc4074","user":{"_id":"677e31bb114aeff62d62bc11","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/JA85qWXtPO_mnK7_Bu2HJ.png","isPro":false,"fullname":"ZIJUN LIN","user":"onepiece1999","type":"user","name":"onepiece1999"},"name":"Zijun Lin","status":"claimed_verified","statusLastChangedAt":"2026-07-30T08:45:04.430Z","hidden":false},{"_id":"6a6af60c4463a8a84bdc4075","user":{"_id":"668e740f1173ab43d9d9ed5e","avatarUrl":"/avatars/caa9b47c2a5f6d6d679759b8b234a0ab.svg","isPro":false,"fullname":"Zeqing Wang","user":"INV-WZQ","type":"user","name":"INV-WZQ"},"name":"Zeqing Wang","status":"claimed_verified","statusLastChangedAt":"2026-07-30T16:45:04.659Z","hidden":false},{"_id":"6a6af60c4463a8a84bdc4076","name":"Cheston Tan","hidden":false},{"_id":"6a6af60c4463a8a84bdc4077","name":"Bihan Wen","hidden":false},{"_id":"6a6af60c4463a8a84bdc4078","name":"Yeying Jin","hidden":false}],"publishedAt":"2026-07-29T00:00:00.000Z","submittedOnDailyAt":"2026-07-30T00:00:00.000Z","title":"StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation","submittedOnDailyBy":{"_id":"677e31bb114aeff62d62bc11","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/JA85qWXtPO_mnK7_Bu2HJ.png","isPro":false,"fullname":"ZIJUN LIN","user":"onepiece1999","type":"user","name":"onepiece1999"},"summary":"Recent game world models can generate visually realistic and interactive environments conditioned on player actions. However, games are not defined by pixels alone; they are governed by explicit mechanics, namely state-dependent rules that control health reduction, skill activation, and game termination. These mechanics depend on precise internal states, such as health points, skill meters, and timers, which are tightly coupled with visual observations and determine how gameplay evolves. Without modeling these state dynamics, existing game world models may generate visually plausible rollouts but violate the underlying game rules. In this paper, we propose StatePlay, a novel state-aware game world model that jointly predicts visual content and game states to promote mechanics-consistent generation. StatePlay adopts a mixture-of-transformers (MoT)-style architecture that preserves specialized visual and state representations while enabling cross-modal interaction, allowing predicted states to guide frame generation. Each branch is further optimized with a distinct objective suited to its modality. Experiments show that StatePlay achieves an average normalized L1 distance below 0.06 for state prediction. Furthermore, compared with models without explicit state modeling, our method improves mechanics fidelity in generated game rollouts by 18.6%. Overall, our work highlights the importance of state-aware game world modeling and advances beyond pixel-level realism toward complete and mechanically faithful game generation.","upvotes":13,"discussionId":"6a6af60d4463a8a84bdc4079","projectPage":"https://jimntu.github.io/stateplay_page/","githubRepo":"https://github.com/Jimntu/StatePlay","githubRepoAddedBy":"user","githubStars":9},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"677e31bb114aeff62d62bc11","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/JA85qWXtPO_mnK7_Bu2HJ.png","isPro":false,"fullname":"ZIJUN LIN","user":"onepiece1999","type":"user"},{"_id":"6696347a0818007c1c7a69be","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6696347a0818007c1c7a69be/k-EEb_x9qASVQn3Pr6BES.png","isPro":true,"fullname":"White Tea","user":"teawhite","type":"user"},{"_id":"6729d1fed3ec5370cb035901","avatarUrl":"/avatars/50f7ce9c635148df76d1c63ebf3efa38.svg","isPro":false,"fullname":"1","user":"DANNY621","type":"user"},{"_id":"668e740f1173ab43d9d9ed5e","avatarUrl":"/avatars/caa9b47c2a5f6d6d679759b8b234a0ab.svg","isPro":false,"fullname":"Zeqing Wang","user":"INV-WZQ","type":"user"},{"_id":"68563e56c09959da82e70968","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68563e56c09959da82e70968/6RfcWlidFVx07aNbKXYgf.jpeg","isPro":false,"fullname":"Yeying Jin","user":"yeyingjin","type":"user"},{"_id":"637f114c1dbae0919108987d","avatarUrl":"/avatars/23d73811b697261ceb80ef1b0806a633.svg","isPro":false,"fullname":"Zizhao Tong","user":"zizhaotong","type":"user"},{"_id":"63ee2fc05f1300034ddb543c","avatarUrl":"/avatars/1e056da18c4db7c3aae5740d0fd10b99.svg","isPro":false,"fullname":"Yixin Yang","user":"yyang181","type":"user"},{"_id":"6918436beff127c6b5564d17","avatarUrl":"/avatars/c71a3e4dbf93df20fd16b859d3be9738.svg","isPro":false,"fullname":"zhangboran","user":"BBoran","type":"user"},{"_id":"65254c565378d720ebb098fa","avatarUrl":"/avatars/10c7e746799754ca5566ce030f812e5f.svg","isPro":false,"fullname":"taylorrr","user":"taylorrr","type":"user"},{"_id":"665c476f052479b276a7239d","avatarUrl":"/avatars/f1ac6fa099efd6141a7382c6ecfced96.svg","isPro":false,"fullname":"yuwei zhang ","user":"zyw2002","type":"user"},{"_id":"624bebf604abc7ebb01789af","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1649143001781-624bebf604abc7ebb01789af.jpeg","isPro":true,"fullname":"Apolinário from multimodal AI art","user":"multimodalart","type":"user"},{"_id":"672e2f4ee4b1d710f7502af5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ecDhbblzDfMJVmSHSy8mq.png","isPro":false,"fullname":"MetrisVailore","user":"Metris","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.26754.md","query":{}}">
StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation
Abstract
Recent game world models can generate visually realistic and interactive environments conditioned on player actions. However, games are not defined by pixels alone; they are governed by explicit mechanics, namely state-dependent rules that control health reduction, skill activation, and game termination. These mechanics depend on precise internal states, such as health points, skill meters, and timers, which are tightly coupled with visual observations and determine how gameplay evolves. Without modeling these state dynamics, existing game world models may generate visually plausible rollouts but violate the underlying game rules. In this paper, we propose StatePlay, a novel state-aware game world model that jointly predicts visual content and game states to promote mechanics-consistent generation. StatePlay adopts a mixture-of-transformers (MoT)-style architecture that preserves specialized visual and state representations while enabling cross-modal interaction, allowing predicted states to guide frame generation. Each branch is further optimized with a distinct objective suited to its modality. Experiments show that StatePlay achieves an average normalized L1 distance below 0.06 for state prediction. Furthermore, compared with models without explicit state modeling, our method improves mechanics fidelity in generated game rollouts by 18.6%. Overall, our work highlights the importance of state-aware game world modeling and advances beyond pixel-level realism toward complete and mechanically faithful game generation.
Community
🚀 StatePlay: Beyond pixel-level realism toward mechanics-consistent Game World Models! Instead of modeling gameplay through visual observations alone, we explicitly predict internal game states and use them to guide frame generation, ensuring consistency with the underlying game mechanics.
✨ Highlights:
State-Aware Generation: Jointly predicts visual content and precise game states, including health points, skill meters, and timers.
Mechanics-Consistent: Couples predicted states with frame generation to enforce state-dependent rules such as valid skill activation, health reduction, and game termination.
Specialized Architecture: A Mixture-of-Transformers-style design preserves modality-specific visual and state representations while enabling effective cross-modal interaction.
Strong Performance: Achieves a normalized state-prediction error below 0.06 and improves mechanics fidelity by 18.6% over models without explicit state modeling.
👇 Dive in:
📄 Paper: https://arxiv.org/abs/2607.26754
🏠 Project: https://jimntu.github.io/stateplay_page/
💻 Code: https://github.com/Jimntu/StatePlay
🤗 Model: https://huggingface.co/onepiece1999/StatePlay
🗄️ Dataset: https://huggingface.co/datasets/onepiece1999/StatePlay-Dataset
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.26754 in a dataset README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.