Hugging Face Daily Papers · · 3 min read

MiniWorld: Democratizing the Training of Video World Models from Scratch

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

MiniWorld: Democratizing the Training of Video World Models from Scratch</p>\n","updatedAt":"2026-08-05T08:23:05.692Z","author":{"_id":"63af9d7f0a2f4a0933956d7f","avatarUrl":"/avatars/7a49a1f628e2b41deaf9bac626d7cace.svg","fullname":"Zhao Yian","name":"zhaoyian01","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7195887565612793},"editors":["zhaoyian01"],"editorAvatarUrls":["/avatars/7a49a1f628e2b41deaf9bac626d7cace.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.01127","authors":[{"_id":"6a71505fec5082b9f872cd01","name":"Yian Zhao","hidden":false},{"_id":"6a71505fec5082b9f872cd02","name":"Ruochong Zheng","hidden":false},{"_id":"6a71505fec5082b9f872cd03","name":"Hongcan Guo","hidden":false},{"_id":"6a71505fec5082b9f872cd04","name":"Yu Yan","hidden":false},{"_id":"6a71505fec5082b9f872cd05","name":"Jian Zhang","hidden":false},{"_id":"6a71505fec5082b9f872cd06","name":"Jie Chen","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/63af9d7f0a2f4a0933956d7f/cgda08v38hchYC4qw8QaX.mp4","https://cdn-uploads.huggingface.co/production/uploads/63af9d7f0a2f4a0933956d7f/1N2I7LEaOB8FNW9M11EKb.mp4"],"publishedAt":"2026-08-02T00:00:00.000Z","submittedOnDailyAt":"2026-08-05T00:00:00.000Z","title":"MiniWorld: Democratizing the Training of Video World Models from Scratch","submittedOnDailyBy":{"_id":"63af9d7f0a2f4a0933956d7f","avatarUrl":"/avatars/7a49a1f628e2b41deaf9bac626d7cace.svg","isPro":false,"fullname":"Zhao Yian","user":"zhaoyian01","type":"user","name":"zhaoyian01"},"summary":"Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through autoregressive state transitions. Unlike conventional video generation models that primarily capture visual appearance and motion, video world models learn the underlying dynamics governing environment evolution under agent actions, providing a foundation for embodied AI and interactive simulation. Recent progress has largely relied on adapting pretrained video generation models through post-training or distillation. Although effective, these approaches often require complex training pipelines, substantial computational resources, and suffer from the mismatch between bidirectional pretraining and causal streaming inference. Recent studies have shown that training autoregressive video world models from scratch is feasible and scalable. However, the community still lacks a lightweight, transparent, and fully reproducible baseline trainable end-to-end with modest computational resources. We present MiniWorld, a reproducible framework for training streaming video world models from scratch. MiniWorld employs a block-causal Video Diffusion Transformer trained with Flow Matching in the latent space of a pretrained Video VAE. Building on Diffusion Forcing, it adopts a chunk-wise non-decreasing noise schedule and two-stage continued training to improve temporal modeling and stability. During inference, MiniWorld combines a rolling KV cache with pipelined asynchronous denoising for efficient streaming generation under bounded computation. The entire model can be trained within several days on a single 8-GPU server. By releasing the training and inference codebase and pretrained checkpoints, we hope MiniWorld will facilitate future research on video world modeling.","upvotes":9,"discussionId":"6a71505fec5082b9f872cd07","projectPage":"https://zhao-yian.github.io/MiniWorld","githubRepo":"https://github.com/zhao-yian/MiniWorld","githubRepoAddedBy":"user","githubStars":21,"organization":{"_id":"61dcd8e344f59573371b5cb6","name":"PekingUniversity","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vavgrBsnkSejriUF4lXDE.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63af9d7f0a2f4a0933956d7f","avatarUrl":"/avatars/7a49a1f628e2b41deaf9bac626d7cace.svg","isPro":false,"fullname":"Zhao Yian","user":"zhaoyian01","type":"user"},{"_id":"65edc863770aa0e25d31e2b8","avatarUrl":"/avatars/7d03ef85ffb617b9e370a6422954a3f6.svg","isPro":false,"fullname":"wuming","user":"wuming251","type":"user"},{"_id":"6a6c7a3d3574d63d5305cd74","avatarUrl":"/avatars/78a0b7bd8c1b9996c19966b58df44488.svg","isPro":false,"fullname":"Richard Wilson","user":"Cobalt-Richard4","type":"user"},{"_id":"6a6c7cc86c4896078fec71b5","avatarUrl":"/avatars/4cf98265c1ead1bdac6e770c404f8814.svg","isPro":false,"fullname":"Michael White","user":"Vector-Beacon","type":"user"},{"_id":"6a6d58512a48f063703b8d3d","avatarUrl":"/avatars/25aa979d4bd497cee90a5429fa13f354.svg","isPro":false,"fullname":"Richard Martinez","user":"Quiet-Richard","type":"user"},{"_id":"6a6de87ef9134eddf85d43da","avatarUrl":"/avatars/48a021580074a097f393f324bad012bb.svg","isPro":false,"fullname":"Timothy Lopez","user":"alex-6968806","type":"user"},{"_id":"69e9fb295e900af460ceda90","avatarUrl":"/avatars/8aebe2ed4286622eb106d2be384f583f.svg","isPro":false,"fullname":"Hongcan Guo","user":"HongcanGuo04","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"64b929308b53fb5dbd059ce3","avatarUrl":"/avatars/564c44cf45db794747a96f79f30ecd91.svg","isPro":false,"fullname":"Liu Songhua","user":"Huage001","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"61dcd8e344f59573371b5cb6","name":"PekingUniversity","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vavgrBsnkSejriUF4lXDE.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.01127.md","query":{}}">
Papers
arxiv:2608.01127

MiniWorld: Democratizing the Training of Video World Models from Scratch

Published on Aug 2
· Submitted by
Zhao Yian
on Aug 5
Authors:
,

Abstract

Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through autoregressive state transitions. Unlike conventional video generation models that primarily capture visual appearance and motion, video world models learn the underlying dynamics governing environment evolution under agent actions, providing a foundation for embodied AI and interactive simulation. Recent progress has largely relied on adapting pretrained video generation models through post-training or distillation. Although effective, these approaches often require complex training pipelines, substantial computational resources, and suffer from the mismatch between bidirectional pretraining and causal streaming inference. Recent studies have shown that training autoregressive video world models from scratch is feasible and scalable. However, the community still lacks a lightweight, transparent, and fully reproducible baseline trainable end-to-end with modest computational resources. We present MiniWorld, a reproducible framework for training streaming video world models from scratch. MiniWorld employs a block-causal Video Diffusion Transformer trained with Flow Matching in the latent space of a pretrained Video VAE. Building on Diffusion Forcing, it adopts a chunk-wise non-decreasing noise schedule and two-stage continued training to improve temporal modeling and stability. During inference, MiniWorld combines a rolling KV cache with pipelined asynchronous denoising for efficient streaming generation under bounded computation. The entire model can be trained within several days on a single 8-GPU server. By releasing the training and inference codebase and pretrained checkpoints, we hope MiniWorld will facilitate future research on video world modeling.

Community

Paper submitter about 10 hours ago

MiniWorld: Democratizing the Training of Video World Models from Scratch

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.01127
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.01127 in a dataset README.md to link it from this page.

Spaces citing this paper

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers