<a href=\"https://cdn-uploads.huggingface.co/production/uploads/6429ac8e8136224fee087253/maJdLYKT4nxMtgZf9Oda_.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/6429ac8e8136224fee087253/maJdLYKT4nxMtgZf9Oda_.png\" alt=\"Screenshot 2026-07-23 at 22.38.14\"></a></p>\n","updatedAt":"2026-07-23T14:38:35.803Z","author":{"_id":"6429ac8e8136224fee087253","avatarUrl":"/avatars/f7e4f0885e8c4a90c75c4e36ae3fba6e.svg","fullname":"Amber Yijia Zheng","name":"amberyzheng","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.35629379749298096},"editors":["amberyzheng"],"editorAvatarUrls":["/avatars/f7e4f0885e8c4a90c75c4e36ae3fba6e.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.18789","authors":[{"_id":"6a6226712ee212ed0e2a128d","name":"Amber Yijia Zheng","hidden":false},{"_id":"6a6226712ee212ed0e2a128e","name":"Lu Liu","hidden":false},{"_id":"6a6226712ee212ed0e2a128f","name":"Raymond A. Yeh","hidden":false},{"_id":"6a6226712ee212ed0e2a1290","name":"Xi Yin","hidden":false}],"publishedAt":"2026-07-21T00:00:00.000Z","submittedOnDailyAt":"2026-07-23T00:00:00.000Z","title":"Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation","submittedOnDailyBy":{"_id":"6429ac8e8136224fee087253","avatarUrl":"/avatars/f7e4f0885e8c4a90c75c4e36ae3fba6e.svg","isPro":false,"fullname":"Amber Yijia Zheng","user":"amberyzheng","type":"user","name":"amberyzheng"},"summary":"Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-world data curation is complex and non-trivial, involving clip selection from raw videos and captioning to create video-text pairs for learning text-to-video mappings. We study how data distribution and caption quality impact text-to-video models. To enable controlled experiments, we introduce Moving Alphabet, a procedural testbed that renders letters with varying fonts, colors, sizes, and positions, moving in different directions and speeds against a black background. This design allows precise control over data distribution and caption quality by corrupting ground-truth metadata. Our experiments yield three findings: a) a diverse and balanced distribution of video content and duration is critical for generalization; b) caption quality significantly affects both model performance and training efficiency, suggesting that text-to-video models are bounded by video understanding capabilities; and c) classifier-free guidance and fine-tuning on high-quality data provide partial recovery from models trained on corrupted captions, but cannot fully compensate for poor pre-training data. We believe these insights can inform the development of large-scale text-to-video models, and we advocate for greater attention to the science of pre-training data.","upvotes":0,"discussionId":"6a6226712ee212ed0e2a1291"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[],"acceptLanguages":["en"],"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.18789.md","query":{}}">
Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
Abstract
Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-world data curation is complex and non-trivial, involving clip selection from raw videos and captioning to create video-text pairs for learning text-to-video mappings. We study how data distribution and caption quality impact text-to-video models. To enable controlled experiments, we introduce Moving Alphabet, a procedural testbed that renders letters with varying fonts, colors, sizes, and positions, moving in different directions and speeds against a black background. This design allows precise control over data distribution and caption quality by corrupting ground-truth metadata. Our experiments yield three findings: a) a diverse and balanced distribution of video content and duration is critical for generalization; b) caption quality significantly affects both model performance and training efficiency, suggesting that text-to-video models are bounded by video understanding capabilities; and c) classifier-free guidance and fine-tuning on high-quality data provide partial recovery from models trained on corrupted captions, but cannot fully compensate for poor pre-training data. We believe these insights can inform the development of large-scale text-to-video models, and we advocate for greater attention to the science of pre-training data.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.18789 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.18789 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.18789 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.