News / #video-gen Tag Video Gen 156 articles archived under #video-gen · RSS Sign in to follow Hugging Face Daily Papers research 2mo ago Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation Abstract Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-world data curation is complex and non-trivial, involving clip selection… 25 Hugging Face Daily Papers research 2mo ago FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation Abstract Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-p routing creates uneven per-head workloads under… 16 Hugging Face Daily Papers research 2mo ago Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Abstract Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through… 16 arXiv — NLP / Computation & Language research 2mo ago Thinking in Video: Can Video Generators Really Reason About the Real World? arXiv:2607.17523v1 Announce Type: cross Abstract: Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about real-world dynamics. We redefine this paradigm… 25 Hugging Face Daily Papers research 2mo ago HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Abstract Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance… 9 Hugging Face Daily Papers research 2mo ago FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications Abstract Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, and intra-model parallelism. Existing… 10 arXiv — Machine Learning research 2mo ago Inference-Time Concept Suppression and Video-Centric Evaluation for Text-to-Video Models arXiv:2607.14194v1 Announce Type: cross Abstract: Text-to-video (T2V) generators can synthesize realistic and temporally coherent videos, but controllably removing a target concept from a generator remains difficult. Unlike text-to-image concept erasure, T2V unlearning must… 5 Hugging Face Daily Papers research 2mo ago MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation Abstract Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated… 11 Hugging Face Daily Papers research 2mo ago KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation Abstract Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the… 14 arXiv — Machine Learning research 2mo ago Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation arXiv:2607.13164v1 Announce Type: cross Abstract: Sign language is a primary communication channel for millions of Deaf and hard-of-hearing people, yet text-to-signer video generation remains costly because video diffusion models are expensive to train and evaluate. This paper… 11 Hugging Face Daily Papers research 2mo ago Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Abstract Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing… 14 arXiv — Machine Learning research 2mo ago Empowering Long-form Omni-modal Understanding with Robust Audio Perception arXiv:2607.10299v1 Announce Type: new Abstract: Recent advances in large-scale multimodal models have drivenremarkable progress in vision-language tasks; however, comprehensiveomni-modal understanding remains under-explored, largely due to thescarcity of datasets with rich,… 8 TechCrunch — AI news-outlet 2mo ago Video generation startup PixVerse raises $439M, valuation soars past $2B Singapore-based video generation startup PixVerse closed a Series C extension on the strength of 15 million monthly active users, it said. 14 Hugging Face Daily Papers research 2mo ago Video Generation Models are General-Purpose Vision Learners Abstract Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale… 28 arXiv — Machine Learning research 2mo ago GenVid2Robot: From Video Generation to Robot Manipulation via Rigid-Geometric Consistency arXiv:2607.09191v1 Announce Type: cross Abstract: Generated videos provide useful visual motion priors for robot manipulation, but their visual plausibility does not imply physical executability. A generated video usually lacks metric geometry, grasp grounding, robot kinematic… 29 Hugging Face Daily Papers research 2mo ago OpenCoF: Learning to Reason Through Video Generation Abstract OpenCoF framework introduces a reasoning video dataset and model that improve temporal reasoning through diverse supervision and explicit reasoning tokens for visual and textual cues. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Reasoning has become a core capability… 38 Hugging Face Daily Papers research 2mo ago CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation Abstract CineMobile enables efficient image-to-video generation on mobile devices through distillation-guided pruning, diffusion distillation, and hybrid quantization techniques while maintaining visual quality and achieving significant speedup. Generated by… 6 Hugging Face Daily Papers research 2mo ago Vidu S1: A Real-Time Interactive Video Generation Model Abstract Vidu S1 is a real-time interactive video generation model that supports voice-controlled digital character animation with infinite-length output and high frame rate on consumer hardware. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We introduce Vidu S1, a real-time… 15 Hugging Face Daily Papers research 2mo ago RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures Abstract RoboTALES introduces a two-stage framework that combines LLM-based planning and VLM-based criticism to improve task-aligned video generation and robotic policy training. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Pretrained video generative models are promising… 22 arXiv — NLP / Computation & Language research 2mo ago LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads? arXiv:2510.09595v3 Announce Type: replace-cross Abstract: Competitive programming problems are increasingly used to evaluate the coding capabilities of large language models (LLMs) due to their complexity and ease of verification. Yet, current coding benchmarks face limitations… 5 Hugging Face Daily Papers research 2mo ago MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing Abstract A video diffusion framework generates long, multi-view consistent videos by combining temporal and view-wise autoregression through 4D geometric bridging and spatio-temporal distillation techniques. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Recent advances in video… 15 Hugging Face Daily Papers research 2mo ago WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory Abstract WorldDirector enables controllable video generation with persistent object memory by decoupling semantic motion planning from visual rendering through LLM coordination of 3D trajectories and camera movements. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We present… 20 Hugging Face Daily Papers research 2mo ago TurboServe: Serving Streaming Video Generation Efficiently and Economically Abstract TurboServe is a specialized serving system for streaming video generation that addresses session state management and dynamic resource allocation challenges through integrated scheduling, autoscaling, and migration mechanisms. Generated by… 5 Hugging Face Daily Papers research 2mo ago AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation Abstract AVTok is a unified tokenizer for audio-video generation that uses a dual-stream transformer architecture with shared encoder-decoder and modal-specific queries to create compact one-dimensional latent representations. Generated by Qwen/Qwen2.5-Coder-32B-Instruct… 21 Hugging Face Daily Papers research 2mo ago DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Model Abstract DreamForge-World 0.1 Preview adapts a video generation architecture with a residual action pathway to enable real-time interactive world simulation on consumer hardware with low computational requirements. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We present… 18 Hugging Face Daily Papers research 3mo ago Walking in the Implicit: Interactive World Exploration via Neural Scene Representation Abstract NeuWorld enables efficient interactive video generation by representing scenes as compact neural implicit states and using a transformer VAE with diffusion transformer for trajectory-conditioned rendering. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Interactive video… 25 Hugging Face Daily Papers research 3mo ago Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation Abstract A vision-language model-based hierarchical question graph framework evaluates video generation models' adherence to physical laws with granular violation detection and human correlation validation. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Video generation models are… 23 Hugging Face Daily Papers research 3mo ago Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models Abstract Autoregressive video diffusion extends diffusion distillation frameworks to real-time streaming generation through causal training paradigms, achieving state-of-the-art performance with fast convergence and interactive world modeling capabilities. Generated by… 4 Hugging Face Daily Papers research 3mo ago MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation Abstract A novel-view video synthesis method that enhances motion-aware diffusion models through multi-view point tracking supervision to improve geometric consistency and motion fidelity. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Synthesizing a novel-view video from a… 37 Hugging Face Daily Papers research 3mo ago UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Abstract UnityShots is a memory-driven audio-video generation system that maintains consistent subject appearance and audio across video cuts using fixed-size long-term and short-term memory slots with boundary-conditioned gates and discrete cut-type priors. Generated by… 7 Hugging Face Daily Papers research 3mo ago TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy Abstract Camera-controllable video virtual try-on framework uses a 4D proxy with explicit human-environment decoupling and DiT-based video generation for omnidirectional viewing. Generated by Qwen/Qwen2.5-Coder-32B-Instruct While Video Virtual Try-on (VVT) has achieved… 4 Hugging Face Daily Papers research 3mo ago DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation Abstract DomainShuttle enables open domain subject-driven text-to-video generation with high fidelity and flexibility across in-domain and cross-domain scenarios through domain-aware modeling and dual RoPE schemes. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Open domain… 10 arXiv — Machine Learning research 3mo ago Information-Theoretic Classifier-Free Guidance with Adaptive Schedule Optimization arXiv:2606.24025v1 Announce Type: new Abstract: Diffusion models have achieved strong performance in image, text-to-image, and video generation, where conditional generation is often controlled by classifier-free guidance (CFG). CFG improves condition consistency by increasing a… 35 arXiv — Machine Learning research 3mo ago Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation arXiv:2606.23743v1 Announce Type: cross Abstract: Modern video diffusion models achieve higher generation quality through scaling, but this also increases inference cost. Although many acceleration methods have been proposed, a central challenge is that the most effective… 30 arXiv — Machine Learning research 3mo ago Autonomous Video Generation with Counterfactual Controllability for Self-Evolving World Models arXiv:2606.24152v1 Announce Type: cross Abstract: Existing literature claims that video generation essentially is world modelling. On the one hand, the claim is productive because it pushes generative AI beyond static images and toward temporally extended physical scenes. On the… 15 Hugging Face Daily Papers research 3mo ago Go-with-the-Track: Video Compositing and Motion Control with Point Tracking Abstract Go-with-the-Track unifies motion control and reference image compositing in video generation by using point-track embeddings with spatial-aware encoding and video diffusion transformers. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Filmmaking demands precise motion… 32 Hugging Face Daily Papers research 3mo ago ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Abstract ImageWAM demonstrates that pretrained image editing models can effectively replace video generation in world action models for robot control, achieving better performance with reduced computational costs. Generated by Qwen/Qwen2.5-Coder-32B-Instruct World Action Models… 25 Hugging Face Daily Papers research 3mo ago LooseControlVideo: Directorial Video Control using Spatial Blocking Abstract LooseControlVideo enables intuitive 3D spatial control in text-to-video generation using sparse oriented 3D boxes as proxies, achieving superior trajectory accuracy and occlusion handling compared to existing methods. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Precise… 10 Hugging Face Daily Papers research 3mo ago MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction Abstract 3D point motion forecasting model predicts object trajectories from visual history and language goals, demonstrating superior performance on benchmarks and transferring effectively to robot manipulation and video generation tasks. Generated by… 4 Hugging Face Daily Papers research 3mo ago Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks Abstract Track2View generates novel camera viewpoints from videos by using 3D point tracks to establish explicit spatiotemporal correspondences, achieving superior visual quality and camera accuracy compared to existing methods. Generated by Qwen/Qwen2.5-Coder-32B-Instruct… 9 Hugging Face Daily Papers research 3mo ago LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies Abstract LaWAM enables efficient robot control by predicting compact latent visual subgoals instead of expensive video generation, achieving high performance with reduced computational latency. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Vision-Language-Action models (VLAs)… 33 Hugging Face Daily Papers research 3mo ago Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Abstract Qwen-RobotWorld is a language-conditioned video world model that predicts future visual trajectories across multiple robotic domains using a double-stream diffusion transformer and embodied world knowledge corpus. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We… 5 Hugging Face Daily Papers research 3mo ago Memento: Reconstruct to Remember for Consistent Long Video Generation Abstract Memento is a subject-reconstruction-guided framework that improves long-form video generation by preserving recurring subjects through memory-based reconstruction and dual-query mechanisms. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Long-form video generation requires… 17 Hugging Face Daily Papers research 3mo ago PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory Abstract PermaVid addresses long-term video consistency after edits by using multi-modal memory banks that separate appearance and geometric structure, enabling coherent video generation across time and viewpoints. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Consistent video… 30 arXiv — NLP / Computation & Language research 3mo ago Helping Figures Tell their Story! Paper-Grounded Video Generation Explaining Complex Scientific Figures arXiv:2606.12576v1 Announce Type: new Abstract: Scientific figures compress complex pipelines into a single canvas, yet understanding them requires paper-grounded, step-by-step narration aligned with visual highlights a capability missing from current video generation systems… 11 Hugging Face Daily Papers research 3mo ago MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation Abstract Video generative models achieve improved long-range consistency through coarse-to-fine token generation using a multi-scale autoencoder and diffusion model architecture. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Video generative models have become increasingly… 28 Hugging Face Daily Papers research 3mo ago Next Forcing: Causal World Modeling with Multi-Chunk Prediction Abstract Next Forcing introduces a multi-chunk prediction framework that accelerates training and inference for autoregressive video generation while improving accuracy and physical law adherence. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Autoregressive video generation has… 19 Hugging Face Daily Papers research 3mo ago FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion Abstract FadeMem introduces a distance-aware key-value memory consolidation mechanism that organizes historical video data into a temporal hierarchy, improving long-video generation by preserving recent context and long-range anchors under fixed cache constraints. Generated by… 36 Hugging Face Daily Papers research 3mo ago Streaming Video Generation with Streaming Force Control Abstract StreamForce is a causal, unified video generation model that provides real-time, physically grounded responses to time-varying forces through a distillation pipeline and autoregressive architecture. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We introduce StreamForce,… 17 Hugging Face Daily Papers research 3mo ago Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation? Abstract Video generation models were evaluated through robotic manipulation tasks to assess their ability to reflect physical reality, revealing that visual quality does not predict executable motion accuracy. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Video generation models… 20 Page 2 of 4 · 156 articles ← Newer Older →