News / #video-gen Tag Video Gen 124 articles archived under #video-gen · RSS Sign in to follow Hugging Face Daily Papers research 21h ago AVA-Encoder: Towards Agent-Native Video Representation Learning Abstract AVA-Encoder learns structured video representations via agentic auto-encoding using knowledge graphs to enable cinematic video generation and reasoning with reduced token usage. Generated by thinkingmachines/Inkling-Small Creative agents still lack an effective way to… 30 Hugging Face Daily Papers research 3d ago Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains Abstract Sci-VBench evaluates video generation requiring scientific reasoning across disciplines, revealing that visual realism advances have not ensured accurate scientific and causal dynamics. Generated by thinkingmachines/Inkling-Small We introduce Sci-VBench, a comprehensive… 19 r/LocalLLaMA community 4d ago MiniMax H3: A New Open-Weight Video Model, Live in ComfyUI MiniMax H3 is an open-weight, general-purpose multimodal video generation model that works across text, images, video, and audio. In ComfyUI, you can use H3 for text-to-video, image-to-video, first- and last-frame generation, and reference-driven creation. H3 jointly generates… 10 Hugging Face Daily Papers research 4d ago SimWAM: A Simple World Action Model for End-to-End Autonomous Driving Abstract World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods require costly future generation at inference. We present SimWAM, a simple yet effective WAM that uses video generation… 19 r/LocalLLaMA community 4d ago Open-weight video gen that actually delivers. Five days with MiniMax H3 on local hardware. H3 weights went live on HuggingFace August 3rd and I started pulling them immediately. An omni-modal video model with native stereo audio in the same forward pass, where audio can actually drive the video generation? On open weights? I had to try it. Five days in, the quality is… 31 Hugging Face Daily Papers research 9d ago MiniWorld: Democratizing the Training of Video World Models from Scratch Abstract Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through autoregressive state transitions. Unlike conventional video generation models that primarily capture visual appearance and… 18 r/LocalLLaMA community 9d ago [Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding] First of all, obviously I took some help from AI to type this post and this is the topic that enabled me to accomplish all that: https://old.reddit.com/r/LocalLLaMA/comments/1veow4b/deepseek_v4flash_284b_moe_at_33_toks_single_68/ This post of mine is based on the link above. My… 11 Hugging Face Daily Papers research 10d ago WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity Abstract Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from… 6 Hugging Face Daily Papers research 14d ago VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System Abstract Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought… 11 Vercel — AI dev-tools 15d ago MiniMax H3 now available on AI Gateway MiniMax H3 is now available on AI Gateway. H3 generates 2K video from a text prompt, a starting image, a pair of first and last frames, or reference material. Alongside text-to-video and first-frame image-to-video, the model supports first-to-last keyframe transitions and… 25 Hugging Face Daily Papers research 16d ago Parallel Decoding Distillation for Fast Image and Video Generation Abstract Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variational score distillation (VSD) and adversarial losses to distill… 31 Hugging Face Daily Papers research 16d ago FilmBench: A Film-Grade Benchmark for Cinematic Video Generation Abstract Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally,… 17 Hugging Face Daily Papers research 17d ago Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification Abstract Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse attention alleviates this bottleneck by computing only selected key-value blocks, yet existing… 19 Hugging Face Daily Papers research 17d ago Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering Abstract Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for gaming and immersive content creation. These applications require long-horizon… 38 TechCrunch — AI news-outlet 20d ago Midjourney acquired the astrology app Co-Star The AI lab Midjourney continues to expand its purview beyond image and video generation. 11 Hugging Face Daily Papers research 20d ago SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Abstract We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while… 19 Hugging Face Daily Papers research 21d ago GraphVid: Interactive Graph-Controllable Video Generation Abstract Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to… 12 TechCrunch — AI news-outlet 21d ago Runway launches AI model router as generative media gets crowded Runway no longer wants to be just another AI model company. It wants to become the infrastructure layer for generative media. On Thursday, the startup launched Runway Media Router through Runway Dev, its developer platform, released earlier this month, that provides API access… 25 Hugging Face Daily Papers research 21d ago Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation Abstract Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-world data curation is complex and non-trivial, involving clip selection… 25 Hugging Face Daily Papers research 22d ago FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation Abstract Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-p routing creates uneven per-head workloads under… 16 Hugging Face Daily Papers research 24d ago Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Abstract Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through… 16 arXiv — NLP / Computation & Language research 24d ago Thinking in Video: Can Video Generators Really Reason About the Real World? arXiv:2607.17523v1 Announce Type: cross Abstract: Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about real-world dynamics. We redefine this paradigm… 25 Hugging Face Daily Papers research 24d ago HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Abstract Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance… 9 Hugging Face Daily Papers research 24d ago FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications Abstract Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, and intra-model parallelism. Existing… 10 arXiv — Machine Learning research 28d ago Inference-Time Concept Suppression and Video-Centric Evaluation for Text-to-Video Models arXiv:2607.14194v1 Announce Type: cross Abstract: Text-to-video (T2V) generators can synthesize realistic and temporally coherent videos, but controllably removing a target concept from a generator remains difficult. Unlike text-to-image concept erasure, T2V unlearning must… 5 Hugging Face Daily Papers research 28d ago MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation Abstract Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated… 11 Hugging Face Daily Papers research 28d ago KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation Abstract Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the… 14 arXiv — Machine Learning research 29d ago Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation arXiv:2607.13164v1 Announce Type: cross Abstract: Sign language is a primary communication channel for millions of Deaf and hard-of-hearing people, yet text-to-signer video generation remains costly because video diffusion models are expensive to train and evaluate. This paper… 11 Hugging Face Daily Papers research 1mo ago Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Abstract Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing… 14 arXiv — Machine Learning research 1mo ago Empowering Long-form Omni-modal Understanding with Robust Audio Perception arXiv:2607.10299v1 Announce Type: new Abstract: Recent advances in large-scale multimodal models have drivenremarkable progress in vision-language tasks; however, comprehensiveomni-modal understanding remains under-explored, largely due to thescarcity of datasets with rich,… 8 TechCrunch — AI news-outlet 1mo ago Video generation startup PixVerse raises $439M, valuation soars past $2B Singapore-based video generation startup PixVerse closed a Series C extension on the strength of 15 million monthly active users, it said. 14 Hugging Face Daily Papers research 1mo ago Video Generation Models are General-Purpose Vision Learners Abstract Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale… 28 arXiv — Machine Learning research 1mo ago GenVid2Robot: From Video Generation to Robot Manipulation via Rigid-Geometric Consistency arXiv:2607.09191v1 Announce Type: cross Abstract: Generated videos provide useful visual motion priors for robot manipulation, but their visual plausibility does not imply physical executability. A generated video usually lacks metric geometry, grasp grounding, robot kinematic… 29 Hugging Face Daily Papers research 1mo ago OpenCoF: Learning to Reason Through Video Generation Abstract OpenCoF framework introduces a reasoning video dataset and model that improve temporal reasoning through diverse supervision and explicit reasoning tokens for visual and textual cues. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Reasoning has become a core capability… 38 Hugging Face Daily Papers research 1mo ago CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation Abstract CineMobile enables efficient image-to-video generation on mobile devices through distillation-guided pruning, diffusion distillation, and hybrid quantization techniques while maintaining visual quality and achieving significant speedup. Generated by… 6 Hugging Face Daily Papers research 1mo ago Vidu S1: A Real-Time Interactive Video Generation Model Abstract Vidu S1 is a real-time interactive video generation model that supports voice-controlled digital character animation with infinite-length output and high frame rate on consumer hardware. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We introduce Vidu S1, a real-time… 15 Hugging Face Daily Papers research 1mo ago RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures Abstract RoboTALES introduces a two-stage framework that combines LLM-based planning and VLM-based criticism to improve task-aligned video generation and robotic policy training. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Pretrained video generative models are promising… 22 arXiv — NLP / Computation & Language research 1mo ago LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads? arXiv:2510.09595v3 Announce Type: replace-cross Abstract: Competitive programming problems are increasingly used to evaluate the coding capabilities of large language models (LLMs) due to their complexity and ease of verification. Yet, current coding benchmarks face limitations… 5 Hugging Face Daily Papers research 1mo ago MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing Abstract A video diffusion framework generates long, multi-view consistent videos by combining temporal and view-wise autoregression through 4D geometric bridging and spatio-temporal distillation techniques. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Recent advances in video… 15 Hugging Face Daily Papers research 1mo ago WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory Abstract WorldDirector enables controllable video generation with persistent object memory by decoupling semantic motion planning from visual rendering through LLM coordination of 3D trajectories and camera movements. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We present… 20 Hugging Face Daily Papers research 1mo ago TurboServe: Serving Streaming Video Generation Efficiently and Economically Abstract TurboServe is a specialized serving system for streaming video generation that addresses session state management and dynamic resource allocation challenges through integrated scheduling, autoscaling, and migration mechanisms. Generated by… 5 Hugging Face Daily Papers research 1mo ago AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation Abstract AVTok is a unified tokenizer for audio-video generation that uses a dual-stream transformer architecture with shared encoder-decoder and modal-specific queries to create compact one-dimensional latent representations. Generated by Qwen/Qwen2.5-Coder-32B-Instruct… 21 Hugging Face Daily Papers research 1mo ago DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Model Abstract DreamForge-World 0.1 Preview adapts a video generation architecture with a residual action pathway to enable real-time interactive world simulation on consumer hardware with low computational requirements. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We present… 18 Hugging Face Daily Papers research 1mo ago Walking in the Implicit: Interactive World Exploration via Neural Scene Representation Abstract NeuWorld enables efficient interactive video generation by representing scenes as compact neural implicit states and using a transformer VAE with diffusion transformer for trajectory-conditioned rendering. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Interactive video… 25 Hugging Face Daily Papers research 1mo ago Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation Abstract A vision-language model-based hierarchical question graph framework evaluates video generation models' adherence to physical laws with granular violation detection and human correlation validation. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Video generation models are… 23 Hugging Face Daily Papers research 1mo ago Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models Abstract Autoregressive video diffusion extends diffusion distillation frameworks to real-time streaming generation through causal training paradigms, achieving state-of-the-art performance with fast convergence and interactive world modeling capabilities. Generated by… 4 Hugging Face Daily Papers research 1mo ago MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation Abstract A novel-view video synthesis method that enhances motion-aware diffusion models through multi-view point tracking supervision to improve geometric consistency and motion fidelity. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Synthesizing a novel-view video from a… 37 Hugging Face Daily Papers research 1mo ago UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Abstract UnityShots is a memory-driven audio-video generation system that maintains consistent subject appearance and audio across video cuts using fixed-size long-term and short-term memory slots with boundary-conditioned gates and discrete cut-type priors. Generated by… 7 Hugging Face Daily Papers research 1mo ago TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy Abstract Camera-controllable video virtual try-on framework uses a 4D proxy with explicit human-environment decoupling and DiT-based video generation for omnidirectional viewing. Generated by Qwen/Qwen2.5-Coder-32B-Instruct While Video Virtual Try-on (VVT) has achieved… 4 Hugging Face Daily Papers research 1mo ago DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation Abstract DomainShuttle enables open domain subject-driven text-to-video generation with high fidelity and flexibility across in-domain and cross-domain scenarios through domain-aware modeling and dual RoPE schemes. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Open domain… 10 Page 1 of 3 · 124 articles Older →