News / #video-gen Tag Video Gen 156 articles archived under #video-gen · RSS Sign in to follow Hugging Face Daily Papers research 3mo ago LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing Abstract LoomVideo presents an efficient 5B-parameter unified architecture for video generation and editing that reduces computational overhead through novel conditioning mechanisms and multi-modal alignment techniques. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Developing… 33 r/MachineLearning community 3mo ago Research in Image/Video Gen AI models [D] I've been going down a rabbit hole with image/video generation/editing models for a few months now, started with playing around with Stable Diffusion and ComfyUI, then got genuinely hooked on understanding why things work, not just that they do. I have an Engineering background… 20 Hugging Face Daily Papers research 3mo ago AAD-1: Asymmetric Adversarial Distillation for One-Step Autoregressive Video Generation Abstract AAD-1 framework improves one-step autoregressive image-to-video generation by breaking generator-discriminator symmetry and using phased training to prevent motion collapse and training instability. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We present AAD-1, an… 17 Hugging Face Daily Papers research 3mo ago Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation Abstract Echo Infinity enables real-time infinite video generation using learnable evolving memory and unified relative RoPE to overcome limitations in existing autoregressive methods. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We present Echo Infinity, an autoregressive (AR)… 18 Hugging Face Daily Papers research 3mo ago NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation Abstract OmniDreams, a foundation generative world model trained from the Cosmos diffusion model, enables real-time action-conditioned video generation for autonomous driving policy evaluation in complex, unseen scenarios. Generated by Qwen/Qwen2.5-Coder-32B-Instruct As… 23 Hugging Face Daily Papers research 3mo ago LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation Abstract LongLive-RAG addresses long-video generation challenges by using retrieval-augmented generation to overcome error accumulation from sliding-window attention, enabling better temporal coherence and quality. AI-generated summary Autoregressive (AR) video diffusion enables… 22 arXiv — Machine Learning research 3mo ago SORA: Free Second-Order Attacks in Fast Adversarial Training arXiv:2606.00738v1 Announce Type: new Abstract: Adversarial Training (AT) is a leading defense against adversarial examples but often suffers from Catastrophic Overfitting (CO) in efficient single-step variants, where robustness to multi-step attacks collapses despite high… 33 Hugging Face Daily Papers research 3mo ago VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization Abstract Video generation models combined with vision-language models acting as test-time teachers through differentiable rewards achieve superior video reasoning performance. AI-generated summary The recent "Reasoning with Video" paradigm utilizes Video Generation Models (VGMs)… 36 Hugging Face Daily Papers research 3mo ago StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration Abstract StreamChar enables real-time streaming audio-video generation for character animation by separating long-horizon orchestration from short-window denoising through an LLM-based orchestrator and joint audio-video DiT, achieving efficient deployment via two-stage… 8 Hugging Face Daily Papers research 3mo ago Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models Abstract A systematic comparison of vision-language models and video generation models reveals complementary strengths for spatial intelligence tasks, with vision-language models excelling in semantic tagging and instance grouping while video generation models perform better in… 21 Hugging Face Daily Papers research 3mo ago One-Forcing: Towards Stable One-Step Autoregressive Video Generation Abstract One-Forcing improves one-step video generation quality and efficiency by combining DMD objective with GAN loss, achieving state-of-the-art results with reduced training costs. AI-generated summary Recent advances have substantially improved real-time interactive video… 32 Hugging Face Daily Papers research 3mo ago DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory Abstract A novel decoupled memory architecture called DecMem is introduced for consistent long-horizon video generation, addressing computational inefficiency and attention dispersion issues in learnable memory systems. AI-generated summary Recent advances in video generative… 30 Hugging Face Daily Papers research 3mo ago Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models Abstract Lumos-Nexus is a training-efficient video generation framework that uses a two-stage approach with a lightweight generator for training and a high-capacity pretrained generator for inference, achieving enhanced visual fidelity through Unified Progressive Frequency… 35 r/MachineLearning community 3mo ago What’s the actual focus in World Models right now? [R] Hey everyone, I'm trying to get back into the loop on world models. The last time I followed SSL closely, the buzz was all about Barlow Twins and DINO, but now everything just looks like scaled-up video generation from big industry labs. What is the actual academic research… 36 r/LocalLLaMA community 4mo ago Keeping multi-GPU rigs cool? As a newbie to building computers, been having issues trying to figure out how to cool my rig. The problem is that as the heat gets shunted upwards, each card gets hotter than the last (eg 31C -> 38C -> 42C -> 44C. At load during things like video generation, the hottest card… 34 Hugging Face Daily Papers research 4mo ago YoCausal: How Far is Video Generation from World Model? A Causality Perspective Abstract Video diffusion models exhibit arrow-of-time perception without true causal understanding, as demonstrated by a novel benchmark measuring causal cognition through reverse surprise and visual language model analysis. AI-generated summary As video diffusion models (VDMs)… 38 Hugging Face Daily Papers research 4mo ago AdaState: Self-Evolving Anchors for Streaming Video Generation Abstract Video diffusion models with adaptive state replacement generate more dynamic videos by evolving scene references rather than fixing to initial frames, using recurrent denoising as transition function. AI-generated summary Autoregressive video diffusion models generate… 24 Hugging Face Daily Papers research 4mo ago SmartDirector: Keyframe-Conditioned Cinematic Video Generation with Narrative Pacing Control Abstract SmartDirector enhances video generation by using multiple keyframes to improve narrative structure and temporal pacing through a two-stage process of low-resolution generation and high-resolution refinement. AI-generated summary The narrative quality of a video… 30 Hugging Face Daily Papers research 4mo ago Native Audio-Visual Alignment for Generation Abstract NAVA enables joint audio-video generation with improved synchronization and controllability through native audio-visual alignment and context-conditioned denoising. AI-generated summary Joint audio-video generation aims to synthesize temporally synchronized and… 38 arXiv — Machine Learning research 4mo ago Refining Multidimensional Video Reward Models via Disentangled Influence Functions arXiv:2605.28203v1 Announce Type: new Abstract: As Text-to-Video (T2V) generation models continue to evolve, the complexity of video evaluation necessitates a fine-grained assessment across various axes. To address this, recent works have focused on developing Multidimensional… 14 Hugging Face Daily Papers research 4mo ago EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration Abstract EverAnimate addresses long-horizon animated video generation challenges through persistent latent propagation and restorative flow matching to maintain visual quality and character identity. AI-generated summary We propose EverAnimate, an efficient post-training method… 4 The Information — AI news-outlet 4mo ago Kuaishou’s Kling AI Video Unit Reaches $500 Million in Annualized Revenue Chinese social media giant Kuaishou Technology said on Wednesday that its Kling AI video business reached an annualized revenue of about $500 million in March. Kling, which develops and sells AI video generation models, generated more than 650 million yuan ($96 million) in… 28 arXiv — NLP / Computation & Language research 4mo ago Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation arXiv:2605.26918v1 Announce Type: new Abstract: Video generation models (VGMs) are rapidly entering classrooms, yet existing benchmarks evaluate only perceptual quality, intrinsic faithfulness, generic safety, or video as a reasoning medium, and none assesses whether the outputs… 27 Hugging Face Daily Papers research 4mo ago EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation Abstract EvalVerse presents a comprehensive evaluation framework for generative video models that bridges the gap between human aesthetic judgment and machine scoring through expert-calibrated vision-language models and multi-stage cinematic assessment. AI-generated summary The… 33 Hugging Face Daily Papers research 4mo ago MotiMotion: Motion-Controlled Video Generation with Visual Reasoning Abstract MotiMotion introduces a reasoning-then-generation framework for motion-controlled video generation that improves plausibility through vision-language reasoning and confidence-aware control mechanisms. AI-generated summary Current motion-controlled image-to-video… 4 Hugging Face Daily Papers research 4mo ago On-Policy Adversarial Flow Distillation for Autoregressive Video Generation Abstract Adversarial Flow Distillation enables efficient distillation of heterogeneous video generation models by using on-policy feedback and forward-process flow-matching updates without requiring teacher scores or detailed trajectory information. AI-generated summary… 37 Hugging Face Daily Papers research 4mo ago Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion Abstract Pantheon360 enables high-fidelity 360° video generation for digital twins by combining 3D-aware diffusion with explicit geometric caching to ensure spatial-temporal consistency. AI-generated summary Generating complete digital twins from videos requires precise camera… 14 Hugging Face Daily Papers research 4mo ago Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration Abstract A multi-agent framework called Soap2Soap is presented for long-horizon video-to-video generation that maintains narrative structure and character identity across extended sequences through consistent semantic backbone and visual reference anchors. AI-generated summary… 25 Hugging Face Daily Papers research 4mo ago Geo-Align: Video Generation Alignment via Metric Geometry Reward Abstract Geo-Align presents a reinforcement learning framework for camera-controlled video re-rendering that improves generalization through scale-aware perceptual rewards and metric 3D estimation for camera trajectory extraction. AI-generated summary Camera-controlled video… 20 r/LocalLLaMA community 4mo ago meituan-longcat/LongCat-Video-Avatar-1.5 · Hugging Face 🚀 Model Introduction We are excited to announce the release of LongCat-Video-Avatar 1.5, an upgraded open-source framework that prioritizes extreme empirical optimization and production-readiness for audio-driven human video generation. Built upon the LongCat-Video foundation… 21 Hugging Face Daily Papers research 4mo ago FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching Abstract A novel inference-time method for long video generation using overlapping sliding windows with Tweedie matching and stochastic early-phase sampling to improve temporal consistency and visual quality. AI-generated summary Extending the generation horizon of video… 11 Hugging Face Daily Papers research 4mo ago Bernini: Latent Semantic Planning for Video Diffusion Abstract A unified video generation and editing framework combines multimodal large language models for semantic planning with diffusion models for pixel rendering, achieving state-of-the-art performance through semantic interface separation and enhanced positional embeddings.… 32 Hugging Face Daily Papers research 4mo ago Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos Abstract MIGA addresses long video generation challenges by reducing training-inference gaps and enhancing temporal consistency through dual consistency mechanisms. AI-generated summary Without incurring significant computational overhead, train-free long video generation aims… 5 Hugging Face Daily Papers research 4mo ago Video Models Can Reason with Verifiable Rewards Abstract VideoRLVR optimizes video diffusion models for verifiable reasoning tasks using reinforcement learning with rule-based rewards, achieving better performance than supervised methods in constraint-satisfying video generation. AI-generated summary Video diffusion models… 11 Hugging Face Daily Papers research 4mo ago CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition Abstract Diffusion models applied in compressed image space generate high-quality images with lower computational cost and support flexible inputs like text or boxes. AI-generated summary Recent diffusion models achieve strong photorealism and fluency in video generation, yet… 37 Hugging Face Daily Papers research 4mo ago MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Abstract MSAVBench presents the first comprehensive benchmark and adaptive evaluation framework for multi-shot audio-video generation, addressing limitations in existing benchmarks through diverse task settings and advanced evaluation mechanisms. AI-generated summary Video… 25 Hugging Face Daily Papers research 4mo ago Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation Abstract Echo-Forcing addresses limitations in interactive long-video generation by decoupling historical memory and recent dynamics through hierarchical temporal memory, scene recall frames, and difference-aware memory decay mechanisms. AI-generated summary Autoregressive video… 5 Hugging Face official-blog 4mo ago Fine-Tuning NVIDIA Cosmos Predict 2.5 with LoRA/DoRA for Robot Video Generation Back to Articles Fine-Tuning NVIDIA Cosmos Predict 2.5 with LoRA/DoRA for Robot Video Generation Enterprise + Article Published May 18, 2026 Upvote - Ting-Yun Chang ting-yunc nvidia Miguel Martin miguelmartin-nv nvidia Jonathan Allen nv-spectralflight nvidia Ke Ding kding1… 11 Hugging Face Daily Papers research 4mo ago OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation Abstract OmniHumanoid enables cross-embodiment video generation by factorizing motion transfer and embodiment-specific adaptation, allowing scalable adaptation to new humanoid embodiments using unpaired data. AI-generated summary Cross-embodiment video generation aims to… 7 TechCrunch — AI news-outlet 4mo ago Runway started by helping filmmakers. Now it wants to beat Google at AI. AI video generation startup Runway is betting that video generation is the path to world models. And that being an AI outsider is an advantage, not a liability. 20 Hugging Face Daily Papers research 4mo ago Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation Abstract A novel causal consistency distillation method enables efficient frame-wise video generation with reduced latency and improved quality compared to existing chunk-wise approaches. AI-generated summary Real-time interactive video generation requires low-latency,… 6 Hugging Face Daily Papers research 4mo ago Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video Abstract A novel approach called Warp-as-History enables camera-controlled video generation by transforming camera-induced warps into pseudo-history representations, achieving zero-shot capability without training or test-time optimization. AI-generated summary Camera-controlled… 18 Hugging Face Daily Papers research 4mo ago PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation Abstract PhyMotion introduces a physics-grounded reward system for human motion generation that evaluates kinematic plausibility, contact consistency, and dynamic feasibility to improve video quality. AI-generated summary Generating realistic human motion is a central yet… 12 Hugging Face Daily Papers research 4mo ago RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO Abstract RAVEN enables real-time video generation through causal autoregressive extrapolation with improved training alignment, while CM-GRPO enhances performance via reinforcement learning applied to consistency model sampling. AI-generated summary Causal autoregressive video… 20 Hugging Face Daily Papers research 4mo ago RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data Abstract RoboEvolve combines vision-language and video generation models in a co-evolutionary framework to enable scalable robotic manipulation with improved data efficiency and continuous learning capabilities. AI-generated summary The scalability of robotic manipulation is… 29 The Information — AI news-outlet 4mo ago Rent the Runway Cofounder to Step Down as CEO Rent the Runway cofounder Jennifer Hyman will step down from the CEO role at the end of this week, the clothing rental company said Wednesday. Hyman will remain an advisor to the company through early 2027, and current board member and former Nordstrom executive Teri Bariquit… 35 Hugging Face Daily Papers research 4mo ago FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation Abstract FaithfulFaces is a pose-faithful facial identity preservation framework that improves identity consistency in text-to-video generation through pose-shared alignment and explicit Euler angle embeddings. AI-generated summary Identity-preserving text-to-video generation… 38 arXiv — NLP / Computation & Language research 4mo ago PresentAgent-2: Towards Generalist Multimodal Presentation Agents arXiv:2605.11363v1 Announce Type: cross Abstract: Presentation generation is moving beyond static slide creation toward end-to-end presentation video generation with research grounding, multimodal media, and interactive delivery. We introduce PresentAgent-2, an agentic framework… 30 Hugging Face Daily Papers research 4mo ago CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives Abstract CausalCine enables interactive, multi-shot video generation by addressing limitations of autoregressive models through causal modeling, dynamic memory routing, and real-time distillation techniques. AI-generated summary Autoregressive video generation aims at real-time,… 38 Vercel — AI dev-tools 5mo ago Seedance 2.0 Video Generation on AI Gateway You can now access Bytedance's latest state-of-the-art video generation model, Seedance 2.0, via AI Gateway with no other provider accounts required. Seedance 2.0 is available on AI Gateway in two variants: Standard and Fast. Both share the same capabilities. Standard produces… 10 Page 3 of 4 · 156 articles ← Newer Older →