In this paper, we propose ReactVAU, a Slow-Fast Decoupled Framework for real-time streaming Video Anomaly Understanding (VAU). Existing VAU methods rely on offline inference with global temporal sampling, which violates causality and prevents deployment in live surveillance streams. Conversely, general streaming video models satisfy causal access but dilute rare transient anomalies during memory compression and often invoke heavyweight MLLMs uniformly over long normal intervals.</p>\n","updatedAt":"2026-09-09T06:50:26.558Z","author":{"_id":"64ae22dd1aee69ece065cdcd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ae22dd1aee69ece065cdcd/JG7QaHIrr4i2k4uwR4pZK.png","fullname":"Min-Hung Chen","name":"cmhungsteve","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":21,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65df9200dc3292a8983e5017/Vs5FPVCH-VZBipV3qKTuy.png","fullname":"NVIDIA","name":"nvidia","type":"org","isHf":false,"plan":"plus"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8134996891021729},"editors":["cmhungsteve"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64ae22dd1aee69ece065cdcd/JG7QaHIrr4i2k4uwR4pZK.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.07941","authors":[{"_id":"6aa10136d0174964227bee7f","name":"Chia-Hui Chen","hidden":false},{"_id":"6aa10136d0174964227bee80","name":"Shih-Ying Yeh","hidden":false},{"_id":"6aa10136d0174964227bee81","name":"Fu-En Yang","hidden":false},{"_id":"6aa10136d0174964227bee82","user":{"_id":"64ae22dd1aee69ece065cdcd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ae22dd1aee69ece065cdcd/JG7QaHIrr4i2k4uwR4pZK.png","isPro":false,"fullname":"Min-Hung Chen","user":"cmhungsteve","type":"user","name":"cmhungsteve"},"name":"Min-Hung Chen","status":"claimed_verified","statusLastChangedAt":"2026-09-09T08:45:04.784Z","hidden":false},{"_id":"6aa10136d0174964227bee83","name":"Shang-Hong Lai","hidden":false}],"publishedAt":"2026-09-07T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding","submittedOnDailyBy":{"_id":"64ae22dd1aee69ece065cdcd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ae22dd1aee69ece065cdcd/JG7QaHIrr4i2k4uwR4pZK.png","isPro":false,"fullname":"Min-Hung Chen","user":"cmhungsteve","type":"user","name":"cmhungsteve"},"summary":"In this paper, we propose ReactVAU, a Slow-Fast Decoupled Framework for real-time streaming Video Anomaly Understanding (VAU). Existing VAU methods rely on offline inference with global temporal sampling, which violates causality and prevents deployment in live surveillance streams. Conversely, general streaming video models satisfy causal access but dilute rare transient anomalies during memory compression and often invoke heavyweight MLLMs uniformly over long normal intervals. React VAU addresses this gap with three synergistic components: a lightweight Fast Detection Module based on Spatial Grid Folding (SGF) for continuous anomaly filtering; an Anomaly-Aware Persistent Memory (AAPM) that protects critical visual cues from temporal decay; and a heavyweight Slow Reasoning Module that remains dormant during normal streams and is awakened only by suspicious events for semantic verification and causal description. Extensive experiments on multiple benchmarks demonstrate that ReactVAU operates under strict streaming constraints while simultaneously achieving competitive performance in both anomaly detection and causal reasoning, alongside significantly enhanced computational efficiency by minimizing heavyweight MLLM invocations. Project page is available at https://huiyuiui.github.io/React_VAU/","upvotes":15,"discussionId":"6aa10137d0174964227bee84","projectPage":"https://huiyuiui.github.io/ReactVAU/","githubRepo":"https://github.com/huiyuiui/ReactVAU-code","githubRepoAddedBy":"user","ai_summary":"ReactVAU enables real-time streaming video anomaly understanding via a fast detection module, persistent anomaly-aware memory, and an on-demand slow reasoning module that minimizes heavy model usage.","ai_keywords":["Slow-Fast Decoupled Framework","Video Anomaly Understanding","Spatial Grid Folding","Anomaly-Aware Persistent Memory","MLLM","streaming video","causal reasoning"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2,"organization":{"_id":"60262b67268c201cdc8b7d43","name":"nvidia","fullname":"NVIDIA","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65df9200dc3292a8983e5017/Vs5FPVCH-VZBipV3qKTuy.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64ae22dd1aee69ece065cdcd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ae22dd1aee69ece065cdcd/JG7QaHIrr4i2k4uwR4pZK.png","isPro":false,"fullname":"Min-Hung Chen","user":"cmhungsteve","type":"user"},{"_id":"6a6aa41d4c287dbb8e805dd3","avatarUrl":"/avatars/f109124053777f7667ff3235cd8aaf79.svg","isPro":false,"fullname":"Sarah Thompson","user":"rapidWing","type":"user"},{"_id":"6a6c8b9dd98e3eb6532ac650","avatarUrl":"/avatars/eb9a91990e1268fd9c43b6885104c63e.svg","isPro":false,"fullname":"John Williams","user":"atlasCraft","type":"user"},{"_id":"6a6de53ec51edbf08f1121ac","avatarUrl":"/avatars/8bcda53af64042ad674b08dd754ab666.svg","isPro":false,"fullname":"Karen Smith","user":"KarenSmith","type":"user"},{"_id":"6a7d46f08aeb2ce6c844a678","avatarUrl":"/avatars/32a54fec747e5b3ee24c098811af1d03.svg","isPro":false,"fullname":"Haoran Sun","user":"graniteVault","type":"user"},{"_id":"6a9ae0f2ba7c2d5974348f30","avatarUrl":"/avatars/c381693352a8d6f1f4090b8bbb730471.svg","isPro":false,"fullname":"中川 里佳","user":"sayuri94","type":"user"},{"_id":"6aa0df63f9e1ac2c921a6e5b","avatarUrl":"/avatars/8e2da222032c04870898d23590acdd96.svg","isPro":false,"fullname":"송명자","user":"Wild-Kgang","type":"user"},{"_id":"6aa0f310fb1ab32096de2d37","avatarUrl":"/avatars/6212882509069006e057446478bdb7b2.svg","isPro":false,"fullname":"박준호","user":"terraTyun","type":"user"},{"_id":"6a6aa2d8df1718450cef87ed","avatarUrl":"/avatars/37b86fdbfaa48d5e154f84d8a232eb15.svg","isPro":false,"fullname":"Joseph Williams","user":"LunarNico","type":"user"},{"_id":"6a6c8cac42585612a497ec3b","avatarUrl":"/avatars/6a54f4673501dbd51c7b5a88a2753c20.svg","isPro":false,"fullname":"Sarah Moore","user":"Sarah-Moore","type":"user"},{"_id":"6a8114f1d45e5232171bb3f2","avatarUrl":"/avatars/e63faa872434b04cb45ba887415f063a.svg","isPro":false,"fullname":"OrbitRidge","user":"OrbitRidge23","type":"user"},{"_id":"6a9aead6ebd110af46eb1371","avatarUrl":"/avatars/085090a50c11b7c96e8551e5f0956254.svg","isPro":false,"fullname":"张飞","user":"Cobalt-Fox","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"60262b67268c201cdc8b7d43","name":"nvidia","fullname":"NVIDIA","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65df9200dc3292a8983e5017/Vs5FPVCH-VZBipV3qKTuy.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.07941.md","query":{}}">
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding
Abstract
ReactVAU enables real-time streaming video anomaly understanding via a fast detection module, persistent anomaly-aware memory, and an on-demand slow reasoning module that minimizes heavy model usage.
In this paper, we propose ReactVAU, a Slow-Fast Decoupled Framework for real-time streaming Video Anomaly Understanding (VAU). Existing VAU methods rely on offline inference with global temporal sampling, which violates causality and prevents deployment in live surveillance streams. Conversely, general streaming video models satisfy causal access but dilute rare transient anomalies during memory compression and often invoke heavyweight MLLMs uniformly over long normal intervals. React VAU addresses this gap with three synergistic components: a lightweight Fast Detection Module based on Spatial Grid Folding (SGF) for continuous anomaly filtering; an Anomaly-Aware Persistent Memory (AAPM) that protects critical visual cues from temporal decay; and a heavyweight Slow Reasoning Module that remains dormant during normal streams and is awakened only by suspicious events for semantic verification and causal description. Extensive experiments on multiple benchmarks demonstrate that ReactVAU operates under strict streaming constraints while simultaneously achieving competitive performance in both anomaly detection and causal reasoning, alongside significantly enhanced computational efficiency by minimizing heavyweight MLLM invocations. Project page is available at https://huiyuiui.github.io/React_VAU/
Community
In this paper, we propose ReactVAU, a Slow-Fast Decoupled Framework for real-time streaming Video Anomaly Understanding (VAU). Existing VAU methods rely on offline inference with global temporal sampling, which violates causality and prevents deployment in live surveillance streams. Conversely, general streaming video models satisfy causal access but dilute rare transient anomalies during memory compression and often invoke heavyweight MLLMs uniformly over long normal intervals.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.07941 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.07941 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.07941 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.