Let's generalize to fully synthetic data!</p>\n","updatedAt":"2026-09-10T05:33:37.646Z","author":{"_id":"6313025e3ef0644e256403c8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/EASeMfkGl6rVvpx-YZbw_.png","fullname":"Andy Yermakov","name":"yermandy","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.818126380443573},"editors":["yermandy"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/EASeMfkGl6rVvpx-YZbw_.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.07369","authors":[{"_id":"6aa0fafdd0174964227bee54","name":"Severyn Shykula","hidden":false},{"_id":"6aa0fafdd0174964227bee55","user":{"_id":"6313025e3ef0644e256403c8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/EASeMfkGl6rVvpx-YZbw_.png","isPro":false,"fullname":"Andy Yermakov","user":"yermandy","type":"user","name":"yermandy"},"name":"Andrii Yermakov","status":"claimed_verified","statusLastChangedAt":"2026-09-09T09:22:43.847Z","hidden":false},{"_id":"6aa0fafdd0174964227bee56","name":"Ivan Samarskyi","hidden":false},{"_id":"6aa0fafdd0174964227bee57","name":"Dmytro Mishkin","hidden":false},{"_id":"6aa0fafdd0174964227bee58","name":"Jan Cech","hidden":false},{"_id":"6aa0fafdd0174964227bee59","name":"Anastasiia Mishchuk","hidden":false}],"publishedAt":"2026-09-07T00:00:00.000Z","submittedOnDailyAt":"2026-09-10T00:00:00.000Z","title":"DF26: We Cannot Tell Fake From Real Anymore","submittedOnDailyBy":{"_id":"6313025e3ef0644e256403c8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/EASeMfkGl6rVvpx-YZbw_.png","isPro":false,"fullname":"Andy Yermakov","user":"yermandy","type":"user","name":"yermandy"},"summary":"We introduce DF26, a novel benchmark for detecting AI-generated videos containing fully synthetic clips produced by recent text-to-video and image-to-video models. The videos capture single-person public-speaking scenarios, spanning direct-to-camera recordings, official statements, and studio interviews - 271 real and 2,420 synthetic videos generated by seven modern video models. The study on DF26 shows that human performance in detecting AI-generated videos, as well as state-of-the-art deepfake detectors, is close to random chance. Our results highlight the limitations of current evaluation protocols and motivate the need for benchmarks that explicitly measure robustness to modern generative model distribution shifts.","upvotes":0,"discussionId":"6aa0fafdd0174964227bee5a","ai_summary":"A new benchmark for AI-generated public-speaking videos reveals that both humans and current detectors perform near chance, underscoring the need for robustness to modern generative distribution shifts.","ai_keywords":["text-to-video","image-to-video","deepfake detectors","generative model distribution shifts","AI-generated videos"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[],"acceptLanguages":["en"],"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.07369.md","query":{}}">
DF26: We Cannot Tell Fake From Real Anymore
Abstract
A new benchmark for AI-generated public-speaking videos reveals that both humans and current detectors perform near chance, underscoring the need for robustness to modern generative distribution shifts.
We introduce DF26, a novel benchmark for detecting AI-generated videos containing fully synthetic clips produced by recent text-to-video and image-to-video models. The videos capture single-person public-speaking scenarios, spanning direct-to-camera recordings, official statements, and studio interviews - 271 real and 2,420 synthetic videos generated by seven modern video models. The study on DF26 shows that human performance in detecting AI-generated videos, as well as state-of-the-art deepfake detectors, is close to random chance. Our results highlight the limitations of current evaluation protocols and motivate the need for benchmarks that explicitly measure robustness to modern generative model distribution shifts.
Community
Let's generalize to fully synthetic data!
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.07369 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.07369 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.