Adaptive video tokenization with learned keep-or-drop selection and sparse-token generation.<br>project page: <a href=\"https://github.com/kakao/KATok\" rel=\"nofollow\">https://github.com/kakao/KATok</a><br>youtube : <a href=\"https://www.youtube.com/watch?v=QCI3hB_UUOc\" rel=\"nofollow\">https://www.youtube.com/watch?v=QCI3hB_UUOc</a></p>\n","updatedAt":"2026-09-01T07:42:28.103Z","author":{"_id":"69e84bfbd1438258816ccb7f","avatarUrl":"/avatars/a29b151aeb9cf8c82803955cda5cf13a.svg","fullname":"lee","name":"yeon-kyeong","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.7957014441490173},"editors":["yeon-kyeong"],"editorAvatarUrls":["/avatars/a29b151aeb9cf8c82803955cda5cf13a.svg"],"reactions":[{"reaction":"👍","users":["tomaspy"],"count":1}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.24293","authors":[{"_id":"6a8e72b67bc881afa25f317f","user":{"_id":"69e84bfbd1438258816ccb7f","avatarUrl":"/avatars/a29b151aeb9cf8c82803955cda5cf13a.svg","isPro":false,"fullname":"lee","user":"yeon-kyeong","type":"user","name":"yeon-kyeong"},"name":"Yeonkyeong Lee","status":"claimed_verified","statusLastChangedAt":"2026-08-28T13:22:11.324Z","hidden":false},{"_id":"6a8e72b67bc881afa25f3180","name":"Hyunsung Go","hidden":false},{"_id":"6a8e72b67bc881afa25f3181","name":"Jongmin Kim","hidden":false},{"_id":"6a8e72b67bc881afa25f3182","name":"Sewoong Lim","hidden":false},{"_id":"6a8e72b67bc881afa25f3183","name":"Donghoon Lee","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/69e84bfbd1438258816ccb7f/L2yuo3BdkuftoB1GrFEZX.mp4"],"publishedAt":"2026-08-25T00:00:00.000Z","submittedOnDailyAt":"2026-09-01T00:00:00.000Z","title":"Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation","submittedOnDailyBy":{"_id":"69e84bfbd1438258816ccb7f","avatarUrl":"/avatars/a29b151aeb9cf8c82803955cda5cf13a.svg","isPro":false,"fullname":"lee","user":"yeon-kyeong","type":"user","name":"yeon-kyeong"},"summary":"Latent diffusion models have emerged as a dominant framework for high-fidelity image and video synthesis, operating in compact latent spaces with variational autoencoders (VAEs) to enhance computational efficiency without compromising visual quality. However, conventional VAEs are suboptimal for video data as they employ fixed compression ratios that cannot adapt to the varying complexity of spatio-temporal content. We present KATok (Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation), a transformer-based VAE that incorporates an adaptive token selector which is jointly learned with latent tokens. By evaluating each token's content-richness as keep-or-drop probability, the token selector effectively discards uninformative tokens, naturally allowing data-dependent compression. Applying adaptive tokenization to diffusion models may cause spatial misalignment, as token dropping can disturb the original spatio-temporal structure. To alleviate this issue, we propose two position-prediction strategies: cascaded and joint generation, to ensure spatial consistency. We empirically show that our model achieves strong reconstruction and generation quality at a state-of-the-art compression ratio. Further analysis on video data reveals that this improvement is primarily achieved by reducing spatio-temporal redundancy and removing uninformative tokens, as supported by both quantitative and qualitative results.","upvotes":7,"discussionId":"6a8e72b67bc881afa25f3184","projectPage":"https://kakao.github.io/KATok/","ai_summary":"KATok is an adaptive transformer-based video tokenizer that selectively drops uninformative tokens to achieve data-dependent compression while preserving spatial consistency for diffusion-based generation.","ai_keywords":["latent diffusion models","variational autoencoders","adaptive token selector","keep-or-drop probability","position-prediction strategies","cascaded generation","joint generation","spatio-temporal redundancy"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"66c2c83e37e5578c16d19500","name":"kakaocorp","fullname":"Kakao Corp.","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66c2a902b83a7e94d517c1ea/6Bo5zjJnRCtIi28zFUPmO.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64eee4ff29b831f265411023","avatarUrl":"/avatars/9ddca0d1c50a8d153ae4f3140caa16c1.svg","isPro":false,"fullname":"Hyunsung Go","user":"hsgo","type":"user"},{"_id":"6718bcba504128ce1d93e197","avatarUrl":"/avatars/d2e6647678882f43f2db71d0b1a76c33.svg","isPro":false,"fullname":"Hyunjoon Lee","user":"malfo-y","type":"user"},{"_id":"6a7bf19cad233873f0ca69b4","avatarUrl":"/avatars/d5454ca570fd4d4cad3405d017258f30.svg","isPro":false,"fullname":"jung","user":"hwaden","type":"user"},{"_id":"69e84bfbd1438258816ccb7f","avatarUrl":"/avatars/a29b151aeb9cf8c82803955cda5cf13a.svg","isPro":false,"fullname":"lee","user":"yeon-kyeong","type":"user"},{"_id":"6528a57bf0042c8301d217dc","avatarUrl":"/avatars/b7e1398aec545a0342c05c67c5493c8b.svg","isPro":false,"fullname":"HanSaem Kim","user":"kensaem","type":"user"},{"_id":"6662b3ee280fb71780b85ef8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/wDr9CDlL40_gp-NBRm7gB.png","isPro":false,"fullname":"Sanghyeon Na","user":"sanghyeonna","type":"user"},{"_id":"631c386bc73939ffc0716a37","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1662793811119-noauth.jpeg","isPro":false,"fullname":"SeongWan Kim","user":"idgmatrix","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"66c2c83e37e5578c16d19500","name":"kakaocorp","fullname":"Kakao Corp.","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66c2a902b83a7e94d517c1ea/6Bo5zjJnRCtIi28zFUPmO.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.24293.md","query":{}}">
Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation
Published on Aug 25
· Submitted by lee on Sep 1 Abstract
KATok is an adaptive transformer-based video tokenizer that selectively drops uninformative tokens to achieve data-dependent compression while preserving spatial consistency for diffusion-based generation.
Latent diffusion models have emerged as a dominant framework for high-fidelity image and video synthesis, operating in compact latent spaces with variational autoencoders (VAEs) to enhance computational efficiency without compromising visual quality. However, conventional VAEs are suboptimal for video data as they employ fixed compression ratios that cannot adapt to the varying complexity of spatio-temporal content. We present KATok (Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation), a transformer-based VAE that incorporates an adaptive token selector which is jointly learned with latent tokens. By evaluating each token's content-richness as keep-or-drop probability, the token selector effectively discards uninformative tokens, naturally allowing data-dependent compression. Applying adaptive tokenization to diffusion models may cause spatial misalignment, as token dropping can disturb the original spatio-temporal structure. To alleviate this issue, we propose two position-prediction strategies: cascaded and joint generation, to ensure spatial consistency. We empirically show that our model achieves strong reconstruction and generation quality at a state-of-the-art compression ratio. Further analysis on video data reveals that this improvement is primarily achieved by reducing spatio-temporal redundancy and removing uninformative tokens, as supported by both quantitative and qualitative results.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.24293 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.24293 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.24293 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.