Pixel-Space Dense Prediction with Pretrained Text-to-Image Diffusion Models</p>\n","updatedAt":"2026-07-13T02:45:10.105Z","author":{"_id":"662a0f2d4bab737c1a279843","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/662a0f2d4bab737c1a279843/fC2p3mjMHkVpDQdEqkuR4.png","fullname":"Dengyang Jiang","name":"DyJiang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":16,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.41693150997161865},"editors":["DyJiang"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/662a0f2d4bab737c1a279843/fC2p3mjMHkVpDQdEqkuR4.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.06553","authors":[{"_id":"6a4ea2a848d70828b718dd33","user":{"_id":"65f2e3d1cec22d29ce41ef94","avatarUrl":"/avatars/a724073aa1ae5d72847955d2f39773fa.svg","isPro":false,"fullname":"Zanyi Wang","user":"xmz111","type":"user","name":"xmz111"},"name":"Zanyi Wang","status":"admin_assigned","statusLastChangedAt":"2026-07-09T19:04:40.333Z","hidden":false},{"_id":"6a4ea2a848d70828b718dd34","name":"Xin Lin","hidden":false},{"_id":"6a4ea2a848d70828b718dd35","name":"Haodong Li","hidden":false},{"_id":"6a4ea2a848d70828b718dd36","name":"Dengyang Jiang","hidden":false},{"_id":"6a4ea2a848d70828b718dd37","name":"Yijiang Li","hidden":false}],"publishedAt":"2026-07-09T00:00:00.000Z","submittedOnDailyAt":"2026-07-13T00:00:00.000Z","title":"From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models","submittedOnDailyBy":{"_id":"662a0f2d4bab737c1a279843","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/662a0f2d4bab737c1a279843/fC2p3mjMHkVpDQdEqkuR4.png","isPro":false,"fullname":"Dengyang Jiang","user":"DyJiang","type":"user","name":"DyJiang"},"summary":"Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining learns rich semantic, structural, and geometric priors. Existing generative and editing approaches reuse these priors by casting dense prediction as target generation: annotations such as depth, normals, alpha mattes, masks, and heatmaps are encoded into an RGB-trained VAE latent space and decoded back as image-like targets. We argue this inherits more of the generative output interface than dense prediction requires: unlike RGB synthesis, dense prediction asks for pixel-correct, task-native fields on the same image plane, not new RGB content to be rendered. Our key observation is that a pretrained DiT already organizes RGB inputs through a patch-to-token-to-patch lattice on the image plane, so each token indexes a fixed output patch whose channels can carry task-native quantities instead of RGB appearance. We instantiate this as ReChannel: we keep the VAE encoder for the DiT's input distribution but drop the target-side decoder, adapt the frozen DiT with task LoRA, and map each token to its p x p x K_t pixel-space patch through a shared token-local linear head--about 33K parameters, no spatial mixing. Using FLUX-Klein, we evaluate on six dense prediction tasks and over a dozen benchmarks. This minimal interface sets new state-of-the-art on trimap-free matting, KITTI depth, and referring segmentation, and stays competitive on normals, saliency, and pose. In a matched 4B setting it is more accurate and 2.48x faster than an edit-plus-latent-decode counterpart--dense perception can benefit from generative pretraining without inheriting its output interface.","upvotes":6,"discussionId":"6a4ea2a848d70828b718dd39","githubRepo":"https://github.com/xmz111/ReChannel","githubRepoAddedBy":"user","ai_summary":"Pretrained diffusion transformers can be adapted for dense prediction tasks by mapping tokens to task-native outputs instead of generating RGB images, achieving state-of-the-art results with minimal additional parameters.","ai_keywords":["text-to-image models","dense prediction","VAE latent space","DiT","task LoRA","token-local linear head","FLUX-Klein","trimap-free matting","KITTI depth","referring segmentation","normals","saliency","pose"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":9},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"662a0f2d4bab737c1a279843","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/662a0f2d4bab737c1a279843/fC2p3mjMHkVpDQdEqkuR4.png","isPro":false,"fullname":"Dengyang Jiang","user":"DyJiang","type":"user"},{"_id":"65f2e3d1cec22d29ce41ef94","avatarUrl":"/avatars/a724073aa1ae5d72847955d2f39773fa.svg","isPro":false,"fullname":"Zanyi Wang","user":"xmz111","type":"user"},{"_id":"6419309f22270b3ccf177c77","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6419309f22270b3ccf177c77/KQa1586iBBKqucUlfpuPp.jpeg","isPro":true,"fullname":"William Li","user":"williamium","type":"user"},{"_id":"641d211e353524fe41f16387","avatarUrl":"/avatars/01d8fec857faac05f15b772f65127565.svg","isPro":true,"fullname":"Haodong Li","user":"haodongli","type":"user"},{"_id":"669d003d84e6a96448daf9f5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/669d003d84e6a96448daf9f5/vxWJg1p91G-OhFeDwtmku.jpeg","isPro":false,"fullname":"Morax Cheng","user":"MoraxCheng","type":"user"},{"_id":"634cfebc350bcee9bed20a4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/634cfebc350bcee9bed20a4d/fN47nN5rhw-HJaFLBZWQy.png","isPro":false,"fullname":"Xingyi Yang","user":"adamdad","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.06553.md","query":{}}">
From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models
Abstract
Pretrained diffusion transformers can be adapted for dense prediction tasks by mapping tokens to task-native outputs instead of generating RGB images, achieving state-of-the-art results with minimal additional parameters.
Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining learns rich semantic, structural, and geometric priors. Existing generative and editing approaches reuse these priors by casting dense prediction as target generation: annotations such as depth, normals, alpha mattes, masks, and heatmaps are encoded into an RGB-trained VAE latent space and decoded back as image-like targets. We argue this inherits more of the generative output interface than dense prediction requires: unlike RGB synthesis, dense prediction asks for pixel-correct, task-native fields on the same image plane, not new RGB content to be rendered. Our key observation is that a pretrained DiT already organizes RGB inputs through a patch-to-token-to-patch lattice on the image plane, so each token indexes a fixed output patch whose channels can carry task-native quantities instead of RGB appearance. We instantiate this as ReChannel: we keep the VAE encoder for the DiT's input distribution but drop the target-side decoder, adapt the frozen DiT with task LoRA, and map each token to its p x p x K_t pixel-space patch through a shared token-local linear head--about 33K parameters, no spatial mixing. Using FLUX-Klein, we evaluate on six dense prediction tasks and over a dozen benchmarks. This minimal interface sets new state-of-the-art on trimap-free matting, KITTI depth, and referring segmentation, and stays competitive on normals, saliency, and pose. In a matched 4B setting it is more accurate and 2.48x faster than an edit-plus-latent-decode counterpart--dense perception can benefit from generative pretraining without inheriting its output interface.
Community
Pixel-Space Dense Prediction with Pretrained Text-to-Image Diffusion Models
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.06553 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.06553 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.06553 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.