Inference-time calibration method for a wide range of diffusion / flow matching models with minimal computational overhead.</p>\n","updatedAt":"2026-07-27T09:03:59.763Z","author":{"_id":"65cf34322a811b355954354b","avatarUrl":"/avatars/fa81874938050b8a57be9b40255d5874.svg","fullname":"YuyaKobayashi","name":"u-kob","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7567868232727051},"editors":["u-kob"],"editorAvatarUrls":["/avatars/fa81874938050b8a57be9b40255d5874.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.22091","authors":[{"_id":"6a670820ab9cdf9be5794b31","user":{"_id":"65cf34322a811b355954354b","avatarUrl":"/avatars/fa81874938050b8a57be9b40255d5874.svg","isPro":false,"fullname":"YuyaKobayashi","user":"u-kob","type":"user","name":"u-kob"},"name":"Yuya Kobayashi","status":"claimed_verified","statusLastChangedAt":"2026-07-27T08:45:04.469Z","hidden":false},{"_id":"6a670820ab9cdf9be5794b32","name":"Masato Ishii","hidden":false},{"_id":"6a670820ab9cdf9be5794b33","name":"Yuhta Takida","hidden":false},{"_id":"6a670820ab9cdf9be5794b34","name":"Takashi Shibuya","hidden":false},{"_id":"6a670820ab9cdf9be5794b35","name":"Yuki Mitsufuji","hidden":false}],"publishedAt":"2026-07-24T00:00:00.000Z","submittedOnDailyAt":"2026-07-27T00:00:00.000Z","title":"Spectral Prior for Reducing Exposure Bias in Diffusion Models","submittedOnDailyBy":{"_id":"65cf34322a811b355954354b","avatarUrl":"/avatars/fa81874938050b8a57be9b40255d5874.svg","isPro":false,"fullname":"YuyaKobayashi","user":"u-kob","type":"user","name":"u-kob"},"summary":"Diffusion models typically suffer from error accumulation during iterative sampling, commonly referred to as exposure bias. We reveal systematic frequency-dependent discrepancies between training and inference, which can be interpreted as frequency-dependent SNR error. Crucially, the direction of this mismatch varies across models and timesteps, indicating that fixed correction rules do not generalize. We propose Spectral Alignment (SPA), a lightweight, guidance-based method that calibrates the power spectrum of intermediate predictions to a pre-computed prior. Our approach consists of two stages: (1) offline fitting of a parametric spectrum model from training data, and (2) inference-time guidance via efficient FFT-based gradient computation. SPA introduces minimal computational overhead (3-4\\%) and is complementary to Classifier-Free Guidance (CFG). We demonstrate consistent improvements across diverse architectures, from pixel-space models (DDPM, ADM) to latent diffusion models (SD2.0, SDXL) and flow-matching models (SD3.5, FLUX). Our implementation is available at https://github.com/SonyResearch/SPA.","upvotes":3,"discussionId":"6a670820ab9cdf9be5794b36","githubRepo":"https://github.com/SonyResearch/SPA","githubRepoAddedBy":"user","githubStars":0,"organization":{"_id":"6304f161c2f4f2d4929d52d7","name":"Sony","fullname":"Sony","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/61ac8f8a00d01045fca0ad2f/zpgaCDJEHxA8kYNDtMKnv.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65cf34322a811b355954354b","avatarUrl":"/avatars/fa81874938050b8a57be9b40255d5874.svg","isPro":false,"fullname":"YuyaKobayashi","user":"u-kob","type":"user"},{"_id":"674545617e0ea169b0c471d1","avatarUrl":"/avatars/a422e3efa5fb1c3f2c6c0997c412b088.svg","isPro":false,"fullname":"Masato Ishii","user":"mi141","type":"user"},{"_id":"6650773ca6acfdd2aba7d486","avatarUrl":"/avatars/d297886ea60dbff98a043caf825820ed.svg","isPro":false,"fullname":"Takashi Shibuya","user":"TakashiShibuyaSony","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6304f161c2f4f2d4929d52d7","name":"Sony","fullname":"Sony","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/61ac8f8a00d01045fca0ad2f/zpgaCDJEHxA8kYNDtMKnv.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.22091.md","query":{}}">
Spectral Prior for Reducing Exposure Bias in Diffusion Models
Abstract
Diffusion models typically suffer from error accumulation during iterative sampling, commonly referred to as exposure bias. We reveal systematic frequency-dependent discrepancies between training and inference, which can be interpreted as frequency-dependent SNR error. Crucially, the direction of this mismatch varies across models and timesteps, indicating that fixed correction rules do not generalize. We propose Spectral Alignment (SPA), a lightweight, guidance-based method that calibrates the power spectrum of intermediate predictions to a pre-computed prior. Our approach consists of two stages: (1) offline fitting of a parametric spectrum model from training data, and (2) inference-time guidance via efficient FFT-based gradient computation. SPA introduces minimal computational overhead (3-4\%) and is complementary to Classifier-Free Guidance (CFG). We demonstrate consistent improvements across diverse architectures, from pixel-space models (DDPM, ADM) to latent diffusion models (SD2.0, SDXL) and flow-matching models (SD3.5, FLUX). Our implementation is available at https://github.com/SonyResearch/SPA.
Community
Inference-time calibration method for a wide range of diffusion / flow matching models with minimal computational overhead.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.22091 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.22091 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.22091 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.