Budget-constrained, verifier-guided best-of-N generation for diffusion models. Given a prompt and a wall-clock budget, Flash-BoN drafts cheap candidates (TaylorSeer-style caching + layer/timestep skipping), selects the best with a pairwise vision-language tournament, refines the top-k to full quality by resuming denoising from a captured state, and scores them with VQAScore.</p>\n","updatedAt":"2026-07-10T02:17:02.145Z","author":{"_id":"5f7fbd813e94f16a85448745","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1649681653581-5f7fbd813e94f16a85448745.jpeg","fullname":"Sayak Paul","name":"sayakpaul","type":"user","isPro":true,"isHf":true,"isHfAdmin":false,"isMod":false,"followerCount":987,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583856921041-5dd96eb166059660ed1ee413.png","fullname":"Hugging Face","name":"huggingface","type":"org","isHf":true,"details":"The AI community building the future.","plan":"team"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8180314898490906},"editors":["sayakpaul"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1649681653581-5f7fbd813e94f16a85448745.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.04461","authors":[{"_id":"6a50558c75fd3d966bd45d6c","name":"Ruchit Rawal","hidden":false},{"_id":"6a50558c75fd3d966bd45d6d","name":"Reza Shirkavand","hidden":false},{"_id":"6a50558c75fd3d966bd45d6e","user":{"_id":"5f7fbd813e94f16a85448745","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1649681653581-5f7fbd813e94f16a85448745.jpeg","isPro":true,"fullname":"Sayak Paul","user":"sayakpaul","type":"user","name":"sayakpaul"},"name":"Sayak Paul","status":"claimed_verified","statusLastChangedAt":"2026-07-10T07:45:04.415Z","hidden":false},{"_id":"6a50558c75fd3d966bd45d6f","name":"Yuxin Wen","hidden":false},{"_id":"6a50558c75fd3d966bd45d70","name":"Heng Huang","hidden":false},{"_id":"6a50558c75fd3d966bd45d71","name":"Yizheng Chen","hidden":false},{"_id":"6a50558c75fd3d966bd45d72","name":"Tom Goldstein","hidden":false},{"_id":"6a50558c75fd3d966bd45d73","name":"Gowthami Somepalli","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/5f7fbd813e94f16a85448745/V0IyG__8_qq6674N_Ri5a.png"],"publishedAt":"2026-07-05T00:00:00.000Z","submittedOnDailyAt":"2026-07-10T00:00:00.000Z","title":"Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models","submittedOnDailyBy":{"_id":"5f7fbd813e94f16a85448745","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1649681653581-5f7fbd813e94f16a85448745.jpeg","isPro":true,"fullname":"Sayak Paul","user":"sayakpaul","type":"user","name":"sayakpaul"},"summary":"Inference-time scaling for text-to-image generation has progressed from simple Best-of-N (BoN) sampling to guided search methods that verify and steer candidate trajectories at intermediate denoising steps. These approaches focus on when and how often to verify during denoising but largely treat the cost of generation itself as fixed. Moreover, the standard practice of comparing methods by number of function evaluations (NFEs) counts only denoising forward passes and ignores verifier overhead, which can distort efficiency rankings. We show that under wall-clock evaluation, simple BoN already matches or outperforms several guided search techniques, suggesting that compute is better spent on broader exploration than on repeated intermediate verification. This motivates Flash-BoN, which generates a large pool of inexpensive draft candidates by combining three complementary acceleration knobs: timestep truncation, layer skipping, and activation proxies into a single configuration optimized once per model. An efficient multi-stage verification procedure then identifies the most promising draft, which is refined at full quality. Across three benchmarks and three model scales, Flash-BoN consistently outperforms all baselines under fixed wall-clock budgets, with gains that grow at larger model scales (+8% AUC). We further show that our strategy combines well and improves existing orthogonal techniques such as reflection-based prompt optimization (+16% AUC). The gains correlate with increased candidate diversity, which also enables draft-guided selection to accelerate RL post-training convergence.","upvotes":2,"discussionId":"6a50558c75fd3d966bd45d74","projectPage":"https://flash-bon.github.io/","githubRepo":"https://github.com/flash-bon/flash-bon","githubRepoAddedBy":"user","ai_summary":"Flash-BoN improves text-to-image generation efficiency by using inexpensive draft candidates generated through timestep truncation, layer skipping, and activation proxies, followed by multi-stage verification that outperforms existing methods under fixed wall-clock budgets.","ai_keywords":["Best-of-N","guided search methods","denoising steps","function evaluations","wall-clock evaluation","Flash-BoN","timestep truncation","layer skipping","activation proxies","multi-stage verification","candidate diversity","RL post-training convergence"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":5,"organization":{"_id":"63dbddbe06e5ca38798321bd","name":"tomg-group-umd","fullname":"Tom Goldstein's Lab at University of Maryland, College Park","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1675353480936-63d98af1897746d6496177df.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"67a385582d7dcbd140ee70d3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67a385582d7dcbd140ee70d3/DSsEv7_qnh8M72iZ_cZo3.webp","isPro":false,"fullname":"Stephen Moore","user":"morongosteve","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63dbddbe06e5ca38798321bd","name":"tomg-group-umd","fullname":"Tom Goldstein's Lab at University of Maryland, College Park","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1675353480936-63d98af1897746d6496177df.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.04461.md","query":{}}">
Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models
Abstract
Flash-BoN improves text-to-image generation efficiency by using inexpensive draft candidates generated through timestep truncation, layer skipping, and activation proxies, followed by multi-stage verification that outperforms existing methods under fixed wall-clock budgets.
Inference-time scaling for text-to-image generation has progressed from simple Best-of-N (BoN) sampling to guided search methods that verify and steer candidate trajectories at intermediate denoising steps. These approaches focus on when and how often to verify during denoising but largely treat the cost of generation itself as fixed. Moreover, the standard practice of comparing methods by number of function evaluations (NFEs) counts only denoising forward passes and ignores verifier overhead, which can distort efficiency rankings. We show that under wall-clock evaluation, simple BoN already matches or outperforms several guided search techniques, suggesting that compute is better spent on broader exploration than on repeated intermediate verification. This motivates Flash-BoN, which generates a large pool of inexpensive draft candidates by combining three complementary acceleration knobs: timestep truncation, layer skipping, and activation proxies into a single configuration optimized once per model. An efficient multi-stage verification procedure then identifies the most promising draft, which is refined at full quality. Across three benchmarks and three model scales, Flash-BoN consistently outperforms all baselines under fixed wall-clock budgets, with gains that grow at larger model scales (+8% AUC). We further show that our strategy combines well and improves existing orthogonal techniques such as reflection-based prompt optimization (+16% AUC). The gains correlate with increased candidate diversity, which also enables draft-guided selection to accelerate RL post-training convergence.
Community
Budget-constrained, verifier-guided best-of-N generation for diffusion models. Given a prompt and a wall-clock budget, Flash-BoN drafts cheap candidates (TaylorSeer-style caching + layer/timestep skipping), selects the best with a pairwise vision-language tournament, refines the top-k to full quality by resuming denoising from a captured state, and scores them with VQAScore.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.04461 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.04461 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.04461 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.