r/LocalLLaMA · · 2 min read

[Paper] Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

[Paper] Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling

Hardware-agnostic strategies for accelerating text-to-image diffusion, such as timestep distillation and feature caching, can reduce inference time without custom kernels or system-level optimization. Among them, multi-resolution generation strategies have recently received broad attention, attaining more than 5x speedup without any training. However, the design of performing upsampling in the latent space, together with the selective modification of partial regions, causes these methods to exhibit noticeable blurring or artifacts. To this end, we propose MrFlow, a training-free multi-resolution acceleration strategy for pretrained flow-matching models built upon a staged low-to-high-resolution pipeline. MrFlow first rapidly generates the main structure at low resolution, then performs super-resolution in the pixel space using a lightweight pretrained GAN-based model, subsequently injects low-strength noise to enable high-frequency resampling, and finally refines the details at high resolution. Quantitative and qualitative results on FLUX.1-dev and Qwen-Image show that MrFlow exploits the quadratic token reduction and reduced step requirement of low-resolution sampling to achieve 10x end-to-end acceleration while keeping OneIG within a 1% gap relative to that before acceleration, significantly surpassing other training-free acceleration strategies, and requiring no training or runtime dynamic identification whatsoever. MrFlow can further be directly combined orthogonally with pre-trained timestep distillation strategies, achieving even higher generation acceleration of up to 25x.

Highlights

  • Training-free deployment. No finetuning, learned upsampler, or model-specific retraining is required.
  • No custom kernels. The implementation uses standard PyTorch, Diffusers pipelines, and scheduler controls.
  • Strong aggressive-speed regime. MrFlow reaches more than 10x end-to-end speedup on Qwen-Image while preserving visual quality.
  • Works with distilled models. The same pipeline can be combined with pretrained timestep-distilled models such as Pi-Flow and FLUX-schnell.
  • Compact staged design. The implementation transfers across Qwen-Image, FLUX.1-dev, FLUX.2 Klein, and Z-Image families.

News

  • [2026/07] 💡 We add a Practical Tips section and encourage everyone to share useful observations and takeaways with each other.
  • [2026/07] 🌱 We add a community contribution area and welcome developers to share MrFlow ports, workflows, and experiments with each other.
  • [2026/07] 📰 MrFlow is featured on Hugging Face Daily Papers.
  • [2026/07] ⚡ We release the MrFlow ComfyUI plugin.
  • [2026/07] 🔥 The MrFlow paper is available on arXiv, and the source code is released.

Representative end-to-end speedups:

Backbone Setting End-to-end speedup
FLUX.1-dev 12 + 1 8.25x
Qwen-Image 12 + 1 10.3x
FLUX.2 Klein Base 9B 12 + 1 8.79x
Z-Image-Turbo 8 + 1 21.0x
Qwen-Image + Pi-Flow 4 + 1 up to 25x

Speedups are measured end to end, including text encoding, VAE encode/decode, super-resolution, noise preparation, and diffusion forward passes.

arXiv : https://arxiv.org/abs/2607.01642

Full Paper : https://arxiv.org/pdf/2607.01642

HuggingFace : https://huggingface.co/Xingyu-Zheng/MrFlow

GitHub : https://github.com/Xingyu-Zheng/MrFlow

submitted by /u/pmttyji
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA