<strong>Mage-Flow</strong> is a compact <strong>4B-scale generative stack</strong> for efficient <strong>text-to-image generation</strong> and <strong>instruction-based image editing</strong>. Instead of scaling to tens of billions of parameters, Mage-Flow reaches state-of-the-art-competitive quality through careful <strong>tokenizer–backbone–system co-design</strong>, so it stays fast, memory-light, and easy to fine-tune under realistic compute budgets.</p>\n<p>The stack is built from <strong>two shared, co-designed components</strong>:</p>\n<ul>\n<li><strong>Mage-VAE</strong> — a lightweight, high-fidelity latent tokenizer (one-step diffusion encode/decode with anchor-latent KL regularization).</li>\n<li><strong>NR-MMDiT</strong> — a shared 4B <strong>Native-Resolution Multimodal Diffusion Transformer</strong>, trained with rectified flow matching in the Mage-VAE latent space.</li>\n</ul>\n<p>Together with native-resolution packing and a fused-kernel training infrastructure, this shared stack powers <strong>two model instantiations</strong>: <strong>Mage-Flow</strong> for text-to-image generation and <strong>Mage-Flow-Edit</strong> for instruction-based image editing. Each ships in <strong>Base</strong>, <strong>RL-aligned</strong>, and <strong>4-step Turbo</strong> variants.</p>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"✨-highlights\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#✨-highlights\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\t✨ Highlights\n\t</span>\n</h2>\n<ul>\n<li><strong>Compact & competitive.</strong> A single 4B family for generation <em>and</em> editing that matches or beats much larger open systems (Qwen-Image 20B, Z-Image 6B, FLUX.2 32B, FireRed-Image-Edit 20B).</li>\n<li><strong>Efficient tokenizer.</strong> Mage-VAE matches FLUX.2-VAE reconstruction fidelity while using <strong>~12× / ~22× fewer encode / decode MACs per pixel</strong>, removing the VAE as the high-resolution bottleneck.</li>\n<li><strong>Native resolution.</strong> One checkpoint generates from <strong>512 to 2048</strong> on any aspect ratio, including extreme <strong>4:1</strong> (e.g. <code>512×2048</code>, <code>2048×512</code>).</li>\n<li><strong>System-level speed.</strong> Native-resolution packing (FlashAttention var-len + per-sample 2D RoPE) + fused CUDA kernels achieve ~2.5× faster training; CFG's conditional/unconditional branches run in <strong>one</strong> packed forward.</li>\n<li><strong>Full family.</strong> <strong>Base</strong>, <strong>RL-aligned</strong>, and <strong>4-step Turbo</strong> variants for both generation and editing.</li>\n<li><strong>Versatile editing.</strong> Mage-Flow-Edit supports semantic content editing, appearance transformation, image restoration, and structure-aware outputs within a unified image-and-text-conditioned model. See the report's editing galleries.</li>\n<li><strong>Interactive latency.</strong> At <code>1024²</code> on a single A100: <strong>Mage-Flow-Turbo 0.59 s/image</strong>, <strong>Mage-Flow-Edit-Turbo 1.02 s/edit</strong>, peak memory <strong>~18–20 GB</strong> (lowest among compared systems).</li>\n</ul>\n","updatedAt":"2026-07-22T06:32:51.777Z","author":{"_id":"64338d1c4521083b9d2d21da","avatarUrl":"/avatars/54b809021d794f1c4b762fbc5d0c7c90.svg","fullname":"Xinjie","name":"Xinjie-Q","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.7605715394020081},"editors":["Xinjie-Q"],"editorAvatarUrls":["/avatars/54b809021d794f1c4b762fbc5d0c7c90.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.19064","authors":[{"_id":"6a6021d57e7f152167e4704a","name":"Xinjie Zhang","hidden":false},{"_id":"6a6021d57e7f152167e4704b","name":"Peng Zhang","hidden":false},{"_id":"6a6021d57e7f152167e4704c","name":"Shicheng Zheng","hidden":false},{"_id":"6a6021d57e7f152167e4704d","name":"Jinghao Guo","hidden":false},{"_id":"6a6021d57e7f152167e4704e","name":"Zhaoyang Jia","hidden":false},{"_id":"6a6021d57e7f152167e4704f","user":{"_id":"649aa367c6cf3cc95bc1b7f6","avatarUrl":"/avatars/4bf5446c261eab08fc06caebf4c5779a.svg","isPro":false,"fullname":"Yifei Shen","user":"yshenaw","type":"user","name":"yshenaw"},"name":"Yifei Shen","status":"claimed_verified","statusLastChangedAt":"2026-07-22T08:45:04.664Z","hidden":false},{"_id":"6a6021d57e7f152167e47050","name":"Xun Guo","hidden":false},{"_id":"6a6021d57e7f152167e47051","user":{"_id":"66e391a5021730e4ead995eb","avatarUrl":"/avatars/43ea3085b77ad9b08d21a8642156e574.svg","isPro":false,"fullname":"Luo Yuxuan","user":"LoYuXrqw","type":"user","name":"LoYuXrqw"},"name":"Yuxuan Luo","status":"claimed_verified","statusLastChangedAt":"2026-07-22T07:39:29.268Z","hidden":false},{"_id":"6a6021d57e7f152167e47052","name":"Jiahao Li","hidden":false},{"_id":"6a6021d57e7f152167e47053","name":"Wenxuan Xie","hidden":false},{"_id":"6a6021d57e7f152167e47054","user":{"_id":"646e1ef5075bbcc48ddf21e8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646e1ef5075bbcc48ddf21e8/g-nFu-plmEdTpnAJh_pUx.png","isPro":false,"fullname":"Pu Fanyi","user":"pufanyi","type":"user","name":"pufanyi"},"name":"Fanyi Pu","status":"claimed_verified","statusLastChangedAt":"2026-07-22T07:39:29.257Z","hidden":false},{"_id":"6a6021d57e7f152167e47055","name":"Xiaoyi Zhang","hidden":false},{"_id":"6a6021d57e7f152167e47056","name":"Kaichen Zhang","hidden":false},{"_id":"6a6021d57e7f152167e47057","name":"Zongyu Guo","hidden":false},{"_id":"6a6021d57e7f152167e47058","user":{"_id":"65a2a384bfaec7e7cae41d27","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65a2a384bfaec7e7cae41d27/pXmEXk_7q9-Y-ceLrR9ef.png","isPro":false,"fullname":"Tianci Bi","user":"tiancibi","type":"user","name":"tiancibi"},"name":"Tianci Bi","status":"claimed_verified","statusLastChangedAt":"2026-07-22T07:39:29.277Z","hidden":false},{"_id":"6a6021d57e7f152167e47059","name":"Dongnan Gui","hidden":false},{"_id":"6a6021d57e7f152167e4705a","name":"Zhening Liu","hidden":false},{"_id":"6a6021d57e7f152167e4705b","name":"Zimo Wen","hidden":false},{"_id":"6a6021d57e7f152167e4705c","name":"Zihan Zheng","hidden":false},{"_id":"6a6021d57e7f152167e4705d","name":"Senqiao Yang","hidden":false},{"_id":"6a6021d57e7f152167e4705e","name":"Xiao Li","hidden":false},{"_id":"6a6021d57e7f152167e4705f","name":"Jinglu Wang","hidden":false},{"_id":"6a6021d57e7f152167e47060","name":"Bin Li","hidden":false},{"_id":"6a6021d57e7f152167e47061","name":"Yan Lu","hidden":false}],"publishedAt":"2026-07-21T00:00:00.000Z","submittedOnDailyAt":"2026-07-22T00:00:00.000Z","title":"Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing","submittedOnDailyBy":{"_id":"64338d1c4521083b9d2d21da","avatarUrl":"/avatars/54b809021d794f1c4b762fbc5d0c7c90.svg","isPro":false,"fullname":"Xinjie","user":"Xinjie-Q","type":"user","name":"Xinjie-Q"},"summary":"Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a lightweight high-fidelity latent tokenizer, and a Native-Resolution Multimodal Diffusion Transformer trained with rectified flow matching. Mage-VAE uses one-step diffusion-style encoding and decoding with anchor-latent regularization, preserving the reconstruction quality of strong public VAEs while reducing tokenization cost by more than an order of magnitude. Together with native-resolution packing and stack-level CUDA kernel fusion, the stack supports flexible-resolution training and improves end-to-end training throughput by about 2.5times. Built on this foundation, we develop a complete model family with Base, RL-aligned, and Turbo variants for both generation and editing. Diffusion-NFT improves prompt following, text rendering, aesthetic quality, and editing fidelity, while few-step distillation with adversarial perceptual guidance produces 4-step Turbo models for low-latency inference. Despite its compact scale, Mage-Flow and Mage-Flow-Edit achieves competitive performance across standard generation and editing benchmarks. More importantly, the Turbo variants make high-resolution generation and editing practical for interactive use: at 1024^2 resolution on a single NVIDIA A100 GPU, Mage-Flow-Turbo generates an image in 0.59s, and Mage-Flow-Edit-Turbo edits an image in 1.02s, while maintaining a small memory footprint. These results show that careful tokenizer--backbone--system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.","upvotes":35,"discussionId":"6a6021d57e7f152167e47062","projectPage":"https://microsoft.github.io/Mage/","githubRepo":"https://github.com/microsoft/Mage","githubRepoAddedBy":"user","githubStars":4,"organization":{"_id":"5e6485f787403103f9f1055e","name":"microsoft","fullname":"Microsoft","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583646260758-5e64858c87403103f9f1055d.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64338d1c4521083b9d2d21da","avatarUrl":"/avatars/54b809021d794f1c4b762fbc5d0c7c90.svg","isPro":false,"fullname":"Xinjie","user":"Xinjie-Q","type":"user"},{"_id":"66e391a5021730e4ead995eb","avatarUrl":"/avatars/43ea3085b77ad9b08d21a8642156e574.svg","isPro":false,"fullname":"Luo Yuxuan","user":"LoYuXrqw","type":"user"},{"_id":"6970c006897d3834ff6bb385","avatarUrl":"/avatars/c627588cac0ae0c6f689800862290219.svg","isPro":false,"fullname":"gg dsf","user":"ggg93949943","type":"user"},{"_id":"649aa367c6cf3cc95bc1b7f6","avatarUrl":"/avatars/4bf5446c261eab08fc06caebf4c5779a.svg","isPro":false,"fullname":"Yifei Shen","user":"yshenaw","type":"user"},{"_id":"64bb77e786e7fb5b8a317a43","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64bb77e786e7fb5b8a317a43/J0jOrlZJ9gazdYaeSH2Bo.png","isPro":false,"fullname":"kcz","user":"kcz358","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"645db15ff4f49de580a10269","avatarUrl":"/avatars/ea1bdd7a478f4c4a7b3e134c4330ec78.svg","isPro":false,"fullname":"snowflakewang","user":"SnowflakeWang","type":"user"},{"_id":"69fae2f570585ad496262741","avatarUrl":"/avatars/f4c0b4a39cf28f35ad1b030248ee6fdc.svg","isPro":false,"fullname":"AnonymousSubmitt","user":"AnonymousSubmitt","type":"user"},{"_id":"652d06833b5997ed71ce5c46","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/652d06833b5997ed71ce5c46/O_D6bpa5mGxLA7uCjmVCG.jpeg","isPro":false,"fullname":"Zhongang Cai","user":"caizhongang","type":"user"},{"_id":"667aec90c06f44945546fc58","avatarUrl":"/avatars/c87ec4f74f6216c99685ace0b9e9080f.svg","isPro":false,"fullname":"Wen","user":"caes0r","type":"user"},{"_id":"646e1ef5075bbcc48ddf21e8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646e1ef5075bbcc48ddf21e8/g-nFu-plmEdTpnAJh_pUx.png","isPro":false,"fullname":"Pu Fanyi","user":"pufanyi","type":"user"},{"_id":"6478679d7b370854241b2ad8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6478679d7b370854241b2ad8/dBczWYYdfEt9tQcnVGhQk.jpeg","isPro":false,"fullname":"xiangan","user":"xiangan","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"5e6485f787403103f9f1055e","name":"microsoft","fullname":"Microsoft","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583646260758-5e64858c87403103f9f1055d.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.19064.md","query":{}}">
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
Published on Jul 21
· Submitted by Xinjie on Jul 22 Abstract
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a lightweight high-fidelity latent tokenizer, and a Native-Resolution Multimodal Diffusion Transformer trained with rectified flow matching. Mage-VAE uses one-step diffusion-style encoding and decoding with anchor-latent regularization, preserving the reconstruction quality of strong public VAEs while reducing tokenization cost by more than an order of magnitude. Together with native-resolution packing and stack-level CUDA kernel fusion, the stack supports flexible-resolution training and improves end-to-end training throughput by about 2.5times. Built on this foundation, we develop a complete model family with Base, RL-aligned, and Turbo variants for both generation and editing. Diffusion-NFT improves prompt following, text rendering, aesthetic quality, and editing fidelity, while few-step distillation with adversarial perceptual guidance produces 4-step Turbo models for low-latency inference. Despite its compact scale, Mage-Flow and Mage-Flow-Edit achieves competitive performance across standard generation and editing benchmarks. More importantly, the Turbo variants make high-resolution generation and editing practical for interactive use: at 1024^2 resolution on a single NVIDIA A100 GPU, Mage-Flow-Turbo generates an image in 0.59s, and Mage-Flow-Edit-Turbo edits an image in 1.02s, while maintaining a small memory footprint. These results show that careful tokenizer--backbone--system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.
Community
Mage-Flow is a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. Instead of scaling to tens of billions of parameters, Mage-Flow reaches state-of-the-art-competitive quality through careful tokenizer–backbone–system co-design, so it stays fast, memory-light, and easy to fine-tune under realistic compute budgets.
The stack is built from two shared, co-designed components:
- Mage-VAE — a lightweight, high-fidelity latent tokenizer (one-step diffusion encode/decode with anchor-latent KL regularization).
- NR-MMDiT — a shared 4B Native-Resolution Multimodal Diffusion Transformer, trained with rectified flow matching in the Mage-VAE latent space.
Together with native-resolution packing and a fused-kernel training infrastructure, this shared stack powers two model instantiations: Mage-Flow for text-to-image generation and Mage-Flow-Edit for instruction-based image editing. Each ships in Base, RL-aligned, and 4-step Turbo variants.
✨ Highlights
- Compact & competitive. A single 4B family for generation and editing that matches or beats much larger open systems (Qwen-Image 20B, Z-Image 6B, FLUX.2 32B, FireRed-Image-Edit 20B).
- Efficient tokenizer. Mage-VAE matches FLUX.2-VAE reconstruction fidelity while using ~12× / ~22× fewer encode / decode MACs per pixel, removing the VAE as the high-resolution bottleneck.
- Native resolution. One checkpoint generates from 512 to 2048 on any aspect ratio, including extreme 4:1 (e.g.
512×2048, 2048×512).
- System-level speed. Native-resolution packing (FlashAttention var-len + per-sample 2D RoPE) + fused CUDA kernels achieve ~2.5× faster training; CFG's conditional/unconditional branches run in one packed forward.
- Full family. Base, RL-aligned, and 4-step Turbo variants for both generation and editing.
- Versatile editing. Mage-Flow-Edit supports semantic content editing, appearance transformation, image restoration, and structure-aware outputs within a unified image-and-text-conditioned model. See the report's editing galleries.
- Interactive latency. At
1024² on a single A100: Mage-Flow-Turbo 0.59 s/image, Mage-Flow-Edit-Turbo 1.02 s/edit, peak memory ~18–20 GB (lowest among compared systems).
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.19064 in a dataset README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.