News / #image-gen Tag Image Gen 191 articles archived under #image-gen · RSS Sign in to follow arXiv — Machine Learning research 1mo ago Flux-OPD: On-Policy Distillation with Evolving Contexts arXiv:2607.28022v1 Announce Type: new Abstract: Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision… 13 Hugging Face Daily Papers research 1mo ago MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing Abstract Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities,… 12 Hugging Face Daily Papers research 1mo ago Flux-OPD: On-Policy Distillation with Evolving Contexts Abstract Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student,… 24 Hugging Face Daily Papers research 2mo ago TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward Abstract Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that modify the sampling trajectory to produce images more faithful to complex compositional prompts. We present TILT, a training-free framework… 27 arXiv — Machine Learning research 2mo ago Learning Sampling Parameters for Diffusion Models arXiv:2607.23488v1 Announce Type: new Abstract: Text-to-image diffusion models expose many inference-time sampling parameters, including prompts, negative prompts, classifier-free guidance scales, and noise schedules. These parameters are typically manually chosen once and then… 29 TechCrunch — AI news-outlet 2mo ago Midjourney acquired the astrology app Co-Star The AI lab Midjourney continues to expand its purview beyond image and video generation. 11 r/LocalLLaMA community 2mo ago FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence Introducing FLUX 3. One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style. Blog Post : https://bfl.ai/blog/flux-3   submitted by   /u/pmttyji [link]   [comments] 23 Hacker News — AI on Front Page community 2mo ago Flux 3 X Mimic: The Next Generation of Video-Action Models Article URL: https://bfl.ai/blog/flux-3-mimic Comments URL: https://news.ycombinator.com/item?id=49033127 Points: 240 # Comments: 32 24 Hacker News — AI on Front Page community 2mo ago Flux 3 Article URL: https://bfl.ai/blog/flux-3 Comments URL: https://news.ycombinator.com/item?id=49031796 Points: 313 # Comments: 77 36 Latent.Space news-outlet 2mo ago [AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model A HUGE win for BFL! 22 arXiv — Machine Learning research 2mo ago Neural Operator Surrogates for Two-Dimensional Neutron Flux Estimation arXiv:2607.19388v1 Announce Type: new Abstract: This work extends our one-dimensional single-sweep neural-operator studies to two dimensions. We consider one-group transport with isotropic scattering. As in the one-dimensional work, we use Fourier neural operators (FNOs) to… 38 r/LocalLLaMA community 2mo ago Mage-Flow - An Efficient Native-Resolution Foundation Model for Image Generation and Editing - Microsoft Models: (Check Model cards for so much sample demo images) https://huggingface.co/microsoft/Mage-Flow https://huggingface.co/microsoft/Mage-Flow-Turbo https://huggingface.co/microsoft/Mage-Flow-Edit Mage-Flow is a compact 4B-scale generative stack for efficient text-to-image… 30 r/MachineLearning community 2mo ago Anyone heading to Jeju for KDD? Let's meet up! 🙋[D] Hey all! Is anyone else going to be at KDD in Jeju? Would love to connect with fellow attendees. I work on interpretability, fairness, and editing of text-to-image models, so I'd especially love to meet people working in these areas. But honestly, we can chat about anything:… 4 Hugging Face Daily Papers research 2mo ago Appearance Pointers -- Multimodal Region Control of Diffusion Transformers Abstract Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers… 20 Hugging Face Daily Papers research 2mo ago Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers Abstract Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computation during denoising remains poorly understood. We introduce a causal interpretability framework for modern large-scale DiTs that combines attention… 8 Hugging Face Daily Papers research 2mo ago Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Abstract Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two… 23 arXiv — Machine Learning research 2mo ago Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation arXiv:2607.14962v1 Announce Type: new Abstract: Text-to-image (T2I) models can synthesize realistic, prompt-aligned images, yet samples generated for the same prompt often cover only a small subset of visually distinct modes. This limits the diversity of images, and for… 4 arXiv — Machine Learning research 2mo ago Inference-Time Concept Suppression and Video-Centric Evaluation for Text-to-Video Models arXiv:2607.14194v1 Announce Type: cross Abstract: Text-to-video (T2V) generators can synthesize realistic and temporally coherent videos, but controllably removing a target concept from a generator remains difficult. Unlike text-to-image concept erasure, T2V unlearning must… 5 Hugging Face Daily Papers research 2mo ago Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation Abstract We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference,… 10 arXiv — Machine Learning research 2mo ago SymbOmni: Evolving Agentic Omni Models via Symbolic Concept Learning arXiv:2607.12042v1 Announce Type: cross Abstract: Visual generation is increasingly ubiquitous in diverse domains, from text-to-image/video synthesis to multimodal interactive creation. Yet prevailing monolithic models remain fundamentally constrained by their inability to learn… 32 Hugging Face Daily Papers research 2mo ago Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Abstract In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of asking the MLLM to judge a generated image or answer decomposed verification… 27 Hugging Face Daily Papers research 2mo ago Latent-Identity Tuning in Text-to-Image Personalization Models Abstract Generating and editing a person's face demands high precision, as even minor modifications can significantly alter a subject's perceived identity. Current personalization and editing methods built on general-purpose text-to-image models, however, often lack the… 36 Hugging Face Daily Papers research 2mo ago From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models Abstract Pretrained diffusion transformers can be adapted for dense prediction tasks by mapping tokens to task-native outputs instead of generating RGB images, achieving state-of-the-art results with minimal additional parameters. Generated by Qwen/Qwen2.5-Coder-32B-Instruct… 6 r/LocalLLaMA community 2mo ago I built Flaxeo Image a local desktop ui for stable diffusion cpp Built around a recent sd.cpp release, aims to expose most of what the backend can do (generate, edit, video paths, models, hardware options), Windows + Linux builds GitHub: https://github.com/fabricio3g/FlaxeoUI   submitted by   /u/fabricio3g [link]   [comments] 6 Vercel — AI dev-tools 2mo ago Seedream 5.0 Pro is now available on AI Gateway Seedream 5.0 Pro is now available on AI Gateway . Seedream 5.0 Pro is an image generation and editing model. It generates images from text, rendering text without spelling errors and following typographic rules, and produces dense infographics with charts, timelines, and layouts… 20 arXiv — Machine Learning research 2mo ago AutoAnchor: Stable Diffusion Unlearning Using Cross-Attention as a Manifold Surrogate arXiv:2607.08337v1 Announce Type: new Abstract: Diffusion unlearning is essential for mitigating the generation of harmful or copyrighted content in text-to-image models. Current diffusion unlearning techniques determine the model update direction by either using alternatives of… 21 arXiv — NLP / Computation & Language research 2mo ago Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing arXiv:2607.08497v1 Announce Type: cross Abstract: Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, they repeatedly feed all historical visual and textual inputs into a shared… 23 Hugging Face Daily Papers research 2mo ago Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models Abstract Flash-BoN improves text-to-image generation efficiency by using inexpensive draft candidates generated through timestep truncation, layer skipping, and activation proxies, followed by multi-stage verification that outperforms existing methods under fixed wall-clock… 9 arXiv — Machine Learning research 2mo ago An Hybrid Quantum-Classical Diffusion Model for Image Generation arXiv:2607.07072v1 Announce Type: new Abstract: Quantum diffusion models provide a physics-consistent route to generative learning by formulating noising and denoising directly on quantum states. However, applying such models to classical high-dimensional data is constrained by… 9 arXiv — NLP / Computation & Language research 2mo ago Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies arXiv:2607.06963v1 Announce Type: cross Abstract: Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic, Google, Meta, Microsoft, Stability AI, respectively, are revolutionizing… 5 arXiv — Machine Learning research 2mo ago TILDE: TILt-based Distributional Erasure for Concept Unlearning arXiv:2607.06432v1 Announce Type: new Abstract: Concept unlearning in text-to-image diffusion models is critical for safe and practical deployment: with rising privacy concerns, copyright disputes, trademark constraints, and safety regulations, deployed systems must be able to… 22 Hugging Face Daily Papers research 2mo ago CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training Abstract A 3D latent diffusion model for chest CT generation that enables controlled synthesis of medical images with clinical attributes through adaptive layer normalization and reinforcement learning post-training. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Controllable… 35 r/LocalLLaMA community 2mo ago [Paper] Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling Hardware-agnostic strategies for accelerating text-to-image diffusion, such as timestep distillation and feature caching, can reduce inference time without custom kernels or system-level optimization. Among them, multi-resolution generation strategies have recently received… 25 TechCrunch — AI news-outlet 2mo ago Midjourney wants Hollywood studios to reveal the details of their AI usage As part of an ongoing legal dispute with three Hollywood studios, Midjourney is seeking to compel those studios to reveal how they use AI themselves. 24 r/MachineLearning community 2mo ago I built my 'first' flow matching image generator, here's what I learned [P] Today I put out my first flow matching image generation model! This is a toy example trained on a 2024 MPS Macbook Pro using a small sample of images—specifically, the Apple emoji library and their text labels. Because of this, it’s not a massive model (clocking in at ~4.7… 20 Hugging Face Daily Papers research 2mo ago Representation Distribution Matching for One-Step Visual Generation Abstract Representation Distribution Matching enables high-quality image generation by matching feature distributions under pretrained encoders, with improved performance through optimized batch sizes and multi-encoder evaluation metrics. Generated by… 33 Hugging Face Daily Papers research 2mo ago Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling Abstract MrFlow accelerates text-to-image diffusion by combining low-resolution generation with pixel-space super-resolution and noise injection, achieving up to 25x speedup without training or runtime modifications. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Hardware-agnostic… 6 Hugging Face Daily Papers research 2mo ago SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation Abstract Scientific image generation faces challenges in semantic alignment and logical reasoning, prompting the creation of SciIR-82k dataset and SciIR-Bench evaluation framework to improve scientific reasoning capabilities in text-to-image models. Generated by… 23 Hugging Face Daily Papers research 2mo ago DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation Abstract DataEvolver is a self-evolving multi-agent framework that improves text-rich image generation by leveraging feedback from rejected samples to iteratively enhance data quality. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Text-rich image generation is one of the most… 11 Hugging Face Daily Papers research 2mo ago InstanceControl: Controllable Complex Image Generation without Instance Labeling Abstract InstanceControl enables multi-instance image generation by using vision-language models to establish instance-level correspondences between text prompts and visual conditions, while employing adaptive mask refinement for improved accuracy. Generated by… 29 arXiv — Machine Learning research 2mo ago Quality-Aware Modulation for Diffusion Transformers arXiv:2606.30934v1 Announce Type: new Abstract: Modern text-to-image diffusion models, such as diffusion transformers (DiT), rely on timestep or prompt embeddings to modulate the strength of the denoising process in each timestep. While this modulation communicates the current… 31 Hugging Face Daily Papers research 3mo ago Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation Abstract ILLUME-X is a unified multimodal paradigm that enhances text-image generation through improved data efficiency, stable training processes, and comprehensive evaluation metrics. Generated by Qwen/Qwen2.5-Coder-32B-Instruct The advancement of generative AI models capable… 17 arXiv — Machine Learning research 3mo ago Can AI Draw Science? A Benchmark for Evaluating Scientific Figure Generation by Text-to-Image and Multimodal Models arXiv:2606.28406v1 Announce Type: new Abstract: Text-to-image and multimodal generative models are increasingly used to produce scientific figures such as mechanism diagrams, experimental-design schematics, conceptual frameworks, and graphical abstracts. Yet existing… 36 Hugging Face Daily Papers research 3mo ago Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting Abstract Flux-GS enables real-time high-fidelity 3D Gaussian Splatting on mobile platforms through efficient lighting representation, attribute-conditioned enhancement, and multi-view densification strategies. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Recent advances in 3D… 10 Hugging Face Daily Papers research 3mo ago Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis Abstract A masked discrete diffusion model for text-to-image synthesis that addresses limitations in token refinement and training efficiency through novel mechanisms and optimizations. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We propose Nemotron-Labs-Diffusion-Image, a… 25 Hugging Face Daily Papers research 3mo ago MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation Abstract MIMFlow combines Normalizing Flows with Masked Image Modeling to improve generative modeling by decoupling semantic representation from pixel-level details, achieving better performance with fewer tokens. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Normalizing Flows… 37 TechCrunch — AI news-outlet 3mo ago Gemini’s personalized AI image generation is now free for US users Google is expanding Gemini’s personalized AI image generation to eligible free users in the U.S., allowing the chatbot to create images based on your interests and data from connected Google apps. 29 r/LocalLLaMA community 3mo ago clark-labs/clark-air-sana-1.6b-1.58bit · Hugging Face A Sana 1.6B text-to-image transformer compressed to ternary (~1.85 bits/weight): 8.6× smaller than FP16, near-FP16 quality. Footprint (measured) Artifact Size vs FP16 What it is FP16 transformer 3.21 GB 1× (100%) reference Clark Air (packed) 374 MB 8.6× (≈12%) packed ternary (… 36 Hugging Face Daily Papers research 3mo ago Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation Abstract A unified agentic framework called Qwen-Image-Agent is proposed to address the context gap in text-to-image generation by progressively constructing complete generation context through planning, reasoning, searching, and memory mechanisms. Generated by… 22 arXiv — NLP / Computation & Language research 3mo ago DanceOPD: On-Policy Generative Field Distillation arXiv:2606.27377v1 Announce Type: cross Abstract: Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing. However, these capabilities are rarely naturally aligned and often conflict. For… 18 Page 2 of 4 · 191 articles ← Newer Older →