News / #image-gen Tag Image Gen 160 articles archived under #image-gen · RSS Sign in to follow r/MachineLearning community 10h ago Reproducible canvas-aligned low-level patterns in somerandomllm-generated images and their possible relation to iterative editing artifacts [D] I may have stumbled onto something interesting while trying to figure out a recurring artifact in ChatGPT image generation and editing (maybe applicable to other models as well?). It started with a very practical problem: After several rounds of generative editing on portraits,… 15 arXiv — NLP / Computation & Language research 2d ago On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation arXiv:2608.11002v1 Announce Type: new Abstract: Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual performance gaps and language-specific effects… 37 Hugging Face Daily Papers research 2d ago Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation Abstract Atelier improves artist-grounded image generation by translating vague artistic intent into explicit control states that separate scene content from style, reducing reliance on stereotypical shortcuts. Generated by thinkingmachines/Inkling-Small Artist-grounded image… 20 arXiv — NLP / Computation & Language research 4d ago Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs arXiv:2608.06967v1 Announce Type: new Abstract: Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide visual-plan failures: an answer may appear creative while repeating familiar… 36 Simon Willison community 6d ago Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) On Wednesday I wrote about One-shotting a Raccoon Heist game using Claude Fable 5 , where I had Claude Fable 5 build a full working game from a premise I generated with GPT-3 and DALL-E four years ago . I decided to… 33 Simon Willison community 6d ago Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) On Wednesday I wrote about One-shotting a Raccoon Heist game using Claude Fable 5 , where I had Claude Fable 5 build a full working game from a premise I generated with GPT-3 and DALL-E four years ago . I decided to… 20 arXiv — NLP / Computation & Language research 7d ago GenGA: Editable and Data-Grounded Graphical Abstract Generation for Academic Papers arXiv:2608.05478v1 Announce Type: cross Abstract: Graphical Abstracts (GAs) visually summarize the key findings of academic papers, playing a crucial role in facilitating the understanding of research content. Recently, advancements in vision-language models and image generation… 38 Hugging Face Daily Papers research 7d ago ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Abstract Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, and the integration of external world knowledge. Existing efforts introduce agent… 36 arXiv — NLP / Computation & Language research 8d ago Simile Understanding in Text-to-Image Models: An Evaluation Framework arXiv:2608.04750v1 Announce Type: cross Abstract: Similes provide a compact and expressive way to describe visual characteristics in text prompts. Recent text-to-image models (t2i models) can produce visually compelling outputs from simile prompts, yet even frontier models… 12 Hugging Face Daily Papers research 8d ago Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models Abstract Leading open text-to-image models often carry complementary strengths: one may lead on preference-aligned aesthetics while another follows compositional instructions more faithfully. However, differences in their autoencoders and noise schedules make it difficult to… 23 Simon Willison community 8d ago One-shotting a Raccoon Heist game using Claude Fable 5 Back in 2024 I tweeted screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in Claude Code for web ) could build the entire game from the content… 29 Simon Willison community 8d ago One-shotting a Raccoon Heist game using Claude Fable 5 Back in 2022 I tweeted screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in Claude Code for web ) could build the entire game from the content… 30 Hugging Face Daily Papers research 9d ago PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs Abstract Scientific poster construction compresses a long multimodal paper into a readable, editable canvas. Existing systems hide request-level failures by scoring only completed outputs; direct image generation is not element-editable, while coding-agent workflows are costly.… 35 arXiv — Machine Learning research 9d ago A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics arXiv:2608.02965v1 Announce Type: new Abstract: Magnetic components in high-frequency, high-power-density converters are increasingly driven by non-sinusoidal flux-density waveforms with fast transitions, minor-loop operation, dc bias, and temperature variation. Under these… 26 Hugging Face Daily Papers research 9d ago Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Abstract Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and… 8 Hugging Face Daily Papers research 9d ago CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Abstract Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct properties: (1) how much visual information… 22 Hugging Face Daily Papers research 9d ago UniWorld-Design: From Pixel Generation to Layer-Native Design Abstract We introduce UniWorld-Design, a framework that redefines image generation from flat pixel synthesis to structured visual composition, with semantic RGBA layers as the atomic units of generation, understanding, and editing. Our key insight is that pixels define how an… 17 arXiv — NLP / Computation & Language research 10d ago Mind the Gap: Zero-Query Jailbreaks via Filter-Generator Discrepancy in Text-to-Image Systems arXiv:2608.00973v1 Announce Type: new Abstract: Text-to-image (T2I) systems typically have prompt-level safety filters before the generator to block unsafe requests, yet such systems remain vulnerable to malicious jailbreak prompts. Transfer-based attacks construct adversarial… 6 arXiv — Machine Learning research 11d ago WaiT for the Signal: Simple Frequency-Aware Flow-Matching arXiv:2607.28760v1 Announce Type: cross Abstract: As image generation models scale to ever higher resolutions, global coherence, local detail, and texture fidelity become critical axes for generation quality. However, standard flow matching treats all spatial frequencies… 37 arXiv — Machine Learning research 14d ago Flux-OPD: On-Policy Distillation with Evolving Contexts arXiv:2607.28022v1 Announce Type: new Abstract: Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision… 13 Hugging Face Daily Papers research 14d ago MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing Abstract Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities,… 12 Hugging Face Daily Papers research 14d ago Flux-OPD: On-Policy Distillation with Evolving Contexts Abstract Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student,… 24 Hugging Face Daily Papers research 16d ago TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward Abstract Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that modify the sampling trajectory to produce images more faithful to complex compositional prompts. We present TILT, a training-free framework… 27 arXiv — Machine Learning research 17d ago Learning Sampling Parameters for Diffusion Models arXiv:2607.23488v1 Announce Type: new Abstract: Text-to-image diffusion models expose many inference-time sampling parameters, including prompts, negative prompts, classifier-free guidance scales, and noise schedules. These parameters are typically manually chosen once and then… 29 TechCrunch — AI news-outlet 20d ago Midjourney acquired the astrology app Co-Star The AI lab Midjourney continues to expand its purview beyond image and video generation. 11 r/LocalLLaMA community 21d ago FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence Introducing FLUX 3. One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style. Blog Post : https://bfl.ai/blog/flux-3   submitted by   /u/pmttyji [link]   [comments] 23 Hacker News — AI on Front Page community 21d ago Flux 3 X Mimic: The Next Generation of Video-Action Models Article URL: https://bfl.ai/blog/flux-3-mimic Comments URL: https://news.ycombinator.com/item?id=49033127 Points: 240 # Comments: 32 24 Hacker News — AI on Front Page community 21d ago Flux 3 Article URL: https://bfl.ai/blog/flux-3 Comments URL: https://news.ycombinator.com/item?id=49031796 Points: 313 # Comments: 77 36 Latent.Space news-outlet 21d ago [AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model A HUGE win for BFL! 22 arXiv — Machine Learning research 22d ago Neural Operator Surrogates for Two-Dimensional Neutron Flux Estimation arXiv:2607.19388v1 Announce Type: new Abstract: This work extends our one-dimensional single-sweep neural-operator studies to two dimensions. We consider one-group transport with isotropic scattering. As in the one-dimensional work, we use Fourier neural operators (FNOs) to… 38 r/LocalLLaMA community 22d ago Mage-Flow - An Efficient Native-Resolution Foundation Model for Image Generation and Editing - Microsoft Models: (Check Model cards for so much sample demo images) https://huggingface.co/microsoft/Mage-Flow https://huggingface.co/microsoft/Mage-Flow-Turbo https://huggingface.co/microsoft/Mage-Flow-Edit Mage-Flow is a compact 4B-scale generative stack for efficient text-to-image… 30 r/MachineLearning community 22d ago Anyone heading to Jeju for KDD? Let's meet up! 🙋[D] Hey all! Is anyone else going to be at KDD in Jeju? Would love to connect with fellow attendees. I work on interpretability, fairness, and editing of text-to-image models, so I'd especially love to meet people working in these areas. But honestly, we can chat about anything:… 4 Hugging Face Daily Papers research 23d ago Appearance Pointers -- Multimodal Region Control of Diffusion Transformers Abstract Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers… 20 Hugging Face Daily Papers research 23d ago Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers Abstract Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computation during denoising remains poorly understood. We introduce a causal interpretability framework for modern large-scale DiTs that combines attention… 8 Hugging Face Daily Papers research 23d ago Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Abstract Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two… 23 arXiv — Machine Learning research 28d ago Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation arXiv:2607.14962v1 Announce Type: new Abstract: Text-to-image (T2I) models can synthesize realistic, prompt-aligned images, yet samples generated for the same prompt often cover only a small subset of visually distinct modes. This limits the diversity of images, and for… 4 arXiv — Machine Learning research 28d ago Inference-Time Concept Suppression and Video-Centric Evaluation for Text-to-Video Models arXiv:2607.14194v1 Announce Type: cross Abstract: Text-to-video (T2V) generators can synthesize realistic and temporally coherent videos, but controllably removing a target concept from a generator remains difficult. Unlike text-to-image concept erasure, T2V unlearning must… 5 Hugging Face Daily Papers research 29d ago Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation Abstract We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference,… 10 arXiv — Machine Learning research 1mo ago SymbOmni: Evolving Agentic Omni Models via Symbolic Concept Learning arXiv:2607.12042v1 Announce Type: cross Abstract: Visual generation is increasingly ubiquitous in diverse domains, from text-to-image/video synthesis to multimodal interactive creation. Yet prevailing monolithic models remain fundamentally constrained by their inability to learn… 32 Hugging Face Daily Papers research 1mo ago Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Abstract In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of asking the MLLM to judge a generated image or answer decomposed verification… 27 Hugging Face Daily Papers research 1mo ago Latent-Identity Tuning in Text-to-Image Personalization Models Abstract Generating and editing a person's face demands high precision, as even minor modifications can significantly alter a subject's perceived identity. Current personalization and editing methods built on general-purpose text-to-image models, however, often lack the… 36 Hugging Face Daily Papers research 1mo ago From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models Abstract Pretrained diffusion transformers can be adapted for dense prediction tasks by mapping tokens to task-native outputs instead of generating RGB images, achieving state-of-the-art results with minimal additional parameters. Generated by Qwen/Qwen2.5-Coder-32B-Instruct… 6 r/LocalLLaMA community 1mo ago I built Flaxeo Image a local desktop ui for stable diffusion cpp Built around a recent sd.cpp release, aims to expose most of what the backend can do (generate, edit, video paths, models, hardware options), Windows + Linux builds GitHub: https://github.com/fabricio3g/FlaxeoUI   submitted by   /u/fabricio3g [link]   [comments] 6 Vercel — AI dev-tools 1mo ago Seedream 5.0 Pro is now available on AI Gateway Seedream 5.0 Pro is now available on AI Gateway . Seedream 5.0 Pro is an image generation and editing model. It generates images from text, rendering text without spelling errors and following typographic rules, and produces dense infographics with charts, timelines, and layouts… 20 arXiv — Machine Learning research 1mo ago AutoAnchor: Stable Diffusion Unlearning Using Cross-Attention as a Manifold Surrogate arXiv:2607.08337v1 Announce Type: new Abstract: Diffusion unlearning is essential for mitigating the generation of harmful or copyrighted content in text-to-image models. Current diffusion unlearning techniques determine the model update direction by either using alternatives of… 21 arXiv — NLP / Computation & Language research 1mo ago Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing arXiv:2607.08497v1 Announce Type: cross Abstract: Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, they repeatedly feed all historical visual and textual inputs into a shared… 23 Hugging Face Daily Papers research 1mo ago Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models Abstract Flash-BoN improves text-to-image generation efficiency by using inexpensive draft candidates generated through timestep truncation, layer skipping, and activation proxies, followed by multi-stage verification that outperforms existing methods under fixed wall-clock… 9 arXiv — Machine Learning research 1mo ago An Hybrid Quantum-Classical Diffusion Model for Image Generation arXiv:2607.07072v1 Announce Type: new Abstract: Quantum diffusion models provide a physics-consistent route to generative learning by formulating noising and denoising directly on quantum states. However, applying such models to classical high-dimensional data is constrained by… 9 arXiv — NLP / Computation & Language research 1mo ago Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies arXiv:2607.06963v1 Announce Type: cross Abstract: Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic, Google, Meta, Microsoft, Stability AI, respectively, are revolutionizing… 5 arXiv — Machine Learning research 1mo ago TILDE: TILt-based Distributional Erasure for Concept Unlearning arXiv:2607.06432v1 Announce Type: new Abstract: Concept unlearning in text-to-image diffusion models is critical for safe and practical deployment: with rising privacy concerns, copyright disputes, trademark constraints, and safety regulations, deployed systems must be able to… 22 Page 1 of 4 · 160 articles Older →