News / #image-gen Tag Image Gen 191 articles archived under #image-gen · RSS Sign in to follow r/LocalLLaMA community 1d ago I added Qwen-Image 2.1 + LoRA support to TensorSharp (GGUF, local inference) I maintain TensorSharp , an open-source inference engine. It can now run Qwen-Image 2.1 locally for text-to-image generation and image editing, with support for its LoRA adapters. I’ve added configs for regular style and editing LoRAs, plus accelerated adapters with their own… 14 r/LocalLLaMA community 4d ago BFL releases FLUX 3 Action: a 7B robot model read more: https://bfl.ai/models/flux-3-action   submitted by   /u/paf1138 [link]   [comments] 36 r/LocalLLaMA community 5d ago What underrated AI tools have actually made you more productive in 2026? I asked this back in 2025, but the AI landscape has changed a lot since then. Not looking for the usual ChatGPT, Claude, Gemini, Midjourney, etc. I'm curious about the lesser-known tools that you actually kept using . Could be for research, coding, design, video, writing,… 23 r/LocalLLaMA community 6d ago [MASSIVE RELEASE] Supra2-IMG - a tiny 100M text-to-image model - SOTA quality and open release! Hey everyone! It has been quite a while since the last SupraLabs model - but today we've something special for y'all: Supra2-IMG It's a 100M parameter DiT text-to-image model trained entirely from scratch in under 10 hours on a single H100 on Runpod. It can generate… 32 arXiv — Machine Learning research 7d ago Assessment of Machine Learning-Based Critical Heat Flux Models in the CTF Subchannel Code for Square Rod Bundle Prediction arXiv:2609.21995v1 Announce Type: new Abstract: The prediction of critical heat flux (CHF), a key safety-related quantity in nuclear thermal hydraulics, remains an important challenge due to its direct relationship with fuel performance and reactor safety. Recent studies have… 9 r/LocalLLaMA community 7d ago Qwen-Image-2.1 released! Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both generation and editing, delivering top-tier quality in a lightweight package. Highlights: - Compact & exceptionally fast: A… 5 arXiv — Machine Learning research 13d ago Attention Is All You Need (to Avoid Spurious Oscillations) arXiv:2609.13531v1 Announce Type: new Abstract: Can attention move a shock across several cells in one update without breaking it? We develop a conservative, fixed grid finite-volume scheme in which a CFL-conditioned attention flux selects upstream information according to the… 18 arXiv — Machine Learning research 14d ago Certifying Concept Unlearning in Text-to-Image Diffusion Models arXiv:2609.12163v1 Announce Type: new Abstract: Existing evaluations of concept unlearning in text-to-image (T2I) diffusion models primarily rely on attack success rates obtained through automated adversarial prompt search. However, these metrics provide only empirical evidence… 16 r/MachineLearning community 16d ago Training a 210M text-to-image DiT from scratch on one GPU: what I measured [P] I trained a 210M-parameter text-to-image diffusion transformer from scratch (3.5 days, one RTX PRO 6000, 4.2M images at 256²) mainly to understand the recipe end to end. Three measurements came out of it that I have not seen stated plainly elsewhere, so I'm posting those rather… 29 Simon Willison community 19d ago Introducing ChatGPT Images 2.5 Introducing ChatGPT Images 2.5 OpenAI's image generation models are apparently used "more than 3 billion images across ChatGPT Images and the GPT‑Image models in the API". This latest release improves their instruction-following ability across multiple turns, responds faster,… 26 arXiv — Machine Learning research 25d ago CAT-Flow: Curvature-Adaptive sTeps for Flow Matching arXiv:2609.01746v1 Announce Type: new Abstract: Flow Matching has emerged as a leading framework for generative modeling, powering state-of-the-art systems such as FLUX and Stable Diffusion 3.5. However, the iterative nature of its ODE-based sampling process creates a… 33 r/MachineLearning community 25d ago Detailed explanation of how to create a text-to-image model from scratch. [R] Jasper Research just released a cookbook on how to build a text-to-image model from scratch. It shares the full reasoning and intermediate results, making it ideal if you want to deep-dive into text-to-image models, or if you are curious about how frontier labs build them. The… 21 arXiv — NLP / Computation & Language research 26d ago Visual Framing for News Stance Detection via Image Generation arXiv:2609.00685v1 Announce Type: new Abstract: Article-level news stance detection aims to identify the perspective of news articles toward social issues. Despite advances in stance detection and its importance for trustworthy media environments, news articles pose distinct… 33 Hugging Face Daily Papers research 26d ago ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models Abstract Text-to-image models exhibit persistent role-linked visual stereotypes across varied contexts, with cross-role attribute concentration increasing rather than diminishing under unrelated prompts. Generated by thinkingmachines/Inkling-Small Text-to-image models learn… 13 arXiv — Machine Learning research 28d ago There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation arXiv:2608.27885v1 Announce Type: new Abstract: Multimodality translation (e.g., text-to-image) is a core generative AI task. However, existing approaches (1) follow generative paths that do not directly represent the source modality, limiting the flexibility of some sampling… 10 r/MachineLearning community 1mo ago I implemented a very tiny image generation model (latent flow transformer) on a RP2350 microcontroller - it can generate 128x128 images of faces [P] Its a 2.4-4 million parameter model, quantized to int8, that can be fully executed on the microcontroller in ~20s with the longest generation. The generated image will then be displayed on a monitor or transferred via usb. Its a latent flow transformer with 12 layers using… 26 arXiv — Machine Learning research 1mo ago Joint Initialization of Flux Networks and Effective Multiplication Factor for Physics-Informed Neural Networks Solving Neutron Diffusion Problems arXiv:2608.25443v1 Announce Type: new Abstract: Efficient determination of the effective multiplication factor (keff) is an important computational task in reactor core neutronics analysis. Physics-informed neural networks (PINNs) incorporate neutron diffusion equations and… 10 TechCrunch — AI news-outlet 1mo ago Stability AI, maker of image generator Stable Diffusion, raises $76 million in fresh funding The company's new fundraising total now stands at $232 million. 29 Hugging Face Daily Papers research 1mo ago UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Abstract A reparameterized pretrained vision transformer unifies semantic understanding, high-fidelity reconstruction, and image generation within a single visual space without requiring a separate VAE. Generated by thinkingmachines/Inkling-Small Semantic vision encoders have… 26 r/MachineLearning community 1mo ago OpenAI advertises "unlimited" image generation in Pro. It is not unlimited. [D] Signed up for ChatGPT Pro for an image generation project. The pricing page says Pro comes with "unlimited and faster image creation" (chatgpt.com/pricing). A few hundred images in, I got rate limited. Reached out to support, and the response was: "Pro access remains subject to… 7 Hugging Face Daily Papers research 1mo ago WithEveryone: Unified Planning and Identity Grounding for Group Image Generation Abstract WithEveryone enables reliable identity-preserving group image generation for up to ten people by grounding identities to layout plans and using region-based identity losses. Generated by thinkingmachines/Inkling-Small Identity-preserving image generation becomes… 17 Hugging Face Daily Papers research 1mo ago DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization Abstract DiSCO is a black-box, zero-shot prompt-level defense that uses distribution-guided suffix expansion and contrastive scoring to reduce harmful image generation without altering the model. Generated by thinkingmachines/Inkling-Small As text-to-image generative models… 27 Hugging Face Daily Papers research 1mo ago Abra: Scaling Diffusion Image Training Abstract Scaling laws for text-to-image diffusion models reveal predictable compute-optimal training requiring far more data per parameter than language models, with robust overtraining behavior and universal curve shapes. Generated by thinkingmachines/Inkling-Small… 22 arXiv — Machine Learning research 1mo ago Abra: Scaling Diffusion Image Training arXiv:2608.17286v1 Announce Type: new Abstract: Compute-optimal scaling laws guide the training of frontier language models yet remain largely unexplored for visual generation. We present a systematic scaling law study for text-to-image diffusion models using Abra, a controlled… 7 Hugging Face Daily Papers research 1mo ago From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation Abstract A capability-driven data infrastructure with curriculum scheduling and specialized data engines trains large multimodal diffusion models on curated heterogeneous supervision for diverse generative tasks. Generated by thinkingmachines/Inkling-Small Large-scale image… 13 r/MachineLearning community 1mo ago Trained an diffusion model that runs on 264KB of RAM [P] I recently bought a Shrike lite which has got 264KB of SRAM. I decided to train an image generation model that generates 32*32 pixel images. The microcontroller also has an FPGA onboard which I used to create two parallel INT8 MAC engines with 16 bit accumulation to speed up… 36 Hugging Face Daily Papers research 1mo ago TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation Abstract This work proposes a compositional operator framework and TRACE-Bench to diagnose multi-reference image generation capabilities across atomic operations. Generated by thinkingmachines/Inkling-Small Despite recent advances in unified multimodal models for multi-reference… 22 Hugging Face Daily Papers research 1mo ago An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models Abstract Researchers propose a latent-to-pixel training strategy that accelerates convergence and improves inference speed for large-scale pixel-space diffusion models. Generated by thinkingmachines/Inkling-Small This paper investigates an increasingly important topic in… 16 Hugging Face Daily Papers research 1mo ago GenRouter: Unified Workflow Routing for Agentic Image Generation Abstract GenRouter is a unified routing framework that adaptively directs prompts to optimal agentic image-generation workflows, cutting costs and latency while improving visual alignment and enabling continuous self-evolution. Generated by thinkingmachines/Inkling-Small The… 34 arXiv — Machine Learning research 1mo ago Adversarial Learning of Classifier-Free Guidance Schedules arXiv:2608.14038v1 Announce Type: new Abstract: Modern text-to-image diffusion models rely on classifier-free guidance (CFG) to achieve high image fidelity and text alignment. However, CFG typically applies a static, global scale across all timesteps, samples, and conditions --… 26 r/LocalLLaMA community 1mo ago [Megathread] Qwen 3.8 27B Release Day Megathread to help with the influx of duplicate / similar posts around the release of the Qwen 3.8 27B release. Quants Fine-Tunes & Abliterations Chat Templates Inference Server Support & Configuration Experiences, Benchmarks & Model Comparisons Official:… 10 r/MachineLearning community 1mo ago Reproducible canvas-aligned low-level patterns in somerandomllm-generated images and their possible relation to iterative editing artifacts [D] I may have stumbled onto something interesting while trying to figure out a recurring artifact in ChatGPT image generation and editing (maybe applicable to other models as well?). It started with a very practical problem: After several rounds of generative editing on portraits,… 15 arXiv — NLP / Computation & Language research 1mo ago On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation arXiv:2608.11002v1 Announce Type: new Abstract: Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual performance gaps and language-specific effects… 37 Hugging Face Daily Papers research 1mo ago Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation Abstract Atelier improves artist-grounded image generation by translating vague artistic intent into explicit control states that separate scene content from style, reducing reliance on stereotypical shortcuts. Generated by thinkingmachines/Inkling-Small Artist-grounded image… 20 arXiv — NLP / Computation & Language research 1mo ago Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs arXiv:2608.06967v1 Announce Type: new Abstract: Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide visual-plan failures: an answer may appear creative while repeating familiar… 36 Simon Willison community 1mo ago Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) On Wednesday I wrote about One-shotting a Raccoon Heist game using Claude Fable 5 , where I had Claude Fable 5 build a full working game from a premise I generated with GPT-3 and DALL-E four years ago . I decided to… 33 Simon Willison community 1mo ago Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) On Wednesday I wrote about One-shotting a Raccoon Heist game using Claude Fable 5 , where I had Claude Fable 5 build a full working game from a premise I generated with GPT-3 and DALL-E four years ago . I decided to… 20 arXiv — NLP / Computation & Language research 1mo ago GenGA: Editable and Data-Grounded Graphical Abstract Generation for Academic Papers arXiv:2608.05478v1 Announce Type: cross Abstract: Graphical Abstracts (GAs) visually summarize the key findings of academic papers, playing a crucial role in facilitating the understanding of research content. Recently, advancements in vision-language models and image generation… 38 Hugging Face Daily Papers research 1mo ago ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Abstract Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, and the integration of external world knowledge. Existing efforts introduce agent… 36 arXiv — NLP / Computation & Language research 1mo ago Simile Understanding in Text-to-Image Models: An Evaluation Framework arXiv:2608.04750v1 Announce Type: cross Abstract: Similes provide a compact and expressive way to describe visual characteristics in text prompts. Recent text-to-image models (t2i models) can produce visually compelling outputs from simile prompts, yet even frontier models… 12 Hugging Face Daily Papers research 1mo ago Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models Abstract Leading open text-to-image models often carry complementary strengths: one may lead on preference-aligned aesthetics while another follows compositional instructions more faithfully. However, differences in their autoencoders and noise schedules make it difficult to… 23 Simon Willison community 1mo ago One-shotting a Raccoon Heist game using Claude Fable 5 Back in 2024 I tweeted screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in Claude Code for web ) could build the entire game from the content… 29 Simon Willison community 1mo ago One-shotting a Raccoon Heist game using Claude Fable 5 Back in 2022 I tweeted screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in Claude Code for web ) could build the entire game from the content… 30 Hugging Face Daily Papers research 1mo ago PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs Abstract Scientific poster construction compresses a long multimodal paper into a readable, editable canvas. Existing systems hide request-level failures by scoring only completed outputs; direct image generation is not element-editable, while coding-agent workflows are costly.… 35 arXiv — Machine Learning research 1mo ago A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics arXiv:2608.02965v1 Announce Type: new Abstract: Magnetic components in high-frequency, high-power-density converters are increasingly driven by non-sinusoidal flux-density waveforms with fast transitions, minor-loop operation, dc bias, and temperature variation. Under these… 26 Hugging Face Daily Papers research 1mo ago Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Abstract Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and… 8 Hugging Face Daily Papers research 1mo ago CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Abstract Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct properties: (1) how much visual information… 22 Hugging Face Daily Papers research 1mo ago UniWorld-Design: From Pixel Generation to Layer-Native Design Abstract We introduce UniWorld-Design, a framework that redefines image generation from flat pixel synthesis to structured visual composition, with semantic RGBA layers as the atomic units of generation, understanding, and editing. Our key insight is that pixels define how an… 17 arXiv — NLP / Computation & Language research 1mo ago Mind the Gap: Zero-Query Jailbreaks via Filter-Generator Discrepancy in Text-to-Image Systems arXiv:2608.00973v1 Announce Type: new Abstract: Text-to-image (T2I) systems typically have prompt-level safety filters before the generator to block unsafe requests, yet such systems remain vulnerable to malicious jailbreak prompts. Transfer-based attacks construct adversarial… 6 arXiv — Machine Learning research 1mo ago WaiT for the Signal: Simple Frequency-Aware Flow-Matching arXiv:2607.28760v1 Announce Type: cross Abstract: As image generation models scale to ever higher resolutions, global coherence, local detail, and texture fidelity become critical axes for generation quality. However, standard flow matching treats all spatial frequencies… 37 Page 1 of 4 · 191 articles Older →