Hugging Face Daily Papers · · 3 min read

LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

model: <a href=\"https://huggingface.co/inclusionAI/LLaDA-UI\">https://huggingface.co/inclusionAI/LLaDA-UI</a></p>\n","updatedAt":"2026-09-15T02:47:26.009Z","author":{"_id":"60d2a2984956988b63753371","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/60d2a2984956988b63753371/apXIcWbi7jnLVH37CdMTV.jpeg","fullname":"Zhangxuan Gu","name":"zhangxgu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":10,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/662e1f9da266499277937d33/fyKuazRifqiaIO34xrhhm.jpeg","fullname":"inclusionAI","name":"inclusionAI","type":"org","isHf":false,"plan":"team","hasPrivateMembersList":true}}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.8652668595314026},"editors":["zhangxgu"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/60d2a2984956988b63753371/apXIcWbi7jnLVH37CdMTV.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.13287","authors":[{"_id":"6aa8a9ea5dd4cb9b4cc024dc","name":"Zhangxuan Gu","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024dd","name":"Haoxing Chen","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024de","name":"Qi Qin","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024df","name":"Yi Xin","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024e0","name":"Kai Gan","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024e1","name":"Lin Liu","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024e2","name":"Long Cui","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024e3","name":"Xiaomei Wang","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024e4","name":"Beitong Zhou","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024e5","name":"Yunzhu Zhang","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024e6","name":"Zhengwen Zeng","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024e7","name":"Changlong Gao","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024e8","name":"Weizhi Chen","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024e9","name":"Rongchao Zhang","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024ea","name":"Haoyuan Wu","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024eb","name":"Shuheng Shen","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024ec","name":"Changhua Meng","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024ed","name":"Weiqiang Wang","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024ee","name":"Jianguo Li","hidden":false},{"_id":"6aa8a9ea5dd4cb9b4cc024ef","name":"Zhenzhong Lan","hidden":false}],"publishedAt":"2026-09-09T00:00:00.000Z","submittedOnDailyAt":"2026-09-15T00:00:00.000Z","title":"LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents","submittedOnDailyBy":{"_id":"60d2a2984956988b63753371","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/60d2a2984956988b63753371/apXIcWbi7jnLVH37CdMTV.jpeg","isPro":false,"fullname":"Zhangxuan Gu","user":"zhangxgu","type":"user","name":"zhangxgu"},"summary":"Diffusion large language models (dLLMs) achieve high decoding efficiency through block-parallel, arbitrary-order generation, making them attractive for latency-sensitive applications. GUI agents represent a natural testbed for this paradigm, as they must repeatedly perceive screen states and emit structured, spatially grounded actions in real time. However, whether dLLMs can be extended into capable multimodal GUI agents while preserving their parallel decoding advantage remains an open question. We present LLaDA-UI, a 16.7B-parameter MoE-based, block-wise diffusion vision-language GUI agent. LLaDA-UI follows a two-stage training pipeline: general multimodal pre-training aligns a native-resolution vision encoder with the LLaDA2.0-mini-base diffusion language backbone, followed by GUI-agent supervised fine-tuning on diverse mobile, desktop, web, and grounding data. Across widely adopted grounding benchmarks and navigation benchmarks spanning multiple platforms, LLaDA-UI substantially outperforms Qwen2.5-VL-7B and surpasses Qwen3-VL-8B on four of six reported GUI benchmarks. These results establish block-wise diffusion as a practical generative paradigm for multimodal GUI agents.","upvotes":11,"discussionId":"6aa8a9ea5dd4cb9b4cc024f0","projectPage":"https://www.inclusion-ai.org/LLaDA-UI/","githubRepo":"https://github.com/inclusionAI/LLaDA-UI","githubRepoAddedBy":"user","ai_summary":"A 16.7B-parameter mixture-of-experts diffusion vision-language agent achieves strong multimodal GUI performance while preserving block-parallel decoding efficiency.","ai_keywords":["diffusion large language models","dLLMs","block-parallel generation","MoE","mixture-of-experts","vision-language model","diffusion vision-language","GUI agent","native-resolution vision encoder","supervised fine-tuning"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":9,"organization":{"_id":"67aea5c8f086ab0f70ed97c9","name":"inclusionAI","fullname":"inclusionAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/662e1f9da266499277937d33/fyKuazRifqiaIO34xrhhm.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"60d2a2984956988b63753371","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/60d2a2984956988b63753371/apXIcWbi7jnLVH37CdMTV.jpeg","isPro":false,"fullname":"Zhangxuan Gu","user":"zhangxgu","type":"user"},{"_id":"63898c562a897944ea5f07a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63898c562a897944ea5f07a2/oG1ia44ICpAk1tReIaX0G.jpeg","isPro":false,"fullname":"Haoxing chen","user":"HaoxingChen","type":"user"},{"_id":"6a880b3ae5c0e08ccd1560d1","avatarUrl":"/avatars/f90cc33a8b2775e9b1cc98d1cd9cd933.svg","isPro":false,"fullname":"ZhuoerXu","user":"ZhuoerX","type":"user"},{"_id":"66bb136002fd8eb58bc84ffb","avatarUrl":"/avatars/122cb8f59c502392768099b3c2afe043.svg","isPro":false,"fullname":"qinqi","user":"Dakerqi","type":"user"},{"_id":"632becb066f28bf34aea2f92","avatarUrl":"/avatars/7ed8111ecf780054e5f852e56b09b5a7.svg","isPro":false,"fullname":"SII-Synbol (辛毅)","user":"XiN0919","type":"user"},{"_id":"64cb238576200ec80fe988f8","avatarUrl":"/avatars/42c48710c7881c9dfbcc075fec3cb600.svg","isPro":false,"fullname":"zeus","user":"zengw","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"68d8ba59f2f999edd0e25eea","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68d8ba59f2f999edd0e25eea/9GLAoaqtxo8c5LQtBt2oX.png","isPro":false,"fullname":"Long Cui","user":"CuiLong7","type":"user"},{"_id":"646def60df618b303b419323","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646def60df618b303b419323/JLJGYen4-5M8ivsLsSk0w.jpeg","isPro":false,"fullname":"Lei Wang","user":"demolei","type":"user"},{"_id":"641e999f326d534e4f275756","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/JMMXiqVWJIExCzD_I3p22.jpeg","isPro":false,"fullname":"FlySugar","user":"SugarVapeur","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"67aea5c8f086ab0f70ed97c9","name":"inclusionAI","fullname":"inclusionAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/662e1f9da266499277937d33/fyKuazRifqiaIO34xrhhm.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.13287.md","query":{}}">
Papers
arxiv:2609.13287

LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents

Published on Sep 9
· Submitted by
Zhangxuan Gu
on Sep 15
Authors:
,

Abstract

A 16.7B-parameter mixture-of-experts diffusion vision-language agent achieves strong multimodal GUI performance while preserving block-parallel decoding efficiency.

Diffusion large language models (dLLMs) achieve high decoding efficiency through block-parallel, arbitrary-order generation, making them attractive for latency-sensitive applications. GUI agents represent a natural testbed for this paradigm, as they must repeatedly perceive screen states and emit structured, spatially grounded actions in real time. However, whether dLLMs can be extended into capable multimodal GUI agents while preserving their parallel decoding advantage remains an open question. We present LLaDA-UI, a 16.7B-parameter MoE-based, block-wise diffusion vision-language GUI agent. LLaDA-UI follows a two-stage training pipeline: general multimodal pre-training aligns a native-resolution vision encoder with the LLaDA2.0-mini-base diffusion language backbone, followed by GUI-agent supervised fine-tuning on diverse mobile, desktop, web, and grounding data. Across widely adopted grounding benchmarks and navigation benchmarks spanning multiple platforms, LLaDA-UI substantially outperforms Qwen2.5-VL-7B and surpasses Qwen3-VL-8B on four of six reported GUI benchmarks. These results establish block-wise diffusion as a practical generative paradigm for multimodal GUI agents.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.13287
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.13287 in a dataset README.md to link it from this page.

Spaces citing this paper

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers