Hugging Face Daily Papers · · 3 min read

InstanceControl: Controllable Complex Image Generation without Instance Labeling

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Accepted by ECCV 2026, project page: <a href=\"https://instancecontrol.github.io/InstanceControl/\" rel=\"nofollow\">https://instancecontrol.github.io/InstanceControl/</a></p>\n","updatedAt":"2026-07-03T06:28:13.921Z","author":{"_id":"633e848d26bad5c614000fbf","avatarUrl":"/avatars/fdb0ed55973613fbd0e8c7e75db4ac59.svg","fullname":"liuxiaoyu","name":"xiaoyu1104","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7136828303337097},"editors":["xiaoyu1104"],"editorAvatarUrls":["/avatars/fdb0ed55973613fbd0e8c7e75db4ac59.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2606.31924","authors":[{"_id":"6a44bcbb41f04ae4d7ad99a0","user":{"_id":"633e848d26bad5c614000fbf","avatarUrl":"/avatars/fdb0ed55973613fbd0e8c7e75db4ac59.svg","isPro":false,"fullname":"liuxiaoyu","user":"xiaoyu1104","type":"user","name":"xiaoyu1104"},"name":"Xiaoyu Liu","status":"claimed_verified","statusLastChangedAt":"2026-07-01T08:44:07.203Z","hidden":false},{"_id":"6a44bcbb41f04ae4d7ad99a1","name":"Huan Wang","hidden":false},{"_id":"6a44bcbb41f04ae4d7ad99a2","name":"Fan Li","hidden":false},{"_id":"6a44bcbb41f04ae4d7ad99a3","name":"Zhixin Wang","hidden":false},{"_id":"6a44bcbb41f04ae4d7ad99a4","name":"Jiaqi Xu","hidden":false},{"_id":"6a44bcbb41f04ae4d7ad99a5","name":"Ming Liu","hidden":false},{"_id":"6a44bcbb41f04ae4d7ad99a6","name":"Wangmeng Zuo","hidden":false}],"publishedAt":"2026-06-30T00:00:00.000Z","submittedOnDailyAt":"2026-07-03T00:00:00.000Z","title":"InstanceControl: Controllable Complex Image Generation without Instance Labeling","submittedOnDailyBy":{"_id":"633e848d26bad5c614000fbf","avatarUrl":"/avatars/fdb0ed55973613fbd0e8c7e75db4ac59.svg","isPro":false,"fullname":"liuxiaoyu","user":"xiaoyu1104","type":"user","name":"xiaoyu1104"},"summary":"Controllable image generation methods, such as ControlNet, have demonstrated a remarkable capacity to introduce visual conditions(e.g., depth maps) to guide image generation. However, these methods often struggle with complex multi-instance scenes, frequently leading to attribute confusion among instances. While recent approaches attempt to mitigate this via manual instance labeling, such requirements are labor-intensive. In this paper, we propose InstanceControl, a novel multi-instance controllable generation method that eliminates the need for instance labeling. We identify the primary bottleneck in existing methods as the inability to accurately associate instance descriptions with their corresponding regions within visual conditions. To address this, we leverage the Vision-Language Model (VLM) to establish instance-level correspondences between text prompts and visual conditions. Specifically, the VLM automatically parses instance descriptions from the text prompts and simultaneously predicts instance masks based on the visual conditions. Furthermore, since the predicted masks may contain noise, we introduce an adaptive mask refinement strategy that dynamically refines these instance masks during the generation process. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods, achieving superior fidelity and precise instance-level control.","upvotes":2,"discussionId":"6a44bcbc41f04ae4d7ad99a7","projectPage":"https://instancecontrol.github.io/InstanceControl/","githubRepo":"https://github.com/liuxiaoyu1104/InstanceControl","githubRepoAddedBy":"user","ai_summary":"InstanceControl enables multi-instance image generation by using vision-language models to establish instance-level correspondences between text prompts and visual conditions, while employing adaptive mask refinement for improved accuracy.","ai_keywords":["ControlNet","Vision-Language Model","instance-level correspondences","instance masks","adaptive mask refinement","multi-instance scenes","visual conditions","text prompts"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":13},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"633e848d26bad5c614000fbf","avatarUrl":"/avatars/fdb0ed55973613fbd0e8c7e75db4ac59.svg","isPro":false,"fullname":"liuxiaoyu","user":"xiaoyu1104","type":"user"},{"_id":"6a147f4f695d577a5249a9c8","avatarUrl":"/avatars/1194f0e6d9c809dd7767c413d64cd889.svg","isPro":false,"fullname":"Emily Brown","user":"emily-brown2025","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2606/2606.31924.md","query":{}}">
Papers
arxiv:2606.31924

InstanceControl: Controllable Complex Image Generation without Instance Labeling

Published on Jun 30
· Submitted by
liuxiaoyu
on Jul 3
Authors:
,
,
,
,
,

Abstract

InstanceControl enables multi-instance image generation by using vision-language models to establish instance-level correspondences between text prompts and visual conditions, while employing adaptive mask refinement for improved accuracy.

Controllable image generation methods, such as ControlNet, have demonstrated a remarkable capacity to introduce visual conditions(e.g., depth maps) to guide image generation. However, these methods often struggle with complex multi-instance scenes, frequently leading to attribute confusion among instances. While recent approaches attempt to mitigate this via manual instance labeling, such requirements are labor-intensive. In this paper, we propose InstanceControl, a novel multi-instance controllable generation method that eliminates the need for instance labeling. We identify the primary bottleneck in existing methods as the inability to accurately associate instance descriptions with their corresponding regions within visual conditions. To address this, we leverage the Vision-Language Model (VLM) to establish instance-level correspondences between text prompts and visual conditions. Specifically, the VLM automatically parses instance descriptions from the text prompts and simultaneously predicts instance masks based on the visual conditions. Furthermore, since the predicted masks may contain noise, we introduce an adaptive mask refinement strategy that dynamically refines these instance masks during the generation process. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods, achieving superior fidelity and precise instance-level control.

Community

Paper author Paper submitter 2 days ago
This comment has been hidden (marked as Resolved)
Paper author Paper submitter about 4 hours ago

Accepted by ECCV 2026, project page: https://instancecontrol.github.io/InstanceControl/

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2606.31924
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2606.31924 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2606.31924 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2606.31924 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers