Hugging Face Daily Papers · · 5 min read

MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities, and interpenetrating bodies. Existing evaluations largely overlook these anatomical and geometric issues, and VLM-as-a-judge checklists often saturate on Interaction while the errors remain obvious to humans. We introduce MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities (C0-C3). We also propose MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction. Anatomy asks whether every human-like mass is explained by a complete set of reconstructed bodies, and Interaction asks whether the penetration and surface distance between those bodies match the contact the instruction asked for. Across ten editors, mesh Anatomy tops out at 0.65 and mesh Interaction at 0.72 on two different models, so no single editor is strong on both, while VLM checklists rate the same images above 0.95. A five-rater study confirms that both axes track human judgement more closely than a zero-shot VLM judge, and the rankings hold under ablation of every weight and threshold.</p>\n","updatedAt":"2026-07-31T03:42:07.205Z","author":{"_id":"646dbba74ad7f907279dd486","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646dbba74ad7f907279dd486/XGHJMFEIpWeDlh0cn9Slu.png","fullname":"Mingxuan Du","name":"Ayanami0730","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":13,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8927974700927734},"editors":["Ayanami0730"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/646dbba74ad7f907279dd486/XGHJMFEIpWeDlh0cn9Slu.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.27616","authors":[{"_id":"6a6c10e4202e2d9e3ffdb742","name":"Jiajia Lin","hidden":false},{"_id":"6a6c10e4202e2d9e3ffdb743","user":{"_id":"646dbba74ad7f907279dd486","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646dbba74ad7f907279dd486/XGHJMFEIpWeDlh0cn9Slu.png","isPro":false,"fullname":"Mingxuan Du","user":"Ayanami0730","type":"user","name":"Ayanami0730"},"name":"Mingxuan Du","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.573Z","hidden":false},{"_id":"6a6c10e4202e2d9e3ffdb744","name":"Tuowen Zhou","hidden":false},{"_id":"6a6c10e4202e2d9e3ffdb745","name":"Benfeng Xu","hidden":false},{"_id":"6a6c10e4202e2d9e3ffdb746","name":"Hongtao Xie","hidden":false}],"publishedAt":"2026-07-30T00:00:00.000Z","submittedOnDailyAt":"2026-07-31T00:00:00.000Z","title":"MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing","submittedOnDailyBy":{"_id":"646dbba74ad7f907279dd486","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646dbba74ad7f907279dd486/XGHJMFEIpWeDlh0cn9Slu.png","isPro":false,"fullname":"Mingxuan Du","user":"Ayanami0730","type":"user","name":"Ayanami0730"},"summary":"Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities, and interpenetrating bodies. Existing evaluations largely overlook these anatomical and geometric issues, and VLM-as-a-judge checklists often saturate on Interaction while the errors remain obvious to humans. We introduce MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities (C0-C3). We also propose MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction. Anatomy asks whether every human-like mass is explained by a complete set of reconstructed bodies, and Interaction asks whether the penetration and surface distance between those bodies match the contact the instruction asked for. Across ten editors, mesh Anatomy tops out at 0.65 and mesh Interaction at 0.72 on two different models, so no single editor is strong on both, while VLM checklists rate the same images above 0.95. A five-rater study confirms that both axes track human judgement more closely than a zero-shot VLM judge, and the rankings hold under ablation of every weight and threshold.","upvotes":34,"discussionId":"6a6c10e4202e2d9e3ffdb747","githubRepo":"https://github.com/AnnLin0628/mpie-bench","githubRepoAddedBy":"user","githubStars":1,"organization":{"_id":"6912993f45f02a20f4c50b8a","name":"muset-ai","fullname":"muset.ai","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/646dbba74ad7f907279dd486/C5j0HEiqEmmxshM61WRON.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a2a342b086133dad7d61c92","avatarUrl":"/avatars/db700b4fc2f762f072e4a22b2f48b757.svg","isPro":false,"fullname":"annlin","user":"annl092723","type":"user"},{"_id":"646dbba74ad7f907279dd486","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646dbba74ad7f907279dd486/XGHJMFEIpWeDlh0cn9Slu.png","isPro":false,"fullname":"Mingxuan Du","user":"Ayanami0730","type":"user"},{"_id":"68959bf9950f311270f0e010","avatarUrl":"/avatars/96f5ac05139954579a97a80e0f29b528.svg","isPro":false,"fullname":"Ruizhe Li","user":"imlrz","type":"user"},{"_id":"68cbf13fd9c1b531d9564757","avatarUrl":"/avatars/e1e26dc251102c5c18ce5fa172e6a8fa.svg","isPro":false,"fullname":"Zhou","user":"Tuowen","type":"user"},{"_id":"68c8ccff2de8edcf8da2fe51","avatarUrl":"/avatars/e55e509d11049691fefe3750a4bed62a.svg","isPro":false,"fullname":"wangpengyu","user":"Bstwpy","type":"user"},{"_id":"62d56f063bf5e059f7cac515","avatarUrl":"/avatars/6bd5cfdc21506ace6176d00a2973d8e5.svg","isPro":false,"fullname":"BenfengXu","user":"SpiketheCowboy","type":"user"},{"_id":"6a6824d5a2bbdde86f7cad76","avatarUrl":"/avatars/6ff679759fc9be5956f1b12799a3b874.svg","isPro":false,"fullname":"Yuhang Zhu","user":"Yuki-Kitayama","type":"user"},{"_id":"6440d49b7663594a126716f2","avatarUrl":"/avatars/bb04af9d1ae9c5bc058ebfbf08f4ebc8.svg","isPro":false,"fullname":"shaohanwang","user":"WShao","type":"user"},{"_id":"62c926c71eba56b9f6212922","avatarUrl":"/avatars/e78d2bceae9b9f2e5295408f7094c4c8.svg","isPro":false,"fullname":"Hal","user":"farawayxxx","type":"user"},{"_id":"6a0fc58cc8676ad292e8a15d","avatarUrl":"/avatars/2649c5c1d64cd72b77265747b89e0cea.svg","isPro":false,"fullname":"Ruizhe Li","user":"imlrz01","type":"user"},{"_id":"64ffd28c4b7e0cd1b7c05fd1","avatarUrl":"/avatars/684ad1ed6d251f519488f1b93bf135e7.svg","isPro":false,"fullname":"Gu","user":"Infinitegu","type":"user"},{"_id":"632db61b78bf224d04280de7","avatarUrl":"/avatars/fe7bbe3045186f71b4632668dabdab4c.svg","isPro":false,"fullname":"wenc","user":"wenc-k","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6912993f45f02a20f4c50b8a","name":"muset-ai","fullname":"muset.ai","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/646dbba74ad7f907279dd486/C5j0HEiqEmmxshM61WRON.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.27616.md","query":{}}">
Papers
arxiv:2607.27616

MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

Published on Jul 30
· Submitted by
Mingxuan Du
on Jul 31
Authors:
,

Abstract

Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities, and interpenetrating bodies. Existing evaluations largely overlook these anatomical and geometric issues, and VLM-as-a-judge checklists often saturate on Interaction while the errors remain obvious to humans. We introduce MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities (C0-C3). We also propose MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction. Anatomy asks whether every human-like mass is explained by a complete set of reconstructed bodies, and Interaction asks whether the penetration and surface distance between those bodies match the contact the instruction asked for. Across ten editors, mesh Anatomy tops out at 0.65 and mesh Interaction at 0.72 on two different models, so no single editor is strong on both, while VLM checklists rate the same images above 0.95. A five-rater study confirms that both axes track human judgement more closely than a zero-shot VLM judge, and the rankings hold under ablation of every weight and threshold.

Community

Paper author Paper submitter about 6 hours ago

Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities, and interpenetrating bodies. Existing evaluations largely overlook these anatomical and geometric issues, and VLM-as-a-judge checklists often saturate on Interaction while the errors remain obvious to humans. We introduce MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities (C0-C3). We also propose MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction. Anatomy asks whether every human-like mass is explained by a complete set of reconstructed bodies, and Interaction asks whether the penetration and surface distance between those bodies match the contact the instruction asked for. Across ten editors, mesh Anatomy tops out at 0.65 and mesh Interaction at 0.72 on two different models, so no single editor is strong on both, while VLM checklists rate the same images above 0.95. A five-rater study confirms that both axes track human judgement more closely than a zero-shot VLM judge, and the rankings hold under ablation of every weight and threshold.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.27616
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.27616 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.27616 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.27616 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers