StyleForge jointly selects visually coherent furniture for fixed 3D room layouts using a dynamic hypergraph style field and counterfactual energy-based test-time training.</p>\n","updatedAt":"2026-08-04T02:51:10.310Z","author":{"_id":"62ebd791fee90fca4742ead8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62ebd791fee90fca4742ead8/8-iJYS9Wk41l6JTGtJrNf.jpeg","fullname":"levon dang","name":"levondang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.858481228351593},"editors":["levondang"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/62ebd791fee90fca4742ead8/8-iJYS9Wk41l6JTGtJrNf.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.01954","authors":[{"_id":"6a715377ec5082b9f872cdda","name":"Lingwei Dang","hidden":false},{"_id":"6a715377ec5082b9f872cddb","name":"Shishuo Shang","hidden":false},{"_id":"6a715377ec5082b9f872cddc","name":"Pan Liu","hidden":false},{"_id":"6a715377ec5082b9f872cddd","name":"Jiajia Cheng","hidden":false},{"_id":"6a715377ec5082b9f872cdde","name":"Ziyan Qiu","hidden":false},{"_id":"6a715377ec5082b9f872cddf","name":"Zhenhao Zhang","hidden":false},{"_id":"6a715377ec5082b9f872cde0","name":"Yufei Zhu","hidden":false},{"_id":"6a715377ec5082b9f872cde1","name":"Shenghui Huang","hidden":false},{"_id":"6a715377ec5082b9f872cde2","name":"Qingxin Xiao","hidden":false},{"_id":"6a715377ec5082b9f872cde3","name":"Yun Hao","hidden":false},{"_id":"6a715377ec5082b9f872cde4","name":"Juntong Li","hidden":false},{"_id":"6a715377ec5082b9f872cde5","name":"Qingyao Wu","hidden":false}],"publishedAt":"2026-08-03T00:00:00.000Z","submittedOnDailyAt":"2026-08-04T00:00:00.000Z","title":"StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field","submittedOnDailyBy":{"_id":"62ebd791fee90fca4742ead8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62ebd791fee90fca4742ead8/8-iJYS9Wk41l6JTGtJrNf.jpeg","isPro":false,"fullname":"levon dang","user":"levondang","type":"user","name":"levondang"},"summary":"Fixed-layout indoor furniture styling requires selecting assets that form a coherent room without changing the prescribed furniture categories, positions, orientations, or scales. Existing approaches typically retrieve each asset independently or rely on static local relations, making them prone to shape, material, and color conflicts after scene composition. We introduce StyleForge, a scene-level structured selection framework built on a dynamic hypergraph style field. A frozen multimodal large language model extracts structured style priors from an open-ended style request and the fixed layout, while StyleForge maintains a learnable candidate distribution for each furniture slot. Conditioned on the target style, the dynamic hypergraph style field adaptively activates and weights layout-induced hyperedges to capture higher-order dependencies among furniture. Counterfactual style preference learning then treats each candidate as a local substitution in the current style field and evaluates its contextual compatibility using Mahalanobis energies. Training alternates between optimizing the style field and the candidate logits. At inference, the model remains frozen and test-time training updates only room-specific candidate logits, progressively correcting cross-slot style conflicts as the global scene context evolves. Experiments on 3D-FRONT demonstrate state-of-the-art furniture retrieval and scene-level style coherence, producing more coherent fixed-layout furniture arrangements than object- and scene-level retrieval baselines.","upvotes":6,"discussionId":"6a715378ec5082b9f872cde6","organization":{"_id":"65d9f8e26b8ab39009e13ba2","name":"cn-scut","fullname":"South China University of Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/636608e78b6bd0d9f663f4cd/8oSX1nc7RpNDTBApOTzWL.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"62ebd791fee90fca4742ead8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62ebd791fee90fca4742ead8/8-iJYS9Wk41l6JTGtJrNf.jpeg","isPro":false,"fullname":"levon dang","user":"levondang","type":"user"},{"_id":"67a474fb3dd995045efb9f3a","avatarUrl":"/avatars/e9fbac5491fdef7b64647ef206f0ce2f.svg","isPro":false,"fullname":"y","user":"ssshang6178096","type":"user"},{"_id":"6571ce9bf469a5d950cbe83e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6571ce9bf469a5d950cbe83e/e4HhzII-e2KqVnl9gvn3x.jpeg","isPro":false,"fullname":"PPL","user":"RichL","type":"user"},{"_id":"67283f5e198c50df99c407bc","avatarUrl":"/avatars/eda5bcaa421d3c05fcae1df83176d66c.svg","isPro":false,"fullname":"jtl","user":"qlzjftm","type":"user"},{"_id":"68eb5e532c5fcae039badaf3","avatarUrl":"/avatars/d9eac0a8a87e4cf05fe3e848def193da.svg","isPro":false,"fullname":"丘梓彦","user":"qzy32","type":"user"},{"_id":"69d369318f9db8cdf1626a66","avatarUrl":"/avatars/598824d5708fb5e1e9e53cc2647b7b19.svg","isPro":false,"fullname":"cjj","user":"cjj135246","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"65d9f8e26b8ab39009e13ba2","name":"cn-scut","fullname":"South China University of Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/636608e78b6bd0d9f663f4cd/8oSX1nc7RpNDTBApOTzWL.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.01954.md","query":{}}">
StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field
Abstract
Fixed-layout indoor furniture styling requires selecting assets that form a coherent room without changing the prescribed furniture categories, positions, orientations, or scales. Existing approaches typically retrieve each asset independently or rely on static local relations, making them prone to shape, material, and color conflicts after scene composition. We introduce StyleForge, a scene-level structured selection framework built on a dynamic hypergraph style field. A frozen multimodal large language model extracts structured style priors from an open-ended style request and the fixed layout, while StyleForge maintains a learnable candidate distribution for each furniture slot. Conditioned on the target style, the dynamic hypergraph style field adaptively activates and weights layout-induced hyperedges to capture higher-order dependencies among furniture. Counterfactual style preference learning then treats each candidate as a local substitution in the current style field and evaluates its contextual compatibility using Mahalanobis energies. Training alternates between optimizing the style field and the candidate logits. At inference, the model remains frozen and test-time training updates only room-specific candidate logits, progressively correcting cross-slot style conflicts as the global scene context evolves. Experiments on 3D-FRONT demonstrate state-of-the-art furniture retrieval and scene-level style coherence, producing more coherent fixed-layout furniture arrangements than object- and scene-level retrieval baselines.
Community
StyleForge jointly selects visually coherent furniture for fixed 3D room layouts using a dynamic hypergraph style field and counterfactual energy-based test-time training.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.01954 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.01954 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.01954 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.