Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed optimization, or single-image generative models that ignore crucial multi-view cues necessary for understanding 3D geometry and material interactions. To address these limitations, we introduce a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation. Adapted from a video foundation model, our architecture features a latent illumination module that dynamically injects target environment maps into spatial features via cross-attention. Furthermore, we employ permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias. To train this robust data-driven model, we construct the massive Laval Objaverse Dataset (LOD), comprising 90K objects and 39K unique illuminations. Extensive experiments demonstrate state-of-the-art visual quality, photorealistic relighting quality, and strong zero-shot generalization across single-view, multi-view, and novel-view relighting tasks.</p>\n","updatedAt":"2026-09-09T08:15:45.459Z","author":{"_id":"6543bb6541540cfc6b9bde66","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/s5afGyHhOTQ1dkQpNYbHJ.png","fullname":"vLAR Group - HK PolyU","name":"vLAR","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8173303008079529},"editors":["vLAR"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/s5afGyHhOTQ1dkQpNYbHJ.png"],"reactions":[{"reaction":"🚀","users":["vLAR"],"count":1}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.07414","authors":[{"_id":"6aa1151dd0174964227beee0","name":"Hejun Wang","hidden":false},{"_id":"6aa1151dd0174964227beee1","name":"Jinxi Li","hidden":false},{"_id":"6aa1151dd0174964227beee2","name":"Junwei Jiang","hidden":false},{"_id":"6aa1151dd0174964227beee3","name":"Shiwei Mao","hidden":false},{"_id":"6aa1151dd0174964227beee4","name":"Hu Cheng","hidden":false},{"_id":"6aa1151dd0174964227beee5","name":"Shouwang Huang","hidden":false},{"_id":"6aa1151dd0174964227beee6","name":"Bo Yang","hidden":false}],"publishedAt":"2026-09-07T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting","submittedOnDailyBy":{"_id":"6543bb6541540cfc6b9bde66","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/s5afGyHhOTQ1dkQpNYbHJ.png","isPro":true,"fullname":"vLAR Group - HK PolyU","user":"vLAR","type":"user","name":"vLAR"},"summary":"Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed optimization, or single-image generative models that ignore crucial multi-view cues necessary for understanding 3D geometry and material interactions. To address these limitations, we introduce a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation. Adapted from a video foundation model, our architecture features a latent illumination module that dynamically injects target environment maps into spatial features via cross-attention. Furthermore, we employ permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias. To train this robust data-driven model, we construct the massive Laval Objaverse Dataset (LOD), comprising 90K objects and 39K unique illuminations. Extensive experiments demonstrate state-of-the-art visual quality, photorealistic relighting quality, and strong zero-shot generalization across single-view, multi-view, and novel-view relighting tasks.","upvotes":7,"discussionId":"6aa1151dd0174964227beee7","projectPage":"https://github.com/vLAR-group/RelightFormer","githubRepo":"https://github.com/vLAR-group/RelightFormer","githubRepoAddedBy":"user","ai_summary":"A feed-forward generative Transformer enables direct single- and multi-view image relighting by injecting target illumination via cross-attention and processing unordered views with permutation-invariant encodings, trained on a large synthetic dataset.","ai_keywords":["feed-forward generative Transformer","cross-attention","latent illumination module","permutation-invariant positional encodings","Laval Objaverse Dataset"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"646ecc368d316fde87b3b6e3","name":"PolyUHK","fullname":"The Hong Kong Polytechnic University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/646ecbc0cbb7bb996513e298/Akb4zKqIP9kb9PQoUPUmj.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6543bb6541540cfc6b9bde66","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/s5afGyHhOTQ1dkQpNYbHJ.png","isPro":true,"fullname":"vLAR Group - HK PolyU","user":"vLAR","type":"user"},{"_id":"6a6aa3fbb58832f7d0fd7e32","avatarUrl":"/avatars/e2b99552157ef3356352ec0af387396e.svg","isPro":false,"fullname":"Sarah Smith","user":"cobaltEvan","type":"user"},{"_id":"6a6d4372a5a4538841b20b30","avatarUrl":"/avatars/cca9481b6e02179085d0574d464d7299.svg","isPro":false,"fullname":"Linda Perez","user":"vectorDawn","type":"user"},{"_id":"6a701f86bc4016943249be07","avatarUrl":"/avatars/928c7a44ae1bcec57c75503d8fe7adb7.svg","isPro":false,"fullname":"Jennifer Clark","user":"RapidBeacon","type":"user"},{"_id":"6a9ae84928cbb1649f95bed8","avatarUrl":"/avatars/72f8dcd0450c38cb6bd3cc13b62e76fd.svg","isPro":false,"fullname":"蒋鑫","user":"VectorBloom75","type":"user"},{"_id":"6aa0e3451f6c4b2d2c287dda","avatarUrl":"/avatars/1986a0eef6b534299b995fb180b108ad.svg","isPro":false,"fullname":"Kevin Thompson","user":"ghays47","type":"user"},{"_id":"6aa0f3cfc471b4ecb36b8394","avatarUrl":"/avatars/374db28d56379c14cc419fe67c80f9be.svg","isPro":false,"fullname":"강시우","user":"Rapid-Glade","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"646ecc368d316fde87b3b6e3","name":"PolyUHK","fullname":"The Hong Kong Polytechnic University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/646ecbc0cbb7bb996513e298/Akb4zKqIP9kb9PQoUPUmj.jpeg"},"query":{}}">
RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting
Abstract
A feed-forward generative Transformer enables direct single- and multi-view image relighting by injecting target illumination via cross-attention and processing unordered views with permutation-invariant encodings, trained on a large synthetic dataset.
Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed optimization, or single-image generative models that ignore crucial multi-view cues necessary for understanding 3D geometry and material interactions. To address these limitations, we introduce a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation. Adapted from a video foundation model, our architecture features a latent illumination module that dynamically injects target environment maps into spatial features via cross-attention. Furthermore, we employ permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias. To train this robust data-driven model, we construct the massive Laval Objaverse Dataset (LOD), comprising 90K objects and 39K unique illuminations. Extensive experiments demonstrate state-of-the-art visual quality, photorealistic relighting quality, and strong zero-shot generalization across single-view, multi-view, and novel-view relighting tasks.
Community
Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed optimization, or single-image generative models that ignore crucial multi-view cues necessary for understanding 3D geometry and material interactions. To address these limitations, we introduce a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation. Adapted from a video foundation model, our architecture features a latent illumination module that dynamically injects target environment maps into spatial features via cross-attention. Furthermore, we employ permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias. To train this robust data-driven model, we construct the massive Laval Objaverse Dataset (LOD), comprising 90K objects and 39K unique illuminations. Extensive experiments demonstrate state-of-the-art visual quality, photorealistic relighting quality, and strong zero-shot generalization across single-view, multi-view, and novel-view relighting tasks.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.07414 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.07414 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.07414 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.