\n\t<a id=\"🔍-overview\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#🔍-overview\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\t🔍 Overview\n\t</span>\n</h2>\n<p><strong>Oxygen-TryOn</strong> is a unified, open-source foundation model for <strong>any-item virtual try-on</strong>. Given one or more reference items — provided either as clean product shots or as in-the-wild photos of someone already wearing them — together with a single target subject image, the model synthesizes a photorealistic image of that subject wearing the referenced items, spanning <strong>virtually any fashion category</strong>: clothing, outerwear, accessories, footwear, bags, and beyond.</p>\n<p>Unlike general-purpose image editors merely <em>prompted</em> for the task, Oxygen-TryOn is <strong>fashion-native</strong>: instead of treating try-on as mask-based inpainting, it reformulates it as a <strong>multi-reference, understanding-driven generation task</strong>, and is built specifically for try-on through a dedicated data engine and try-on-specific training. It accepts a variable number of references, composes multiple items in a single generation pass, and reasons holistically about layering and occlusion across full-body and half-body views, diverse poses, and non-standard subjects.</p>\n<p>Under the hood, Oxygen-TryOn is built on the <strong>JoyAI-Image-Edit</strong> architecture and initialized from its pretrained weights, coupling a multimodal large language model (MLLM) for reference and instruction understanding with a multimodal diffusion transformer (MMDiT) for high-fidelity synthesis. It is trained with a three-stage recipe — continued pre-training (CPT), large-scale supervised fine-tuning (SFT), and reinforcement learning (RL) under a hybrid reward — and retains the general instruction-based editing ability of its foundation (e.g., pose change) within the same generation pass.</p>\n<p>To our knowledge, Oxygen-TryOn is the <strong>first open-source system</strong> to deliver any-item, multi-reference try-on at this level of fidelity, achieving state-of-the-art consistency and realism that surpasses strong proprietary systems such as Nano Banana Pro, GPT-Image-2, and Seedream5 Lite, as well as leading open-source models such as FLUX.2.</p>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"✨-key-features\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#✨-key-features\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\t✨ Key Features\n\t</span>\n</h2>\n<ul>\n<li>🧥 <strong>Any item, any combination</strong> — garments, outerwear, accessories, shoes, bags, and more; from a single item to free multi-item outfits, with the model resolving layering and occlusion (\"OOTD\"-style full-outfit composition).</li>\n<li>🖼️ <strong>Heterogeneous references</strong> — accepts both clean product shots and in-the-wild worn-on photos; full- or half-body subjects with a variable number of references.</li>\n<li>🧍 <strong>Faithful preservation</strong> — keeps both the subject's identity and the referenced items' appearance intact.</li>\n<li>✏️ <strong>Built-in editing</strong> — general instruction-based edits (e.g., pose change) within the same generation pass, with no second model or pass.</li>\n<li>🎭 <strong>Cross-domain generalization</strong> — even dresses stylized 3D avatars, illustrated characters, statues, or posters while respecting the original style and geometry.</li>\n<li>🏆 <strong>State-of-the-art</strong> single-item consistency & realism, surpassing strong proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and leading open-source models (FLUX.2).</li>\n</ul>\n","updatedAt":"2026-07-28T03:52:54.561Z","author":{"_id":"63993093afe0d224cf2ba54a","avatarUrl":"/avatars/67268a3bdfa0ffbe7111fe3ab1cb83da.svg","fullname":"YongLiu","name":"LYAWWH","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":9,"identifiedLanguage":{"language":"en","probability":0.8417012691497803},"editors":["LYAWWH"],"editorAvatarUrls":["/avatars/67268a3bdfa0ffbe7111fe3ab1cb83da.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.21694","authors":[{"_id":"6a67548d620d9a1bd3b2d407","user":{"_id":"63993093afe0d224cf2ba54a","avatarUrl":"/avatars/67268a3bdfa0ffbe7111fe3ab1cb83da.svg","isPro":false,"fullname":"YongLiu","user":"LYAWWH","type":"user","name":"LYAWWH"},"name":"Yong Liu","status":"claimed_verified","statusLastChangedAt":"2026-07-27T16:45:05.497Z","hidden":false},{"_id":"6a67548d620d9a1bd3b2d408","user":{"_id":"670ce8badeb97b8b72e6e6e6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/_x-e95IYnkVQLb2KV7nIv.png","isPro":false,"fullname":"fuxiaolong","user":"fxlong","type":"user","name":"fxlong"},"name":"Xiaolong Fu","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:45:04.513Z","hidden":false},{"_id":"6a67548d620d9a1bd3b2d409","user":{"_id":"64b8ff3eee98fa83542c8284","avatarUrl":"/avatars/63fa3312e8ec5533c222211b83414091.svg","isPro":false,"fullname":"Zihang Xu","user":"xxxzh","type":"user","name":"xxxzh"},"name":"Zihang Xu","status":"claimed_verified","statusLastChangedAt":"2026-07-28T16:45:04.613Z","hidden":false},{"_id":"6a67548d620d9a1bd3b2d40a","user":{"_id":"65edc5b64f2eb015855afc0c","avatarUrl":"/avatars/12c03399feb8d5f3a10fb8e01eda1987.svg","isPro":false,"fullname":"xue wen","user":"Davidscut","type":"user","name":"Davidscut"},"name":"Wen Xue","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:45:04.517Z","hidden":false},{"_id":"6a67548d620d9a1bd3b2d40b","name":"Xueheng Li","hidden":false},{"_id":"6a67548d620d9a1bd3b2d40c","name":"Lin Song","hidden":false},{"_id":"6a67548d620d9a1bd3b2d40d","name":"Yuan Zhang","hidden":false},{"_id":"6a67548d620d9a1bd3b2d40e","name":"Chuyang Zhao","hidden":false},{"_id":"6a67548d620d9a1bd3b2d40f","name":"Haoyang Huang","hidden":false},{"_id":"6a67548d620d9a1bd3b2d410","name":"Nan Duan","hidden":false},{"_id":"6a67548d620d9a1bd3b2d411","name":"Yipeng Sun","hidden":false},{"_id":"6a67548d620d9a1bd3b2d412","name":"Yan Li","hidden":false},{"_id":"6a67548d620d9a1bd3b2d413","name":"Simiu Gu","hidden":false}],"publishedAt":"2026-07-23T00:00:00.000Z","submittedOnDailyAt":"2026-07-28T00:00:00.000Z","title":"Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On","submittedOnDailyBy":{"_id":"63993093afe0d224cf2ba54a","avatarUrl":"/avatars/67268a3bdfa0ffbe7111fe3ab1cb83da.svg","isPro":false,"fullname":"YongLiu","user":"LYAWWH","type":"user","name":"LYAWWH"},"summary":"We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for try-on through a dedicated data engine and try-on-specific training. Given one or more reference items (clean product shots or in-the-wild worn-on photos) and a single target subject image, it synthesizes a photorealistic image of the subject wearing the items across virtually any fashion category. Prior systems handle a single garment category in a studio setting, and recent multi-reference methods remain garment-centric; in contrast, Oxygen-TryOn supports diverse items and scenarios, including full- and half-body views, a variable number of references, and free multi-item composition, while faithfully preserving both subject identity and item appearance. Instead of mask-based inpainting, we reformulate try-on as a multi-reference, understanding-driven generation task. We build a data engine that collects, manufactures, annotates, and filters high-quality try-on data at scale, and design a three-stage recipe of continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). The RL stage uses a hybrid reward combining an in-house try-on reward model with a proprietary, rubric-guided general-purpose model, jointly supervising fine-grained consistency and instruction-level quality. It also follows general editing instructions (e.g., pose changes) in the same pass. Across public benchmarks and our in-house Oxygen-TryOn Bench, it achieves state-of-the-art consistency and realism on single-item try-on and leads on multi-item try-on, matching or surpassing both leading proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and open-source models (FLUX.2).","upvotes":20,"discussionId":"6a67548e620d9a1bd3b2d414","projectPage":"https://oxygenvision.github.io/Oxygen-TryOn/","organization":{"_id":"648146059860cd75c2614575","name":"JD-company","fullname":"JD.com","avatar":"https://www.gravatar.com/avatar/cb23650a1f9439629ffd5b495119abea?d=retro&size=100"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63993093afe0d224cf2ba54a","avatarUrl":"/avatars/67268a3bdfa0ffbe7111fe3ab1cb83da.svg","isPro":false,"fullname":"YongLiu","user":"LYAWWH","type":"user"},{"_id":"670ce8badeb97b8b72e6e6e6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/_x-e95IYnkVQLb2KV7nIv.png","isPro":false,"fullname":"fuxiaolong","user":"fxlong","type":"user"},{"_id":"68946c01cac26deedc53dd1a","avatarUrl":"/avatars/f0410ec479e14adc36f19ebbcc44e6cc.svg","isPro":false,"fullname":"zyj","user":"zyjcode","type":"user"},{"_id":"65edc5b64f2eb015855afc0c","avatarUrl":"/avatars/12c03399feb8d5f3a10fb8e01eda1987.svg","isPro":false,"fullname":"xue wen","user":"Davidscut","type":"user"},{"_id":"6926f87c2fe1e0de176f4162","avatarUrl":"/avatars/7bdc515214283c1080f71c404c2d7ff6.svg","isPro":false,"fullname":"Weidi Zhang","user":"csweidi","type":"user"},{"_id":"68500c915e818395b2a7f37e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/XduJSQCU1hGJP0CkREfNL.png","isPro":false,"fullname":"Zhang Jun Cheng","user":"zhangjuncheng","type":"user"},{"_id":"64b8ff3eee98fa83542c8284","avatarUrl":"/avatars/63fa3312e8ec5533c222211b83414091.svg","isPro":false,"fullname":"Zihang Xu","user":"xxxzh","type":"user"},{"_id":"61791bfa9e61d7ff71b52b98","avatarUrl":"/avatars/1af34650f2326c6a46bb31b38a3be165.svg","isPro":false,"fullname":"P. Sun","user":"tuniao","type":"user"},{"_id":"651ed7ef755e92f7f12742e6","avatarUrl":"/avatars/57a9cc189b4a59299aad6c96191b18d8.svg","isPro":true,"fullname":"yu li","user":"lyabc","type":"user"},{"_id":"667f51ad74fb1736a43bd32e","avatarUrl":"/avatars/55a25104883942ff7dfe080283064623.svg","isPro":false,"fullname":"Jiahe Guo","user":"MuyuenLP","type":"user"},{"_id":"64b34404520fa3a154dc4e4c","avatarUrl":"/avatars/f1fa3ab19d35f20bb9a30bf2ae1a54bc.svg","isPro":false,"fullname":"xinwei yang","user":"xw7777777","type":"user"},{"_id":"682d81e389b5f79a1d7b9682","avatarUrl":"/avatars/4e0b4f1654de8c74da5b54b154de3e65.svg","isPro":false,"fullname":"SuleBai","user":"SuleBai","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"648146059860cd75c2614575","name":"JD-company","fullname":"JD.com","avatar":"https://www.gravatar.com/avatar/cb23650a1f9439629ffd5b495119abea?d=retro&size=100"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.21694.md","query":{}}">
Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On
Abstract
We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for try-on through a dedicated data engine and try-on-specific training. Given one or more reference items (clean product shots or in-the-wild worn-on photos) and a single target subject image, it synthesizes a photorealistic image of the subject wearing the items across virtually any fashion category. Prior systems handle a single garment category in a studio setting, and recent multi-reference methods remain garment-centric; in contrast, Oxygen-TryOn supports diverse items and scenarios, including full- and half-body views, a variable number of references, and free multi-item composition, while faithfully preserving both subject identity and item appearance. Instead of mask-based inpainting, we reformulate try-on as a multi-reference, understanding-driven generation task. We build a data engine that collects, manufactures, annotates, and filters high-quality try-on data at scale, and design a three-stage recipe of continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). The RL stage uses a hybrid reward combining an in-house try-on reward model with a proprietary, rubric-guided general-purpose model, jointly supervising fine-grained consistency and instruction-level quality. It also follows general editing instructions (e.g., pose changes) in the same pass. Across public benchmarks and our in-house Oxygen-TryOn Bench, it achieves state-of-the-art consistency and realism on single-item try-on and leads on multi-item try-on, matching or surpassing both leading proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and open-source models (FLUX.2).
Community
🔍 Overview
Oxygen-TryOn is a unified, open-source foundation model for any-item virtual try-on. Given one or more reference items — provided either as clean product shots or as in-the-wild photos of someone already wearing them — together with a single target subject image, the model synthesizes a photorealistic image of that subject wearing the referenced items, spanning virtually any fashion category: clothing, outerwear, accessories, footwear, bags, and beyond.
Unlike general-purpose image editors merely prompted for the task, Oxygen-TryOn is fashion-native: instead of treating try-on as mask-based inpainting, it reformulates it as a multi-reference, understanding-driven generation task, and is built specifically for try-on through a dedicated data engine and try-on-specific training. It accepts a variable number of references, composes multiple items in a single generation pass, and reasons holistically about layering and occlusion across full-body and half-body views, diverse poses, and non-standard subjects.
Under the hood, Oxygen-TryOn is built on the JoyAI-Image-Edit architecture and initialized from its pretrained weights, coupling a multimodal large language model (MLLM) for reference and instruction understanding with a multimodal diffusion transformer (MMDiT) for high-fidelity synthesis. It is trained with a three-stage recipe — continued pre-training (CPT), large-scale supervised fine-tuning (SFT), and reinforcement learning (RL) under a hybrid reward — and retains the general instruction-based editing ability of its foundation (e.g., pose change) within the same generation pass.
To our knowledge, Oxygen-TryOn is the first open-source system to deliver any-item, multi-reference try-on at this level of fidelity, achieving state-of-the-art consistency and realism that surpasses strong proprietary systems such as Nano Banana Pro, GPT-Image-2, and Seedream5 Lite, as well as leading open-source models such as FLUX.2.
✨ Key Features
- 🧥 Any item, any combination — garments, outerwear, accessories, shoes, bags, and more; from a single item to free multi-item outfits, with the model resolving layering and occlusion ("OOTD"-style full-outfit composition).
- 🖼️ Heterogeneous references — accepts both clean product shots and in-the-wild worn-on photos; full- or half-body subjects with a variable number of references.
- 🧍 Faithful preservation — keeps both the subject's identity and the referenced items' appearance intact.
- ✏️ Built-in editing — general instruction-based edits (e.g., pose change) within the same generation pass, with no second model or pass.
- 🎭 Cross-domain generalization — even dresses stylized 3D avatars, illustrated characters, statues, or posters while respecting the original style and geometry.
- 🏆 State-of-the-art single-item consistency & realism, surpassing strong proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and leading open-source models (FLUX.2).
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.21694 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.21694 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.21694 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.