Industrial recommenders increasingly adopt the pretrain-then-transfer paradigm, yet behavioral distribution drift raises two questions: what to learn from behavior sequences, and how to transfer the learned knowledge while the pretrained model is continually refreshed. To resolve them, we propose Knowledge-Geometry Decoupling (KGD). For what to learn, conventional next-token prediction treats adjacency as dependency and may encode spurious transitions across unrelated sessions. We introduce Behavioral Multi-Token Prediction (BMTP) to retain only collaboratively or semantically related future items as supervision, yielding cleaner and more transferable behavioral knowledge. For how to transfer, pretrained knowledge and task-specific geometry impose conflicting optimization demands on shared parameters. To handle it, KGD assigns them to separate parameter sets: a refreshable encoder owns behavioral knowledge, while a task learner reads contextualized encoder states through read-only cross-attention and writes task-specific geometry through Anchored Calibration Residual (ACR) orthogonal to the pretrained embedding. The decoupled ownership enables continual knowledge refresh without task-gradient interference or invalidating downstream adaptation. KGD improves over strong pretrain-transfer baselines by 4-12% on eight public benchmarks and sustains its advantage over a 90-day production stream where baselines show no gains. KGD has been fully deployed in Shopee. In a live A/B test on Shopee Homepage Search, it increases GMV per user by 1.75% and advertising revenue by 1.53%, demonstrating its high practical value. We provide the core implementation of KGD at <a href=\"https://github.com/FuCongResearchSquad/KGD4REC\" rel=\"nofollow\">https://github.com/FuCongResearchSquad/KGD4REC</a></p>\n","updatedAt":"2026-08-05T03:34:50.177Z","author":{"_id":"665e8515045fcbf12b99558a","avatarUrl":"/avatars/e4ea26703d9dd2c8f19412ccd81d2cae.svg","fullname":"Fu Cong","name":"fcthebrave","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8637149930000305},"editors":["fcthebrave"],"editorAvatarUrls":["/avatars/e4ea26703d9dd2c8f19412ccd81d2cae.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.02738","authors":[{"_id":"6a7296061a375f948521c330","name":"Zixuan Wang","hidden":false},{"_id":"6a7296061a375f948521c331","name":"Yuhong Chen","hidden":false},{"_id":"6a7296061a375f948521c332","name":"Yuxuan Zhu","hidden":false},{"_id":"6a7296061a375f948521c333","name":"Guidong Lei","hidden":false},{"_id":"6a7296061a375f948521c334","name":"Zhiluohan Guo","hidden":false},{"_id":"6a7296061a375f948521c335","name":"Yu Zhao","hidden":false},{"_id":"6a7296061a375f948521c336","name":"Kun Wang","hidden":false},{"_id":"6a7296061a375f948521c337","name":"Bangyang Hong","hidden":false},{"_id":"6a7296061a375f948521c338","name":"Kangle Wu","hidden":false},{"_id":"6a7296061a375f948521c339","name":"Yabo Ni","hidden":false},{"_id":"6a7296061a375f948521c33a","name":"Anxiang Zeng","hidden":false},{"_id":"6a7296061a375f948521c33b","name":"Cong Fu","hidden":false},{"_id":"6a7296061a375f948521c33c","name":"Hui Li","hidden":false}],"publishedAt":"2026-08-03T00:00:00.000Z","submittedOnDailyAt":"2026-08-05T00:00:00.000Z","title":"Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation","submittedOnDailyBy":{"_id":"665e8515045fcbf12b99558a","avatarUrl":"/avatars/e4ea26703d9dd2c8f19412ccd81d2cae.svg","isPro":false,"fullname":"Fu Cong","user":"fcthebrave","type":"user","name":"fcthebrave"},"summary":"Industrial recommenders increasingly adopt the pretrain-then-transfer paradigm, yet behavioral distribution drift raises two questions: what to learn from behavior sequences, and how to transfer the learned knowledge while the pretrained model is continually refreshed. To resolve them, we propose Knowledge-Geometry Decoupling (KGD). For what to learn, conventional next-token prediction treats adjacency as dependency and may encode spurious transitions across unrelated sessions. We introduce Behavioral Multi-Token Prediction (BMTP) to retain only collaboratively or semantically related future items as supervision, yielding cleaner and more transferable behavioral knowledge. For how to transfer, pretrained knowledge and task-specific geometry impose conflicting optimization demands on shared parameters. To handle it, KGD assigns them to separate parameter sets: a refreshable encoder owns behavioral knowledge, while a task learner reads contextualized encoder states through read-only cross-attention and writes task-specific geometry through Anchored Calibration Residual (ACR) orthogonal to the pretrained embedding. The decoupled ownership enables continual knowledge refresh without task-gradient interference or invalidating downstream adaptation. KGD improves over strong pretrain-transfer baselines by 4-12% on eight public benchmarks and sustains its advantage over a 90-day production stream where baselines show no gains. KGD has been fully deployed in Shopee. In a live A/B test on Shopee Homepage Search, it increases GMV per user by 1.75% and advertising revenue by 1.53%, demonstrating its high practical value. We provide the core implementation of KGD at https://github.com/FuCongResearchSquad/KGD4REC.","upvotes":31,"discussionId":"6a7296061a375f948521c33d","githubRepo":"https://github.com/FuCongResearchSquad/KGD4REC","githubRepoAddedBy":"user","githubStars":0,"organization":{"_id":"693bcc3aca53da32fcf95a61","name":"PIIR","fullname":"Personalized&Inclusive Intelligence Research","avatar":"https://www.gravatar.com/avatar/cf1ab1e4c4a4b41ef5692c9a8f6e6d4f?d=retro&size=100"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"665e8515045fcbf12b99558a","avatarUrl":"/avatars/e4ea26703d9dd2c8f19412ccd81d2cae.svg","isPro":false,"fullname":"Fu Cong","user":"fcthebrave","type":"user"},{"_id":"67e3c591d8c74e4fbf726755","avatarUrl":"/avatars/056cf8cba50e52affb98c9e6fc0fed8f.svg","isPro":false,"fullname":"wangzixuan","user":"williamzx","type":"user"},{"_id":"641f18a673cfc036ddbaeccf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/RtOK_qV71NM87A67XFJCN.png","isPro":false,"fullname":"Yuxuan Zhu","user":"Ethan7","type":"user"},{"_id":"6877568698c692330d8c7241","avatarUrl":"/avatars/297d46b560701dcdb0d09286177a75f3.svg","isPro":false,"fullname":"CC","user":"ccc131","type":"user"},{"_id":"64a6b2adacb2711f38fe646b","avatarUrl":"/avatars/a11f3db3284ebe826dc6fbc53fce9f66.svg","isPro":false,"fullname":"Chen Bingjujn","user":"vnkyi","type":"user"},{"_id":"68d214fe2e59f503dbffb8c9","avatarUrl":"/avatars/fef3c37dc9dd9e4833d7ba05c9e293de.svg","isPro":false,"fullname":"guo","user":"guoehan","type":"user"},{"_id":"672dda02ee49faac3ac69510","avatarUrl":"/avatars/d264e8f9f8a9b2787ff4768dfc371a2e.svg","isPro":false,"fullname":"Lehua He","user":"yycloudywind","type":"user"},{"_id":"69ddb030c25e2efd3b4d2bb5","avatarUrl":"/avatars/21dd3d6d611d3937b09577fcfcc5ba04.svg","isPro":false,"fullname":"boat","user":"excitinnnnnng","type":"user"},{"_id":"68d218c230c41a74dbee3564","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/SjHI-qhzku51Sod3AUKxm.png","isPro":false,"fullname":"Jingzhi.Ding","user":"Manderin","type":"user"},{"_id":"67a32b2307690f2a5728dd01","avatarUrl":"/avatars/9c631045c47c82c8b403321dd3f8f0ec.svg","isPro":false,"fullname":"TheXiang","user":"TheXiang2000","type":"user"},{"_id":"694e1505a624c194a198a20c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/694e1505a624c194a198a20c/ZS98Y8FcGS3CQOH2wnqkd.jpeg","isPro":false,"fullname":"Yu Liang","user":"Ricardoly","type":"user"},{"_id":"69703d5be3c0eaed790701e9","avatarUrl":"/avatars/d0dc083df5df5db91fb29566372a974e.svg","isPro":false,"fullname":"zhoutao","user":"zhoutao-shopee","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"693bcc3aca53da32fcf95a61","name":"PIIR","fullname":"Personalized&Inclusive Intelligence Research","avatar":"https://www.gravatar.com/avatar/cf1ab1e4c4a4b41ef5692c9a8f6e6d4f?d=retro&size=100"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.02738.md","query":{}}">
Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation
Abstract
Industrial recommenders increasingly adopt the pretrain-then-transfer paradigm, yet behavioral distribution drift raises two questions: what to learn from behavior sequences, and how to transfer the learned knowledge while the pretrained model is continually refreshed. To resolve them, we propose Knowledge-Geometry Decoupling (KGD). For what to learn, conventional next-token prediction treats adjacency as dependency and may encode spurious transitions across unrelated sessions. We introduce Behavioral Multi-Token Prediction (BMTP) to retain only collaboratively or semantically related future items as supervision, yielding cleaner and more transferable behavioral knowledge. For how to transfer, pretrained knowledge and task-specific geometry impose conflicting optimization demands on shared parameters. To handle it, KGD assigns them to separate parameter sets: a refreshable encoder owns behavioral knowledge, while a task learner reads contextualized encoder states through read-only cross-attention and writes task-specific geometry through Anchored Calibration Residual (ACR) orthogonal to the pretrained embedding. The decoupled ownership enables continual knowledge refresh without task-gradient interference or invalidating downstream adaptation. KGD improves over strong pretrain-transfer baselines by 4-12% on eight public benchmarks and sustains its advantage over a 90-day production stream where baselines show no gains. KGD has been fully deployed in Shopee. In a live A/B test on Shopee Homepage Search, it increases GMV per user by 1.75% and advertising revenue by 1.53%, demonstrating its high practical value. We provide the core implementation of KGD at https://github.com/FuCongResearchSquad/KGD4REC.
Community
Industrial recommenders increasingly adopt the pretrain-then-transfer paradigm, yet behavioral distribution drift raises two questions: what to learn from behavior sequences, and how to transfer the learned knowledge while the pretrained model is continually refreshed. To resolve them, we propose Knowledge-Geometry Decoupling (KGD). For what to learn, conventional next-token prediction treats adjacency as dependency and may encode spurious transitions across unrelated sessions. We introduce Behavioral Multi-Token Prediction (BMTP) to retain only collaboratively or semantically related future items as supervision, yielding cleaner and more transferable behavioral knowledge. For how to transfer, pretrained knowledge and task-specific geometry impose conflicting optimization demands on shared parameters. To handle it, KGD assigns them to separate parameter sets: a refreshable encoder owns behavioral knowledge, while a task learner reads contextualized encoder states through read-only cross-attention and writes task-specific geometry through Anchored Calibration Residual (ACR) orthogonal to the pretrained embedding. The decoupled ownership enables continual knowledge refresh without task-gradient interference or invalidating downstream adaptation. KGD improves over strong pretrain-transfer baselines by 4-12% on eight public benchmarks and sustains its advantage over a 90-day production stream where baselines show no gains. KGD has been fully deployed in Shopee. In a live A/B test on Shopee Homepage Search, it increases GMV per user by 1.75% and advertising revenue by 1.53%, demonstrating its high practical value. We provide the core implementation of KGD at https://github.com/FuCongResearchSquad/KGD4REC
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.02738 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.02738 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.