Hugging Face Daily Papers · · 7 min read

RecGPT-V3 Technical Report

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

🚀 <strong>RecGPT-V3 Is Here: Making LLMs the “Intelligent Brain” of Industrial-Scale Recommender Systems</strong></p>\n<p>From <strong>RecGPT-V1</strong>, which pioneered intent understanding in Taobao’s recommendation pipeline, to <strong>RecGPT-V2</strong>, which expanded the breadth and efficiency of intent understanding through multi-expert collaboration, our mission has remained the same: to move recommender systems beyond <strong>“predicting the next click”</strong> toward <strong>“understanding the real needs behind every choice.”</strong></p>\n<p>Today, we introduce <strong>RecGPT-V3</strong>. It brings together <strong>continually evolving stateful user memory, hybrid-modal understanding that connects natural language with the item space, and efficient latent reasoning that can be decoded back into human-readable explanations</strong>. Together, these capabilities enable recommender systems to remember users, understand their intent more precisely, and perform deep reasoning at significantly lower cost.</p>\n<p>🧠 <strong>From Repeated Analysis to Continual Memory</strong><br>The <strong>Memory Hub</strong> distills long-term user behavior into structured memory that evolves over time. Instead of repeatedly processing the entire interaction history from scratch, the system continuously accumulates and refines its understanding of each user, reducing user-modeling compute by <strong>55.8%</strong>.</p>\n<p>🎯 <strong>From Ambiguous Intent to Precise Item Grounding</strong><br>The model jointly reasons over natural language and <strong>Semantic IDs (SIDs)</strong>. Natural language captures open-ended and complex user needs, while SIDs ground those needs directly in concrete items, bridging the gap between “understanding the user” and “finding the right products.”</p>\n<p>⚡ <strong>From Lengthy Explicit Reasoning to Efficient Latent Reasoning</strong><br><strong>Latent Intent Reasoning</strong> internalizes thousands of explicit reasoning tokens into a compact set of latent tokens, substantially reducing online inference costs. These latent reasoning traces can also be <strong>decoded back into human-readable explanations on demand</strong>, preserving interpretability while delivering a <strong>3.46× end-to-end inference speedup</strong>.</p>\n<p>RecGPT-V3 has now been deployed in the <strong>“Guess What You Like”</strong> feed on Taobao’s homepage. Large-scale online A/B tests have delivered significant improvements:</p>\n<p>📈 <strong>IPV +1.28%, CTR +1.00%, TC +1.97%, and GMV +3.97%</strong><br>📉 <strong>52.4% reduction in end-to-end serving resource consumption</strong></p>\n<p>From RecGPT-V1 to RecGPT-V3, we are steadily advancing recommendation from “replaying a user’s past” toward <strong>“continually remembering users, deeply understanding their intent, and efficiently delivering truly personalized experiences.”</strong> ✨</p>\n","updatedAt":"2026-07-20T01:58:30.609Z","author":{"_id":"65acfb3a14e6582c30b4ce76","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65acfb3a14e6582c30b4ce76/RhEhePggBtyM0RIIqXQen.jpeg","fullname":"TangJiakai","name":"TangJiakai5704","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8656882047653198},"editors":["TangJiakai5704"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/65acfb3a14e6582c30b4ce76/RhEhePggBtyM0RIIqXQen.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.15591","authors":[{"_id":"6a5d80c56a69ce099f4d6d41","name":"Bowen Zheng","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d42","name":"Chao Yi","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d43","name":"Dian Chen","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d44","name":"Gaoyang Guo","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d45","name":"Han Zhu","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d46","name":"Jiakai Tang","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d47","name":"Jian Wu","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d48","name":"Mao Zhang","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d49","name":"Wen Chen","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d4a","name":"Yifan Lu","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d4b","name":"Yujie Luo","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d4c","name":"Yuning Jiang","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d4d","name":"Zhujin Gao","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d4e","name":"Bo Zheng","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d4f","name":"Dixuan Wang","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d50","name":"Hao Fang","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d51","name":"Jiancai Liu","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d52","name":"Jing Yu","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d53","name":"Ke Chen","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d54","name":"Kewei Zhu","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d55","name":"Mingke Xu","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d56","name":"Wenjun Yang","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d57","name":"Xunke Xi","hidden":false},{"_id":"6a5d80c56a69ce099f4d6d58","name":"Zile Zhou","hidden":false}],"publishedAt":"2026-07-17T00:00:00.000Z","submittedOnDailyAt":"2026-07-20T00:00:00.000Z","title":"RecGPT-V3 Technical Report","submittedOnDailyBy":{"_id":"65acfb3a14e6582c30b4ce76","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65acfb3a14e6582c30b4ce76/RhEhePggBtyM0RIIqXQen.jpeg","isPro":false,"fullname":"TangJiakai","user":"TangJiakai5704","type":"user","name":"TangJiakai5704"},"summary":"Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in historical behavior toward reasoning about the intent that drives it. RecGPT-V1 pioneered this paradigm on Taobao by centering user understanding, and RecGPT-V2 scaled it via coordinated multi-agent reasoning; both are deployed in production with consistent gains in user experience and commercial outcomes. However, operating RecGPT at scale reveals three challenges: (1) stateless behavior modeling, where each request reprocesses full user history, wasting computation and discarding prior analysis; (2) a tag-to-item information bottleneck, where natural-language tags form a lossy channel between user understanding and item grounding; and (3) inefficient explicit reasoning, whose lengthy chain-of-thought incurs untenable latency and compute overhead.\n We present RecGPT-V3, a stateful, hybrid-modal recommender that reasons over natural language for open-world knowledge and Semantic IDs (SIDs) for concrete item grounding. A Memory Hub maintains structured, continually evolving user memory that distills long-horizon behavior into condensed units, cutting user-modeling computation by 55.8%. A Hybrid-modal Foundation Model allows the LLM jointly reason over text tags and SIDs, opening a high-bandwidth channel into the item space. Latent Intent Reasoning internalizes verbose rationales into compact learnable latent tokens that remain decodable into readable explanations, lowering output token cost by 200x. Deployed in Taobao's \"Guess What You Like\" feed, RecGPT-V3 achieves consistent gains in large-scale online A/B tests: IPV +1.28%, CTR +1.00%, TC +1.97%, GMV +3.97%, while cutting end-to-end serving resource consumption by 52.4%.","upvotes":22,"discussionId":"6a5d80c56a69ce099f4d6d59"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65acfb3a14e6582c30b4ce76","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65acfb3a14e6582c30b4ce76/RhEhePggBtyM0RIIqXQen.jpeg","isPro":false,"fullname":"TangJiakai","user":"TangJiakai5704","type":"user"},{"_id":"64ba481acd40b08a682e3579","avatarUrl":"/avatars/9a1b2569b368f446ad42a5716825e5f0.svg","isPro":false,"fullname":"wenyuer","user":"wenyuer","type":"user"},{"_id":"6a4539f015e14ca75b302a8a","avatarUrl":"/avatars/4f17dc791457b794f9daed0d41856c77.svg","isPro":false,"fullname":"Yezi Chu","user":"yezichu","type":"user"},{"_id":"667e6c4f1cb740c97401f127","avatarUrl":"/avatars/39ff8af6460790fa1a65fc02be45081c.svg","isPro":false,"fullname":"Yifan Lu","user":"lyfcan","type":"user"},{"_id":"643433e7546e16f17a13a35d","avatarUrl":"/avatars/01fb6adcd3c7aa87f38dea52e28fbbb0.svg","isPro":false,"fullname":"yang","user":"wenjunyang","type":"user"},{"_id":"642682cf13a9e5d9675a613f","avatarUrl":"/avatars/55d2397f94af8a413acfbaa6b9be5236.svg","isPro":false,"fullname":"chaoyi","user":"chao-yi","type":"user"},{"_id":"61dc205641c14d1ae1f47e2a","avatarUrl":"/avatars/36bb569eccab913ab193a5f1918f0448.svg","isPro":false,"fullname":"Luo","user":"sculuo96","type":"user"},{"_id":"66960f262dcf962073b2ad19","avatarUrl":"/avatars/27a284ec917b9fa4ca47a2683c5eb8b3.svg","isPro":false,"fullname":"Mi Yan","user":"M1YAN","type":"user"},{"_id":"65e2e646a0681de63059ce40","avatarUrl":"/avatars/64174bce24ca5827a3402b90e37aec5c.svg","isPro":false,"fullname":"Dixuan Wang","user":"Dixuan","type":"user"},{"_id":"645469c7363bb3aaf9ca9caf","avatarUrl":"/avatars/0a456a16c1447dd1dcd8d45b807af77c.svg","isPro":false,"fullname":"Gaoyang Guo","user":"hairlatic","type":"user"},{"_id":"64df7baf2aeee67969559893","avatarUrl":"/avatars/53d9a86c24e0389ff1fa0c0b7e1cf5b9.svg","isPro":false,"fullname":"Bowen Zheng","user":"bwzheng0324","type":"user"},{"_id":"674467a0e500cbe102885921","avatarUrl":"/avatars/33425251da8b2a38e6f72036d25953b4.svg","isPro":false,"fullname":"LJW","user":"axdyer","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.15591.md","query":{}}">
Papers
arxiv:2607.15591

RecGPT-V3 Technical Report

Published on Jul 17
· Submitted by
TangJiakai
on Jul 20
Authors:
,

Abstract

Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in historical behavior toward reasoning about the intent that drives it. RecGPT-V1 pioneered this paradigm on Taobao by centering user understanding, and RecGPT-V2 scaled it via coordinated multi-agent reasoning; both are deployed in production with consistent gains in user experience and commercial outcomes. However, operating RecGPT at scale reveals three challenges: (1) stateless behavior modeling, where each request reprocesses full user history, wasting computation and discarding prior analysis; (2) a tag-to-item information bottleneck, where natural-language tags form a lossy channel between user understanding and item grounding; and (3) inefficient explicit reasoning, whose lengthy chain-of-thought incurs untenable latency and compute overhead. We present RecGPT-V3, a stateful, hybrid-modal recommender that reasons over natural language for open-world knowledge and Semantic IDs (SIDs) for concrete item grounding. A Memory Hub maintains structured, continually evolving user memory that distills long-horizon behavior into condensed units, cutting user-modeling computation by 55.8%. A Hybrid-modal Foundation Model allows the LLM jointly reason over text tags and SIDs, opening a high-bandwidth channel into the item space. Latent Intent Reasoning internalizes verbose rationales into compact learnable latent tokens that remain decodable into readable explanations, lowering output token cost by 200x. Deployed in Taobao's "Guess What You Like" feed, RecGPT-V3 achieves consistent gains in large-scale online A/B tests: IPV +1.28%, CTR +1.00%, TC +1.97%, GMV +3.97%, while cutting end-to-end serving resource consumption by 52.4%.

Community

🚀 RecGPT-V3 Is Here: Making LLMs the “Intelligent Brain” of Industrial-Scale Recommender Systems

From RecGPT-V1, which pioneered intent understanding in Taobao’s recommendation pipeline, to RecGPT-V2, which expanded the breadth and efficiency of intent understanding through multi-expert collaboration, our mission has remained the same: to move recommender systems beyond “predicting the next click” toward “understanding the real needs behind every choice.”

Today, we introduce RecGPT-V3. It brings together continually evolving stateful user memory, hybrid-modal understanding that connects natural language with the item space, and efficient latent reasoning that can be decoded back into human-readable explanations. Together, these capabilities enable recommender systems to remember users, understand their intent more precisely, and perform deep reasoning at significantly lower cost.

🧠 From Repeated Analysis to Continual Memory
The Memory Hub distills long-term user behavior into structured memory that evolves over time. Instead of repeatedly processing the entire interaction history from scratch, the system continuously accumulates and refines its understanding of each user, reducing user-modeling compute by 55.8%.

🎯 From Ambiguous Intent to Precise Item Grounding
The model jointly reasons over natural language and Semantic IDs (SIDs). Natural language captures open-ended and complex user needs, while SIDs ground those needs directly in concrete items, bridging the gap between “understanding the user” and “finding the right products.”

From Lengthy Explicit Reasoning to Efficient Latent Reasoning
Latent Intent Reasoning internalizes thousands of explicit reasoning tokens into a compact set of latent tokens, substantially reducing online inference costs. These latent reasoning traces can also be decoded back into human-readable explanations on demand, preserving interpretability while delivering a 3.46× end-to-end inference speedup.

RecGPT-V3 has now been deployed in the “Guess What You Like” feed on Taobao’s homepage. Large-scale online A/B tests have delivered significant improvements:

📈 IPV +1.28%, CTR +1.00%, TC +1.97%, and GMV +3.97%
📉 52.4% reduction in end-to-end serving resource consumption

From RecGPT-V1 to RecGPT-V3, we are steadily advancing recommendation from “replaying a user’s past” toward “continually remembering users, deeply understanding their intent, and efficiently delivering truly personalized experiences.”

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.15591
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.15591 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.15591 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.15591 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers