Hugging Face Daily Papers · · 6 min read

DREAM Technical Report

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<a href=\"https://cdn-uploads.huggingface.co/production/uploads/64e6fde9c0cc3e95d1d382f8/xJbrCvqbEcOJmE3vSctO0.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/64e6fde9c0cc3e95d1d382f8/xJbrCvqbEcOJmE3vSctO0.png\" alt=\"image\"></a></p>\n<p>Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across modules, rely on rigid rules, and have limited awareness of real-time intent, leaving session-level shifts among browsing, comparison, and purchase insufficiently addressed. We present DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them. DREAM has two core components. First, a three-tier Intent Engine fuses on-device signals into structured L0/L1/L2 intent representations; its edge-cloud trigger chain reduces reporting volume to approximately 8.7%. Second, a Meta Engine uses a MetaModel for layered M1-to-M2-to-M3 reasoning: intent summarization, strategy planning informed by Strategy Memory, and parameter translation. It dispatches the resulting parameters through a unified outlet with safety guardrails. A Reward Dual Loop continuously optimizes both components by combining offline simulation for strategy-space exploration with online feedback for outcome calibration, forming a cycle of generation, execution, evaluation, and experience accumulation. Large-scale A/B tests on Taobao's homepage feed show that re-ranking control alone improves IPV by 2.06%, Core IPV by 2.39%, and GMV by 0.88%. Extending control to fine ranking raises these gains to 2.71%, 3.06%, and 1.31%, respectively, while consistently improving PV by more than 1%. These gains require neither replacement of pipeline models nor compromise of serving stability, supporting agentic meta-control as a viable paradigm for industrial recommendation.</p>\n","updatedAt":"2026-08-26T10:31:16.425Z","author":{"_id":"64e6fde9c0cc3e95d1d382f8","avatarUrl":"/avatars/d14e323402a4b23197e84335dfa5e6ce.svg","fullname":"Yipeng Yu","name":"dev23user","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8766741156578064},"editors":["dev23user"],"editorAvatarUrls":["/avatars/d14e323402a4b23197e84335dfa5e6ce.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.09408","authors":[{"_id":"6a8d2a3a5add2537c32e97e1","name":"Bin Zhang","hidden":false},{"_id":"6a8d2a3a5add2537c32e97e2","name":"Bowen Zheng","hidden":false},{"_id":"6a8d2a3a5add2537c32e97e3","name":"Chao Yi","hidden":false},{"_id":"6a8d2a3a5add2537c32e97e4","name":"Chengyu Lai","hidden":false},{"_id":"6a8d2a3a5add2537c32e97e5","name":"Dian Chen","hidden":false},{"_id":"6a8d2a3a5add2537c32e97e6","name":"Dimin Wang","hidden":false},{"_id":"6a8d2a3a5add2537c32e97e7","name":"Gaoyang Guo","hidden":false},{"_id":"6a8d2a3a5add2537c32e97e8","name":"Jialin Zhu","hidden":false},{"_id":"6a8d2a3a5add2537c32e97e9","name":"Jian Wu","hidden":false},{"_id":"6a8d2a3a5add2537c32e97ea","name":"Jing Yu","hidden":false},{"_id":"6a8d2a3a5add2537c32e97eb","name":"Jiuning Lin","hidden":false},{"_id":"6a8d2a3a5add2537c32e97ec","name":"Lingqing Zhang","hidden":false},{"_id":"6a8d2a3a5add2537c32e97ed","name":"Lingyun Zheng","hidden":false},{"_id":"6a8d2a3a5add2537c32e97ee","name":"Mao Zhang","hidden":false},{"_id":"6a8d2a3a5add2537c32e97ef","name":"Mingming Pan","hidden":false},{"_id":"6a8d2a3a5add2537c32e97f0","name":"Ruiquan Lan","hidden":false},{"_id":"6a8d2a3a5add2537c32e97f1","name":"Shuai Zhong","hidden":false},{"_id":"6a8d2a3a5add2537c32e97f2","name":"Wen Chen","hidden":false},{"_id":"6a8d2a3a5add2537c32e97f3","name":"Wendong Zhang","hidden":false},{"_id":"6a8d2a3a5add2537c32e97f4","name":"Xiaodong Zhu","hidden":false},{"_id":"6a8d2a3a5add2537c32e97f5","name":"Xuan Chen","hidden":false},{"_id":"6a8d2a3a5add2537c32e97f6","name":"Xunke Xi","hidden":false},{"_id":"6a8d2a3a5add2537c32e97f7","name":"Yifan Lu","hidden":false},{"_id":"6a8d2a3a5add2537c32e97f8","name":"Yiheng Wang","hidden":false},{"_id":"6a8d2a3a5add2537c32e97f9","name":"Yue Zeng","hidden":false},{"_id":"6a8d2a3a5add2537c32e97fa","name":"Yujie Luo","hidden":false},{"_id":"6a8d2a3a5add2537c32e97fb","name":"Yuning Jiang","hidden":false},{"_id":"6a8d2a3a5add2537c32e97fc","name":"Zhe Hu","hidden":false},{"_id":"6a8d2a3a5add2537c32e97fd","name":"Zhibo Xiao","hidden":false},{"_id":"6a8d2a3a5add2537c32e97fe","name":"Zihong Huang","hidden":false},{"_id":"6a8d2a3a5add2537c32e97ff","name":"Binbin Cao","hidden":false},{"_id":"6a8d2a3a5add2537c32e9800","name":"Bo Zheng","hidden":false},{"_id":"6a8d2a3a5add2537c32e9801","name":"Danning Wang","hidden":false},{"_id":"6a8d2a3a5add2537c32e9802","name":"Dixuan Wang","hidden":false},{"_id":"6a8d2a3a5add2537c32e9803","name":"Ge Fan","hidden":false},{"_id":"6a8d2a3a5add2537c32e9804","name":"Haixia Wu","hidden":false},{"_id":"6a8d2a3a5add2537c32e9805","name":"Han Zhu","hidden":false},{"_id":"6a8d2a3a5add2537c32e9806","name":"Hao Fang","hidden":false},{"_id":"6a8d2a3a5add2537c32e9807","name":"Haoming Chen","hidden":false},{"_id":"6a8d2a3a5add2537c32e9808","name":"Huiping Chu","hidden":false},{"_id":"6a8d2a3a5add2537c32e9809","name":"Jian Wang","hidden":false},{"_id":"6a8d2a3a5add2537c32e980a","name":"Jianjun Wu","hidden":false},{"_id":"6a8d2a3a5add2537c32e980b","name":"Jiawei Wu","hidden":false},{"_id":"6a8d2a3a5add2537c32e980c","name":"Jiaxin Yu","hidden":false},{"_id":"6a8d2a3a5add2537c32e980d","name":"Jingwen Liu","hidden":false},{"_id":"6a8d2a3a5add2537c32e980e","name":"Jinzhe Shan","hidden":false},{"_id":"6a8d2a3a5add2537c32e980f","name":"Kai Meng","hidden":false},{"_id":"6a8d2a3a5add2537c32e9810","name":"Kai Zhang","hidden":false},{"_id":"6a8d2a3a5add2537c32e9811","name":"Keqin Xu","hidden":false},{"_id":"6a8d2a3a5add2537c32e9812","name":"Kewei Zhu","hidden":false},{"_id":"6a8d2a3a5add2537c32e9813","name":"Lang Tian","hidden":false},{"_id":"6a8d2a3a5add2537c32e9814","name":"Leihui Chen","hidden":false},{"_id":"6a8d2a3a5add2537c32e9815","name":"Li Chen","hidden":false},{"_id":"6a8d2a3a5add2537c32e9816","name":"Licheng Xu","hidden":false},{"_id":"6a8d2a3a5add2537c32e9817","name":"Lide Xiao","hidden":false},{"_id":"6a8d2a3a5add2537c32e9818","name":"Ruitong Zhang","hidden":false},{"_id":"6a8d2a3a5add2537c32e9819","name":"Shiyao Peng","hidden":false},{"_id":"6a8d2a3a5add2537c32e981a","name":"Silu Zhou","hidden":false},{"_id":"6a8d2a3a5add2537c32e981b","name":"Tao Wang","hidden":false},{"_id":"6a8d2a3a5add2537c32e981c","name":"Wei Shi","hidden":false},{"_id":"6a8d2a3a5add2537c32e981d","name":"Wenjun Yang","hidden":false},{"_id":"6a8d2a3a5add2537c32e981e","name":"Xiang Chen","hidden":false},{"_id":"6a8d2a3a5add2537c32e981f","name":"Xiang Gao","hidden":false},{"_id":"6a8d2a3a5add2537c32e9820","name":"Xiao Ren","hidden":false},{"_id":"6a8d2a3a5add2537c32e9821","name":"Xu Liu","hidden":false},{"_id":"6a8d2a3a5add2537c32e9822","name":"Xuwen Wang","hidden":false},{"_id":"6a8d2a3a5add2537c32e9823","name":"Yang Li","hidden":false},{"_id":"6a8d2a3a5add2537c32e9824","name":"Yeqiu Yang","hidden":false},{"_id":"6a8d2a3a5add2537c32e9825","name":"Yi Hu","hidden":false},{"_id":"6a8d2a3a5add2537c32e9826","name":"Yichen Yuan","hidden":false},{"_id":"6a8d2a3a5add2537c32e9827","name":"Yinnan Song","hidden":false},{"_id":"6a8d2a3a5add2537c32e9828","name":"Yipeng Yu","hidden":false},{"_id":"6a8d2a3a5add2537c32e9829","name":"Yuan Liu","hidden":false},{"_id":"6a8d2a3a5add2537c32e982a","name":"Yunqi Gao","hidden":false},{"_id":"6a8d2a3a5add2537c32e982b","name":"Zhiliang Huang","hidden":false},{"_id":"6a8d2a3a5add2537c32e982c","name":"Zhujin Gao","hidden":false},{"_id":"6a8d2a3a5add2537c32e982d","name":"Zongyuan Wu","hidden":false}],"publishedAt":"2026-08-13T00:00:00.000Z","submittedOnDailyAt":"2026-08-26T00:00:00.000Z","title":"DREAM Technical Report","submittedOnDailyBy":{"_id":"64e6fde9c0cc3e95d1d382f8","avatarUrl":"/avatars/d14e323402a4b23197e84335dfa5e6ce.svg","isPro":false,"fullname":"Yipeng Yu","user":"dev23user","type":"user","name":"dev23user"},"summary":"Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across modules, rely on rigid rules, and have limited awareness of real-time intent, leaving session-level shifts among browsing, comparison, and purchase insufficiently addressed. We present DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them. DREAM has two core components. First, a three-tier Intent Engine fuses on-device signals into structured L0/L1/L2 intent representations; its edge-cloud trigger chain reduces reporting volume to approximately 8.7%. Second, a Meta Engine uses a MetaModel for layered M1-to-M2-to-M3 reasoning: intent summarization, strategy planning informed by Strategy Memory, and parameter translation. It dispatches the resulting parameters through a unified outlet with safety guardrails. A Reward Dual Loop continuously optimizes both components by combining offline simulation for strategy-space exploration with online feedback for outcome calibration, forming a cycle of generation, execution, evaluation, and experience accumulation. Large-scale A/B tests on Taobao's homepage feed show that re-ranking control alone improves IPV by 2.06%, Core IPV by 2.39%, and GMV by 0.88%. Extending control to fine ranking raises these gains to 2.71%, 3.06%, and 1.31%, respectively, while consistently improving PV by more than 1%. These gains require neither replacement of pipeline models nor compromise of serving stability, supporting agentic meta-control as a viable paradigm for industrial recommendation.","upvotes":3,"discussionId":"6a8d2a3b5add2537c32e982e","ai_summary":"DREAM introduces an agentic meta-control layer over industrial recommender pipelines that uses intent reasoning and dual-loop optimization to improve session-level recommendations without replacing existing models.","ai_keywords":["agentic meta-control","Intent Engine","MetaModel","Strategy Memory","Reward Dual Loop","L0/L1/L2 intent representations","M1-to-M2-to-M3 reasoning","edge-cloud trigger chain"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"68be41370a3fcebdcad6516a","name":"alibabagroup","fullname":"alibaba","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68be3ab7e52df040b2cf80dc/li4G29u_EGswyTN1Sm_Kq.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64e6fde9c0cc3e95d1d382f8","avatarUrl":"/avatars/d14e323402a4b23197e84335dfa5e6ce.svg","isPro":false,"fullname":"Yipeng Yu","user":"dev23user","type":"user"},{"_id":"6406d8f4d684369027164e41","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6406d8f4d684369027164e41/vHRepecrzKtxdm0q52lVd.jpeg","isPro":false,"fullname":"sunyuhan","user":"yuuhan","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68be41370a3fcebdcad6516a","name":"alibabagroup","fullname":"alibaba","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68be3ab7e52df040b2cf80dc/li4G29u_EGswyTN1Sm_Kq.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.09408.md","query":{}}">
Papers
arxiv:2608.09408

DREAM Technical Report

Published on Aug 13
· Submitted by
Yipeng Yu
on Aug 26
Authors:
,

Abstract

DREAM introduces an agentic meta-control layer over industrial recommender pipelines that uses intent reasoning and dual-loop optimization to improve session-level recommendations without replacing existing models.

Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across modules, rely on rigid rules, and have limited awareness of real-time intent, leaving session-level shifts among browsing, comparison, and purchase insufficiently addressed. We present DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them. DREAM has two core components. First, a three-tier Intent Engine fuses on-device signals into structured L0/L1/L2 intent representations; its edge-cloud trigger chain reduces reporting volume to approximately 8.7%. Second, a Meta Engine uses a MetaModel for layered M1-to-M2-to-M3 reasoning: intent summarization, strategy planning informed by Strategy Memory, and parameter translation. It dispatches the resulting parameters through a unified outlet with safety guardrails. A Reward Dual Loop continuously optimizes both components by combining offline simulation for strategy-space exploration with online feedback for outcome calibration, forming a cycle of generation, execution, evaluation, and experience accumulation. Large-scale A/B tests on Taobao's homepage feed show that re-ranking control alone improves IPV by 2.06%, Core IPV by 2.39%, and GMV by 0.88%. Extending control to fine ranking raises these gains to 2.71%, 3.06%, and 1.31%, respectively, while consistently improving PV by more than 1%. These gains require neither replacement of pipeline models nor compromise of serving stability, supporting agentic meta-control as a viable paradigm for industrial recommendation.

Community

Paper submitter about 4 hours ago

image

Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across modules, rely on rigid rules, and have limited awareness of real-time intent, leaving session-level shifts among browsing, comparison, and purchase insufficiently addressed. We present DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them. DREAM has two core components. First, a three-tier Intent Engine fuses on-device signals into structured L0/L1/L2 intent representations; its edge-cloud trigger chain reduces reporting volume to approximately 8.7%. Second, a Meta Engine uses a MetaModel for layered M1-to-M2-to-M3 reasoning: intent summarization, strategy planning informed by Strategy Memory, and parameter translation. It dispatches the resulting parameters through a unified outlet with safety guardrails. A Reward Dual Loop continuously optimizes both components by combining offline simulation for strategy-space exploration with online feedback for outcome calibration, forming a cycle of generation, execution, evaluation, and experience accumulation. Large-scale A/B tests on Taobao's homepage feed show that re-ranking control alone improves IPV by 2.06%, Core IPV by 2.39%, and GMV by 0.88%. Extending control to fine ranking raises these gains to 2.71%, 3.06%, and 1.31%, respectively, while consistently improving PV by more than 1%. These gains require neither replacement of pipeline models nor compromise of serving stability, supporting agentic meta-control as a viable paradigm for industrial recommendation.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.09408
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.09408 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.09408 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.09408 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers