Hugging Face Daily Papers · · 3 min read

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

To recursively self-improve, agents need to dream. History can be the world they dream in.</p>\n<p>More Details: <a href=\"https://dream-rsi.com/\" rel=\"nofollow\">https://dream-rsi.com/</a></p>\n","updatedAt":"2026-09-15T02:11:27.250Z","author":{"_id":"6623ea65b642e29cdf90a1b4","avatarUrl":"/avatars/e32e90574c1162b2be87ed78604e3e4d.svg","fullname":"TongZheng","name":"TongZheng1999","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":12,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8560773134231567},"editors":["TongZheng1999"],"editorAvatarUrls":["/avatars/e32e90574c1162b2be87ed78604e3e4d.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.14858","authors":[{"_id":"6aa8a7435dd4cb9b4cc024aa","name":"Tong Zheng","hidden":false},{"_id":"6aa8a7435dd4cb9b4cc024ab","name":"Xidong Wu","hidden":false},{"_id":"6aa8a7435dd4cb9b4cc024ac","name":"Zheng Zhang","hidden":false},{"_id":"6aa8a7435dd4cb9b4cc024ad","name":"Zhankui He","hidden":false},{"_id":"6aa8a7435dd4cb9b4cc024ae","name":"Chaoyi Zhang","hidden":false},{"_id":"6aa8a7435dd4cb9b4cc024af","name":"Benjamin Coleman","hidden":false},{"_id":"6aa8a7435dd4cb9b4cc024b0","name":"Ruoqiao Wei","hidden":false},{"_id":"6aa8a7435dd4cb9b4cc024b1","name":"Di Bai","hidden":false},{"_id":"6aa8a7435dd4cb9b4cc024b2","name":"Haolin Liu","hidden":false},{"_id":"6aa8a7435dd4cb9b4cc024b3","name":"Rui Liu","hidden":false},{"_id":"6aa8a7435dd4cb9b4cc024b4","name":"Xue Wang","hidden":false},{"_id":"6aa8a7435dd4cb9b4cc024b5","name":"Yue Zhuan","hidden":false},{"_id":"6aa8a7435dd4cb9b4cc024b6","name":"Wang-Cheng Kang","hidden":false},{"_id":"6aa8a7435dd4cb9b4cc024b7","name":"Renkai Xiang","hidden":false},{"_id":"6aa8a7435dd4cb9b4cc024b8","name":"Heng Huang","hidden":false},{"_id":"6aa8a7435dd4cb9b4cc024b9","name":"Xinwu Cheng","hidden":false},{"_id":"6aa8a7435dd4cb9b4cc024ba","name":"Yunsong Guo","hidden":false}],"publishedAt":"2026-09-14T00:00:00.000Z","submittedOnDailyAt":"2026-09-15T00:00:00.000Z","title":"Dream-RSI: Recursive Self-Improvement through Evolving Worlds","submittedOnDailyBy":{"_id":"6623ea65b642e29cdf90a1b4","avatarUrl":"/avatars/e32e90574c1162b2be87ed78604e3e4d.svg","isPro":true,"fullname":"TongZheng","user":"TongZheng1999","type":"user","name":"TongZheng1999"},"summary":"Recursive self-improvement is becoming increasingly vital for autonomous AI agents, where progress hinges on discovering high-value solutions across complex domains. The driver of this process is effective exploration, however, managing and improving exploration strategies remains a major bottleneck. Current systems face a fundamental dilemma: fixed strategies fail to adapt as search spaces scale, while online policy optimization requires navigating vast meta-search spaces under delayed and expensive feedback over long-horizon rollouts. We introduce Dream-RSI, a framework for scalable and recursively self-improving exploration. A lightweight orchestration layer makes exploration explicit and programmable while leaving the underlying coding agent unchanged. Our key insight is that accumulated discovery history can serve as a replay simulator over the realized search space. By performing dreaming in the replay simulator constructed from historical discovery trees, Dream-RSI secures immediate, low-cost off-policy feedback to evaluate and refine exploration policies without invoking repetitive, expensive online evaluations. The improved policy is subsequently redeployed online to drive further discovery, continuously expanding the simulator pool in a self-improving loop. Across algorithm engineering, mathematical optimization, and GPU kernel engineering, Dream-RSI achieves competitive or improved discovery quality while substantially reducing discovery cost in several settings.","upvotes":157,"discussionId":"6aa8a7445dd4cb9b4cc024bb","projectPage":"https://dream-rsi.com/","githubRepo":"https://github.com/zhengkid/Dream-RSI","githubRepoAddedBy":"user","ai_summary":"Dream-RSI enables scalable recursive self-improvement by using historical discovery replay to evaluate exploration policies offline, reducing costly online evaluations.","ai_keywords":["recursive self-improvement","exploration strategies","replay simulator","off-policy feedback","discovery trees","Dream-RSI"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":4,"organization":{"_id":"5e6aca39878b8b2bf9806447","name":"google","fullname":"Google","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/WtA3YYitedOr9n02eHfJe.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6623ea65b642e29cdf90a1b4","avatarUrl":"/avatars/e32e90574c1162b2be87ed78604e3e4d.svg","isPro":true,"fullname":"TongZheng","user":"TongZheng1999","type":"user"},{"_id":"65037565da2d88e201f63b7a","avatarUrl":"/avatars/d1b6ce17236360e9583b8bb4cb87e506.svg","isPro":false,"fullname":"Runpeng Dai","user":"Leo-Dai","type":"user"},{"_id":"6623e265609d7a39a1107cc9","avatarUrl":"/avatars/887f1f5b74f6669ce7d560d95b05b530.svg","isPro":false,"fullname":"Dawn","user":"LegendaryDawn","type":"user"},{"_id":"6850ab3fb73a72cdc6159f6f","avatarUrl":"/avatars/74c58336655cdfe7d3fb9854d426240d.svg","isPro":false,"fullname":"Haolin Liu","user":"lhl616","type":"user"},{"_id":"62ea79dd01ed9b0e8f61ccd3","avatarUrl":"/avatars/70af83e0e267be39fcd5f23b85e2dafa.svg","isPro":false,"fullname":"Chengsong Huang","user":"ChengsongHuang","type":"user"},{"_id":"6656bf615b203a05a1f0968c","avatarUrl":"/avatars/1ee0b0099c10dd76c8e3b7d312221b15.svg","isPro":false,"fullname":"Rui Liu","user":"lr10260","type":"user"},{"_id":"6488a4039d6109a4dd058e44","avatarUrl":"/avatars/2eac1e3a3e8189bbf8f712ef96290401.svg","isPro":false,"fullname":"YANSHUO CHEN","user":"alegendaryfish","type":"user"},{"_id":"65346df299d8bba29428970a","avatarUrl":"/avatars/f701ecd98be449758bd9ecdfd58f646d.svg","isPro":false,"fullname":"Louis Liu","user":"Louis0324","type":"user"},{"_id":"63c1699e40a26dd2db32400d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63c1699e40a26dd2db32400d/3N0-Zp8igv8-52mXAdiiq.jpeg","isPro":false,"fullname":"Chroma","user":"Chroma111","type":"user"},{"_id":"6683fc5344a65be1aab25dc0","avatarUrl":"/avatars/e13cde3f87b59e418838d702807df3b5.svg","isPro":false,"fullname":"hjkim","user":"hojie11","type":"user"},{"_id":"64ec2759bfb2aa06a469eccd","avatarUrl":"/avatars/552a3f67952e33252d24d3e1c100e837.svg","isPro":false,"fullname":"X C","user":"cavosamir","type":"user"},{"_id":"661eb07769dd840451351ea0","avatarUrl":"/avatars/fc024fc573b01ae88a027a81c1b74c42.svg","isPro":false,"fullname":"Chenhui Xu","user":"miniHui","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"5e6aca39878b8b2bf9806447","name":"google","fullname":"Google","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/WtA3YYitedOr9n02eHfJe.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.14858.md","query":{}}">
Papers
arxiv:2609.14858

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

Published on Sep 14
· Submitted by
TongZheng
on Sep 15
#2 Paper of the day
Authors:
,

Abstract

Dream-RSI enables scalable recursive self-improvement by using historical discovery replay to evaluate exploration policies offline, reducing costly online evaluations.

Recursive self-improvement is becoming increasingly vital for autonomous AI agents, where progress hinges on discovering high-value solutions across complex domains. The driver of this process is effective exploration, however, managing and improving exploration strategies remains a major bottleneck. Current systems face a fundamental dilemma: fixed strategies fail to adapt as search spaces scale, while online policy optimization requires navigating vast meta-search spaces under delayed and expensive feedback over long-horizon rollouts. We introduce Dream-RSI, a framework for scalable and recursively self-improving exploration. A lightweight orchestration layer makes exploration explicit and programmable while leaving the underlying coding agent unchanged. Our key insight is that accumulated discovery history can serve as a replay simulator over the realized search space. By performing dreaming in the replay simulator constructed from historical discovery trees, Dream-RSI secures immediate, low-cost off-policy feedback to evaluate and refine exploration policies without invoking repetitive, expensive online evaluations. The improved policy is subsequently redeployed online to drive further discovery, continuously expanding the simulator pool in a self-improving loop. Across algorithm engineering, mathematical optimization, and GPU kernel engineering, Dream-RSI achieves competitive or improved discovery quality while substantially reducing discovery cost in several settings.

Community

To recursively self-improve, agents need to dream. History can be the world they dream in.

More Details: https://dream-rsi.com/

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.14858
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.14858 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.14858 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.14858 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers