We’re excited to share <strong>RSIAgent: Autonomous Exploration for Recursive Self-Improvement in New Environments</strong>!</p>\n<p>Instead of continuing to scale model parameters, we explore a different path: <strong>Scaling Experience</strong>. RSIAgent enables agents to autonomously decide what to explore, execute tasks, verify outcomes, and consolidate stable action–condition–outcome relationships into reusable memory, all while keeping the underlying model weights frozen.</p>\n<p>RSIAgent coordinates curriculum, actor, and verifier agents through a <strong>Broad-to-Deep</strong> exploration strategy: broad exploration builds coverage across tools, workflows, and failure modes, while deep exploration targets difficult cases, hidden constraints, and boundary conditions.</p>\n<p>When applied to Kimi-K3 and GLM-5.3, RSIAgent achieves:</p>\n<ul>\n<li><strong>78.98% Partial Score</strong> on OSWorld 2.0 (0808 offline), compared with <strong>72.60%</strong> for GPT-6 Astra</li>\n<li><strong>84.82%</strong> on Agents’ Last Exam (Near-term), compared with <strong>82.26%</strong> for GPT-6 Astra</li>\n</ul>\n<p>These results suggest that agents can continue improving by autonomously acquiring, verifying, and reusing their own experience—even without updating model parameters.</p>\n<ul>\n<li>Blog: <a href=\"https://aetherlabs.ai/articles/rsiagent-autonomous-exploration-for-recursive-self-improvement\" rel=\"nofollow\">https://aetherlabs.ai/articles/rsiagent-autonomous-exploration-for-recursive-self-improvement</a></li>\n<li>Code: <a href=\"https://github.com/AetherLabsAI/RSIAgent\" rel=\"nofollow\">https://github.com/AetherLabsAI/RSIAgent</a></li>\n<li>Project page: <a href=\"https://aetherlabsai.github.io/RSIAgent/\" rel=\"nofollow\">https://aetherlabsai.github.io/RSIAgent/</a></li>\n<li>Paper: <a href=\"https://arxiv.org/abs/2609.15364\" rel=\"nofollow\">https://arxiv.org/abs/2609.15364</a></li>\n</ul>\n<p>We’d love to hear your feedback and are happy to answer any questions!</p>\n","updatedAt":"2026-09-15T05:26:38.428Z","author":{"_id":"691d4123f8321286ee15a131","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/691d4123f8321286ee15a131/0Hfhodl-2KNYCH1RClWvR.jpeg","fullname":"shicheng","name":"Shichengf","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8224727511405945},"editors":["Shichengf"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/691d4123f8321286ee15a131/0Hfhodl-2KNYCH1RClWvR.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.15364","authors":[{"_id":"6aa8d4675dd4cb9b4cc0292e","name":"Sibo Zhu","hidden":false},{"_id":"6aa8d4675dd4cb9b4cc0292f","name":"Shicheng Fan","hidden":false},{"_id":"6aa8d4675dd4cb9b4cc02930","name":"Xinyue Wang","hidden":false},{"_id":"6aa8d4675dd4cb9b4cc02931","name":"Wenyi Wu","hidden":false},{"_id":"6aa8d4675dd4cb9b4cc02932","name":"Kun Zhou","hidden":false},{"_id":"6aa8d4675dd4cb9b4cc02933","name":"Biwei Huang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/691d4123f8321286ee15a131/UejP-CxK7t4b1hMxR79Sh.png","https://cdn-uploads.huggingface.co/production/uploads/691d4123f8321286ee15a131/B1koMQVD3NcsONFXA6Tml.mp4"],"publishedAt":"2026-09-14T00:00:00.000Z","submittedOnDailyAt":"2026-09-15T00:00:00.000Z","title":"RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments","submittedOnDailyBy":{"_id":"691d4123f8321286ee15a131","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/691d4123f8321286ee15a131/0Hfhodl-2KNYCH1RClWvR.jpeg","isPro":false,"fullname":"shicheng","user":"Shichengf","type":"user","name":"Shichengf"},"summary":"Digital agents must often adapt to new environments whose interfaces, tools, and failure modes are not fully captured by pretrained models. We introduce RSIAgent, a training-free multi-agent framework for recursive self-improvement through autonomous memory construction. RSIAgent coordinates curriculum, actor, and verifier agents to continually explore the environment, validate outcomes, and retain environment-specific knowledge, including reusable causal relationships between actions, conditions, and consequences. It further adopts a broad-then-deep exploration strategy, combining parallel broad recursive self-exploration for discovering diverse environment structures with focused deep self-exploration for uncovering hard cases, hidden constraints, boundary conditions, and previously unknown causal dependencies. The resulting memory is frozen and can be directly reused for downstream tasks without updating model parameters. Experiments on OSWorld-v2 and Agent's Last Exam show that RSIAgent substantially improves strong open-source models, enabling Kimi-K3 and GLM-5.3 to outperform frontier closed-source models including GPT-6.","upvotes":13,"discussionId":"6aa8d4675dd4cb9b4cc02934","projectPage":"https://aetherlabsai.github.io/RSIAgent/","githubRepo":"https://github.com/AetherLabsAI/RSIAgent","githubRepoAddedBy":"user","ai_summary":"RSIAgent is a training-free multi-agent framework that enables recursive self-improvement via autonomous memory construction and broad-then-deep exploration to adapt digital agents to new environments.","ai_keywords":["RSIAgent","recursive self-improvement","multi-agent framework","autonomous memory construction","causal relationships","broad-then-deep exploration","recursive self-exploration"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":50,"organization":{"_id":"6aa0ef5950271eb539c5db53","name":"AetherLabs-AI","fullname":"AetherLabs-AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66f4fcbc29b10ae4c990a2e0/48L5Xsw_rF8WKtMtWvlMm.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"691d4123f8321286ee15a131","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/691d4123f8321286ee15a131/0Hfhodl-2KNYCH1RClWvR.jpeg","isPro":false,"fullname":"shicheng","user":"Shichengf","type":"user"},{"_id":"66f4fcbc29b10ae4c990a2e0","avatarUrl":"/avatars/5d6df3aa5792a031c28274d428c46d84.svg","isPro":false,"fullname":"Kun Zhou","user":"FrancisKunZhou","type":"user"},{"_id":"674abd2829a3bb873fc57cc9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ydrol4HKoYQvO6sABPfPp.png","isPro":false,"fullname":"Nick Sibo Zhu","user":"Nick0907","type":"user"},{"_id":"686263d4cfa12e6d9d95f5db","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/686263d4cfa12e6d9d95f5db/UFmER9oV14v7kDDCW_9V2.png","isPro":false,"fullname":"Ziming Xu","user":"Thrcle","type":"user"},{"_id":"652344742470ec6d1aa5d57b","avatarUrl":"/avatars/3d9c4fda428b10cefb4e13a086ce14fb.svg","isPro":true,"fullname":"Xinyue Wang","user":"XinyueWangg","type":"user"},{"_id":"679923cba2b48413dde2b3bb","avatarUrl":"/avatars/28c3f05649a9e412c9c59491f08a0954.svg","isPro":false,"fullname":"Yufan Wei","user":"yufanwei","type":"user"},{"_id":"67d65a1eb20ded10ed9b9164","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/iAFQ6frGvE6Ock6tGuJBI.png","isPro":false,"fullname":"jijun","user":"xujijun","type":"user"},{"_id":"672aac577ec01ef4af3367d5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/uNgXH3uMBFt-hI3cCeQ-5.jpeg","isPro":false,"fullname":"Wenyi Wu","user":"WenyiWU0111","type":"user"},{"_id":"6aa8d9cf922e0db0cbd4db16","avatarUrl":"/avatars/d6464ed7712fb79f10e67741e3828666.svg","isPro":false,"fullname":"Xuan Chen","user":"Mikrokosmoser","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"67a5c0ca39ece2ce518c4825","avatarUrl":"/avatars/e1c7d5865671177fbcd7438473d49d85.svg","isPro":false,"fullname":"Zijian Zhou","user":"Kailai1104","type":"user"},{"_id":"6a2905845160513fef821e8c","avatarUrl":"/avatars/e2657de0cddabffdfd4a6d744770de38.svg","isPro":false,"fullname":"Nelson Ashley","user":"Foggywhale","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6aa0ef5950271eb539c5db53","name":"AetherLabs-AI","fullname":"AetherLabs-AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66f4fcbc29b10ae4c990a2e0/48L5Xsw_rF8WKtMtWvlMm.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.15364.md","query":{}}">
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
Abstract
RSIAgent is a training-free multi-agent framework that enables recursive self-improvement via autonomous memory construction and broad-then-deep exploration to adapt digital agents to new environments.
Digital agents must often adapt to new environments whose interfaces, tools, and failure modes are not fully captured by pretrained models. We introduce RSIAgent, a training-free multi-agent framework for recursive self-improvement through autonomous memory construction. RSIAgent coordinates curriculum, actor, and verifier agents to continually explore the environment, validate outcomes, and retain environment-specific knowledge, including reusable causal relationships between actions, conditions, and consequences. It further adopts a broad-then-deep exploration strategy, combining parallel broad recursive self-exploration for discovering diverse environment structures with focused deep self-exploration for uncovering hard cases, hidden constraints, boundary conditions, and previously unknown causal dependencies. The resulting memory is frozen and can be directly reused for downstream tasks without updating model parameters. Experiments on OSWorld-v2 and Agent's Last Exam show that RSIAgent substantially improves strong open-source models, enabling Kimi-K3 and GLM-5.3 to outperform frontier closed-source models including GPT-6.
Community
We’re excited to share RSIAgent: Autonomous Exploration for Recursive Self-Improvement in New Environments!
Instead of continuing to scale model parameters, we explore a different path: Scaling Experience. RSIAgent enables agents to autonomously decide what to explore, execute tasks, verify outcomes, and consolidate stable action–condition–outcome relationships into reusable memory, all while keeping the underlying model weights frozen.
RSIAgent coordinates curriculum, actor, and verifier agents through a Broad-to-Deep exploration strategy: broad exploration builds coverage across tools, workflows, and failure modes, while deep exploration targets difficult cases, hidden constraints, and boundary conditions.
When applied to Kimi-K3 and GLM-5.3, RSIAgent achieves:
- 78.98% Partial Score on OSWorld 2.0 (0808 offline), compared with 72.60% for GPT-6 Astra
- 84.82% on Agents’ Last Exam (Near-term), compared with 82.26% for GPT-6 Astra
These results suggest that agents can continue improving by autonomously acquiring, verifying, and reusing their own experience—even without updating model parameters.
We’d love to hear your feedback and are happy to answer any questions!
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.15364 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.15364 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.15364 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.