Hugging Face Daily Papers · · 6 min read

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evolution proceeds in two modes. In recursive free evolution, improvement is itself a task, and completing one evolution cycle can schedule the next. In experience-driven core evolution, ordinary work and social interaction expose bugs, rough edges, and inefficient context construction that lead to reviewed structural changes.<br>On Terminal-Bench 2.1, an Opus 5 run scores 86.74%, the best result reported on the benchmark. On OSWorld-Verified, an Opus 5 run reaches 90.69%, exceeding the best previously reported score. A five-rollout CL-Bench campaign achieves a normalized reward of 0.23, setting a new state of the art.<br>Hope is the longest-running publicly documented Ouroboros deployment. It is a 161-day living agent experiment in free evolution under governed human communication across seven surfaces. Human interaction surfaces faults and generates proposals, but the agent decides which changes to pursue. Because a self-developing agent may rewrite its own code and select new model APIs, operational safety becomes a primary design problem: guardrails must remain authoritative under evolutionary and public social pressure. Benchmark campaigns use frozen system snapshots, while Hope continues live evolution on a separate lineage.</p>\n","updatedAt":"2026-08-11T04:13:16.538Z","author":{"_id":"6172aaeec8e66e2aa84c06b9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6172aaeec8e66e2aa84c06b9/ZdRZSp3P1SU6CIDbvQwkv.jpeg","fullname":"Anton Razzhigaev","name":"razzant","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":24,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8829264640808105},"editors":["razzant"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6172aaeec8e66e2aa84c06b9/ZdRZSp3P1SU6CIDbvQwkv.jpeg"],"reactions":[],"isReport":false}},{"id":"6a7aedd9cbec0637f0abf540","author":{"_id":"6172aaeec8e66e2aa84c06b9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6172aaeec8e66e2aa84c06b9/ZdRZSp3P1SU6CIDbvQwkv.jpeg","fullname":"Anton Razzhigaev","name":"razzant","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":24,"isUserFollowing":false},"createdAt":"2026-08-11T09:39:37.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"Hi @mfielding92, @yaoyuzhao, @Sam1818AZ, and @dram023. You maintain collections about self-evolving agents and agent harnesses, so this paper may fit. Ouroboros is a coding-agent harness whose tools, prompts, context assembly, and core implementation evolve through reviewed commits. If it fits your scope, would you consider adding arXiv:2608.08311? The code and public benchmark evidence are linked above. Thanks for curating these collections.","html":"<p>Hi <span class=\"SVELTE_PARTIAL_HYDRATER contents\" data-target=\"UserMention\" data-props=\"{&quot;user&quot;:&quot;mfielding92&quot;}\"><span class=\"inline-block\"><span class=\"contents\"><a href=\"/mfielding92\">@<span class=\"underline\">mfielding92</span></a></span> </span></span>, <span class=\"SVELTE_PARTIAL_HYDRATER contents\" data-target=\"UserMention\" data-props=\"{&quot;user&quot;:&quot;yaoyuzhao&quot;}\"><span class=\"inline-block\"><span class=\"contents\"><a href=\"/yaoyuzhao\">@<span class=\"underline\">yaoyuzhao</span></a></span> </span></span>, <span class=\"SVELTE_PARTIAL_HYDRATER contents\" data-target=\"UserMention\" data-props=\"{&quot;user&quot;:&quot;Sam1818AZ&quot;}\"><span class=\"inline-block\"><span class=\"contents\"><a href=\"/Sam1818AZ\">@<span class=\"underline\">Sam1818AZ</span></a></span> </span></span>, and <span class=\"SVELTE_PARTIAL_HYDRATER contents\" data-target=\"UserMention\" data-props=\"{&quot;user&quot;:&quot;dram023&quot;}\"><span class=\"inline-block\"><span class=\"contents\"><a href=\"/dram023\">@<span class=\"underline\">dram023</span></a></span> </span></span>. You maintain collections about self-evolving agents and agent harnesses, so this paper may fit. Ouroboros is a coding-agent harness whose tools, prompts, context assembly, and core implementation evolve through reviewed commits. If it fits your scope, would you consider adding arXiv:2608.08311? The code and public benchmark evidence are linked above. Thanks for curating these collections.</p>\n","updatedAt":"2026-08-11T09:39:37.932Z","author":{"_id":"6172aaeec8e66e2aa84c06b9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6172aaeec8e66e2aa84c06b9/ZdRZSp3P1SU6CIDbvQwkv.jpeg","fullname":"Anton Razzhigaev","name":"razzant","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":24,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8820592761039734},"editors":["razzant"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6172aaeec8e66e2aa84c06b9/ZdRZSp3P1SU6CIDbvQwkv.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.08311","authors":[{"_id":"6a7a9fea019ce76dc7b3aacf","user":{"_id":"6172aaeec8e66e2aa84c06b9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6172aaeec8e66e2aa84c06b9/ZdRZSp3P1SU6CIDbvQwkv.jpeg","isPro":false,"fullname":"Anton Razzhigaev","user":"razzant","type":"user","name":"razzant"},"name":"Anton Razzhigaev","status":"claimed_verified","statusLastChangedAt":"2026-08-11T08:45:04.508Z","hidden":false},{"_id":"6a7a9fea019ce76dc7b3aad0","name":"Andrei Gritsaev","hidden":false},{"_id":"6a7a9fea019ce76dc7b3aad1","name":"Andrei Kaznacheev","hidden":false},{"_id":"6a7a9fea019ce76dc7b3aad2","name":"Nikita Dragunov","hidden":false},{"_id":"6a7a9fea019ce76dc7b3aad3","name":"Roman Yampolskiy","hidden":false},{"_id":"6a7a9fea019ce76dc7b3aad4","name":"Andrei Kuznetsov","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6172aaeec8e66e2aa84c06b9/uLvkxBoMmC9OrSJuho-m8.png","https://cdn-uploads.huggingface.co/production/uploads/6172aaeec8e66e2aa84c06b9/qW9VmSHez6N4hMZ93Jntv.png"],"publishedAt":"2026-08-08T00:00:00.000Z","submittedOnDailyAt":"2026-08-11T00:00:00.000Z","title":"Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution","submittedOnDailyBy":{"_id":"6172aaeec8e66e2aa84c06b9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6172aaeec8e66e2aa84c06b9/ZdRZSp3P1SU6CIDbvQwkv.jpeg","isPro":false,"fullname":"Anton Razzhigaev","user":"razzant","type":"user","name":"razzant"},"summary":"We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evolution proceeds in two modes. In recursive free evolution, improvement is itself a task, and completing one evolution cycle can schedule the next. In experience-driven core evolution, ordinary work and social interaction expose bugs, rough edges, and inefficient context construction that lead to reviewed structural changes.\n On Terminal-Bench 2.1, an Opus 5 run scores 86.74%, the best result reported on the benchmark. On OSWorld-Verified, an Opus 5 run reaches 90.69%, exceeding the best previously reported score. A five-rollout CL-Bench campaign achieves a normalized reward of 0.2301, setting a new state of the art.\n Hope is the longest-running publicly documented Ouroboros deployment. It is a 161-day living agent experiment in free evolution under governed human communication across seven surfaces. Human interaction surfaces faults and generates proposals, but the agent decides which changes to pursue. Because a self-developing agent may rewrite its own code and select new model APIs, operational safety becomes a primary design problem: guardrails must remain authoritative under evolutionary and public social pressure. Benchmark campaigns use frozen system snapshots, while Hope continues live evolution on a separate lineage.","upvotes":65,"discussionId":"6a7a9fea019ce76dc7b3aad5","projectPage":"https://ouroboros-agent.ai","githubRepo":"https://github.com/razzant/ouroboros","githubRepoAddedBy":"user","githubStars":1089},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6172aaeec8e66e2aa84c06b9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6172aaeec8e66e2aa84c06b9/ZdRZSp3P1SU6CIDbvQwkv.jpeg","isPro":false,"fullname":"Anton Razzhigaev","user":"razzant","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"691ddfc82af40fd526b31481","avatarUrl":"/avatars/d34835bd4c9c4a0fe48fa0eee5be67c0.svg","isPro":false,"fullname":"Andrei Gritsaev","user":"PUFL","type":"user"},{"_id":"67d5a331eab66ce9cb01bae4","avatarUrl":"/avatars/3ed437c874889e8a6db66c3ef88a60c0.svg","isPro":false,"fullname":"DMITRII ZHEMCHUZHNIKOV","user":"zhemchuzhnikov","type":"user"},{"_id":"649a98378362132404508317","avatarUrl":"/avatars/377fbdfbf95fbacac975b33a7bda02e2.svg","isPro":true,"fullname":"Andrei Filatov","user":"anvilarth","type":"user"},{"_id":"67bcb1012906865678a11f91","avatarUrl":"/avatars/80fb0cc24f0d16c4740f9115b680df0f.svg","isPro":false,"fullname":"Vladimir Korviakov","user":"korviakov","type":"user"},{"_id":"643984dceb7c5616ef3f5d54","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/643984dceb7c5616ef3f5d54/10JRkblrRIEVci6UJwvPz.jpeg","isPro":false,"fullname":"Andrey Kuznetsov","user":"kuznetsoffandrey","type":"user"},{"_id":"6310ff34bc152fa3e810c186","avatarUrl":"/avatars/bfd63bcd81548283f5e496e3693bf143.svg","isPro":true,"fullname":"Elizaveta Goncharova","user":"Elizaveta","type":"user"},{"_id":"623a3b20fa4890c51b04cba7","avatarUrl":"/avatars/ebc71284fab04d506471aa5b1792157b.svg","isPro":false,"fullname":"Andrey Moskalenko","user":"ANDRYHA","type":"user"},{"_id":"6616719945336ca7746eaa38","avatarUrl":"/avatars/ac77ebda8507d75376973144263beb83.svg","isPro":false,"fullname":"Dmitrii Mikhailov","user":"Botsman11","type":"user"},{"_id":"65c4eb7cd1dcbd30d86febec","avatarUrl":"/avatars/001c8f02e8ce794b2c21883628b2da72.svg","isPro":false,"fullname":"free-bit","user":"free-bit","type":"user"},{"_id":"6603240583986596681300e1","avatarUrl":"/avatars/58aab44ca5191327cb70282a3c01ae07.svg","isPro":false,"fullname":"Igor Gromov ","user":"Transformator","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.08311.md","query":{}}">
Papers
arxiv:2608.08311

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution

Published on Aug 8
· Submitted by
Anton Razzhigaev
on Aug 11
Authors:

Abstract

We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evolution proceeds in two modes. In recursive free evolution, improvement is itself a task, and completing one evolution cycle can schedule the next. In experience-driven core evolution, ordinary work and social interaction expose bugs, rough edges, and inefficient context construction that lead to reviewed structural changes. On Terminal-Bench 2.1, an Opus 5 run scores 86.74%, the best result reported on the benchmark. On OSWorld-Verified, an Opus 5 run reaches 90.69%, exceeding the best previously reported score. A five-rollout CL-Bench campaign achieves a normalized reward of 0.2301, setting a new state of the art. Hope is the longest-running publicly documented Ouroboros deployment. It is a 161-day living agent experiment in free evolution under governed human communication across seven surfaces. Human interaction surfaces faults and generates proposals, but the agent decides which changes to pursue. Because a self-developing agent may rewrite its own code and select new model APIs, operational safety becomes a primary design problem: guardrails must remain authoritative under evolutionary and public social pressure. Benchmark campaigns use frozen system snapshots, while Hope continues live evolution on a separate lineage.

Community

Paper author Paper submitter about 15 hours ago

We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evolution proceeds in two modes. In recursive free evolution, improvement is itself a task, and completing one evolution cycle can schedule the next. In experience-driven core evolution, ordinary work and social interaction expose bugs, rough edges, and inefficient context construction that lead to reviewed structural changes.
On Terminal-Bench 2.1, an Opus 5 run scores 86.74%, the best result reported on the benchmark. On OSWorld-Verified, an Opus 5 run reaches 90.69%, exceeding the best previously reported score. A five-rollout CL-Bench campaign achieves a normalized reward of 0.23, setting a new state of the art.
Hope is the longest-running publicly documented Ouroboros deployment. It is a 161-day living agent experiment in free evolution under governed human communication across seven surfaces. Human interaction surfaces faults and generates proposals, but the agent decides which changes to pursue. Because a self-developing agent may rewrite its own code and select new model APIs, operational safety becomes a primary design problem: guardrails must remain authoritative under evolutionary and public social pressure. Benchmark campaigns use frozen system snapshots, while Hope continues live evolution on a separate lineage.

Paper author Paper submitter about 9 hours ago

Hi @mfielding92 , @yaoyuzhao , @Sam1818AZ , and @dram023 . You maintain collections about self-evolving agents and agent harnesses, so this paper may fit. Ouroboros is a coding-agent harness whose tools, prompts, context assembly, and core implementation evolve through reviewed commits. If it fits your scope, would you consider adding arXiv:2608.08311? The code and public benchmark evidence are linked above. Thanks for curating these collections.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.08311
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.08311 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers