Hugging Face Daily Papers · · 6 min read

FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

This paper addresses a practical bottleneck in agent systems: useful workflows discovered during inference are usually discarded, while existing skill libraries are often static or built offline. FlowEvo closes this loop by turning verified successful workflows into persistent executable skills. These skills can be reused directly or guide the creation of new workflows, while harmful skills are identified and suppressed. Using GPT 4o mini, FlowEvo outperforms eight baselines across five benchmarks. On ALFWorld, it achieves an 85.6% success rate, 26.4 points above the strongest baseline, while using roughly one third of the tokens. Results across 10 base models further demonstrate its potential for building agents that improve continuously without additional training.</p>\n","updatedAt":"2026-08-21T23:09:40.396Z","author":{"_id":"66739484a660dfbb2643eb3d","avatarUrl":"/avatars/c9a6c6a1e295e448dacfccb1a2860e8a.svg","fullname":"Leo Y","name":"LeoYML","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9082251191139221},"editors":["LeoYML"],"editorAvatarUrls":["/avatars/c9a6c6a1e295e448dacfccb1a2860e8a.svg"],"reactions":[],"isReport":false}},{"id":"6a88fc5b6c1ecc2d8cdcab58","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false},"createdAt":"2026-08-22T01:33:15.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills](https://huggingface.co/papers/2606.26669) (2026)\n* [Living-Harness Is an Interactive-Agent Evolver](https://huggingface.co/papers/2607.26598) (2026)\n* [Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents](https://huggingface.co/papers/2608.15071) (2026)\n* [Self-Evolving Coding Agents](https://huggingface.co/papers/2608.03392) (2026)\n* [MemoHarness: Agent Harnesses That Learn from Experience](https://huggingface.co/papers/2607.14159) (2026)\n* [Object-Centric Environment Modeling for Agentic Tasks](https://huggingface.co/papers/2607.02846) (2026)\n* [SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries](https://huggingface.co/papers/2608.05604) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2606.26669\">SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.26598\">Living-Harness Is an Interactive-Agent Evolver</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.15071\">Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.03392\">Self-Evolving Coding Agents</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.14159\">MemoHarness: Agent Harnesses That Learn from Experience</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.02846\">Object-Centric Environment Modeling for Agentic Tasks</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.05604\">SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-08-22T01:33:15.752Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7436361908912659},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.21596","authors":[{"_id":"6a88d9023d26296ea30913c6","name":"Zeyu Ren","hidden":false},{"_id":"6a88d9023d26296ea30913c7","name":"Ling Yue","hidden":false},{"_id":"6a88d9023d26296ea30913c8","name":"Ran Li","hidden":false},{"_id":"6a88d9023d26296ea30913c9","name":"Yishu Wang","hidden":false},{"_id":"6a88d9023d26296ea30913ca","name":"Shengxiang Xu","hidden":false},{"_id":"6a88d9023d26296ea30913cb","name":"Hanmo Liu","hidden":false},{"_id":"6a88d9023d26296ea30913cc","name":"Shaowu Pan","hidden":false},{"_id":"6a88d9023d26296ea30913cd","name":"Shimin Di","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/66739484a660dfbb2643eb3d/oCWfxqHgnJhnOUMsB0KeY.png"],"publishedAt":"2026-08-20T00:00:00.000Z","submittedOnDailyAt":"2026-08-21T00:00:00.000Z","title":"FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills","submittedOnDailyBy":{"_id":"66739484a660dfbb2643eb3d","avatarUrl":"/avatars/c9a6c6a1e295e448dacfccb1a2860e8a.svg","isPro":false,"fullname":"Leo Y","user":"LeoYML","type":"user","name":"LeoYML"},"summary":"Large language model agents can adapt to complex tasks by constructing workflows at inference time, but procedures discovered in one episode are usually discarded after execution. Existing skill libraries provide reusable executable routines, but are typically assembled offline and do not grow from the agent's own workflows. We introduce FlowEvo, a training-free framework in which workflows and skills co-evolve at inference time. FlowEvo compiles successful workflows into callable skills, stores them in a persistent bank, and uses retrieved skills either through direct execution or as context for constructing new workflows. It also tracks each skill's downstream utility and suppresses skills that cause negative transfer. Using a shared GPT-4o-mini backbone, FlowEvo achieves the highest accuracy among 8 baselines on the full standard splits of ALFWorld, HumanEval, MBPP, GSM8K, and MATH-500. On ALFWorld, it reaches 85.6%, 26.4 points above the strongest baseline, while using roughly one third as many tokens. Across 10 base models spanning 7B to 671B parameters, FlowEvo outperforms ExpeL in 49 of 50 model-dataset comparisons. Code is available at https://github.com/DEFENSE-SEU/FlowEvo.","upvotes":2,"discussionId":"6a88d9033d26296ea30913ce","githubRepo":"https://github.com/DEFENSE-SEU/FlowEvo","githubRepoAddedBy":"user","ai_summary":"FlowEvo enables large language model agents to co-evolve reusable skills and workflows during inference, improving accuracy and efficiency across diverse benchmarks.","ai_keywords":["FlowEvo","skill libraries","workflow construction","inference-time adaptation","negative transfer","persistent skill bank","cross-model evaluation"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":7,"organization":{"_id":"654366fbeb97ea6fae4661e4","name":"SoutheastU","fullname":"Southeast Univeristy","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68590547473738a3ca2c203c/l4z89rV91d8-t22AgCOEA.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66739484a660dfbb2643eb3d","avatarUrl":"/avatars/c9a6c6a1e295e448dacfccb1a2860e8a.svg","isPro":false,"fullname":"Leo Y","user":"LeoYML","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"654366fbeb97ea6fae4661e4","name":"SoutheastU","fullname":"Southeast Univeristy","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68590547473738a3ca2c203c/l4z89rV91d8-t22AgCOEA.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.21596.md","query":{}}">
Papers
arxiv:2607.21596

FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills

Published on Aug 20
· Submitted by
Leo Y
on Aug 21
Authors:
,

Abstract

FlowEvo enables large language model agents to co-evolve reusable skills and workflows during inference, improving accuracy and efficiency across diverse benchmarks.

Large language model agents can adapt to complex tasks by constructing workflows at inference time, but procedures discovered in one episode are usually discarded after execution. Existing skill libraries provide reusable executable routines, but are typically assembled offline and do not grow from the agent's own workflows. We introduce FlowEvo, a training-free framework in which workflows and skills co-evolve at inference time. FlowEvo compiles successful workflows into callable skills, stores them in a persistent bank, and uses retrieved skills either through direct execution or as context for constructing new workflows. It also tracks each skill's downstream utility and suppresses skills that cause negative transfer. Using a shared GPT-4o-mini backbone, FlowEvo achieves the highest accuracy among 8 baselines on the full standard splits of ALFWorld, HumanEval, MBPP, GSM8K, and MATH-500. On ALFWorld, it reaches 85.6%, 26.4 points above the strongest baseline, while using roughly one third as many tokens. Across 10 base models spanning 7B to 671B parameters, FlowEvo outperforms ExpeL in 49 of 50 model-dataset comparisons. Code is available at https://github.com/DEFENSE-SEU/FlowEvo.

Community

Paper submitter about 3 hours ago

This paper addresses a practical bottleneck in agent systems: useful workflows discovered during inference are usually discarded, while existing skill libraries are often static or built offline. FlowEvo closes this loop by turning verified successful workflows into persistent executable skills. These skills can be reused directly or guide the creation of new workflows, while harmful skills are identified and suppressed. Using GPT 4o mini, FlowEvo outperforms eight baselines across five benchmarks. On ALFWorld, it achieves an 85.6% success rate, 26.4 points above the strongest baseline, while using roughly one third of the tokens. Results across 10 base models further demonstrate its potential for building agents that improve continuously without additional training.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.21596
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.21596 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.21596 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.21596 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers