We release GDPevo, the first benchmark for evaluating agent self-evolution on GDP-related enterprise tasks.</p>\n<p>We also open-source the fully automated pipeline that generates evolution-native benchmark data, providing a practical response to contamination.</p>\n<p>We found the best evolved agents remain far below a fully informed oracle ceiling of 91.6%, indicating that the self-evolution ability of current agents remains far from fully realized.</p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/642fef28a043f0ac7defa8a9/LrGPJKQmneZNwDIZh7XM7.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/642fef28a043f0ac7defa8a9/LrGPJKQmneZNwDIZh7XM7.png\" alt=\"image\"></a></p>\n","updatedAt":"2026-08-06T07:01:02.126Z","author":{"_id":"642fef28a043f0ac7defa8a9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642fef28a043f0ac7defa8a9/RwOEkuj3fOnOA54tGR7Ea.png","fullname":"Yaowei Zheng","name":"hiyouga","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3700,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8097880482673645},"editors":["hiyouga"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/642fef28a043f0ac7defa8a9/RwOEkuj3fOnOA54tGR7Ea.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.03764","authors":[{"_id":"6a742fb1c5e410d076869bf8","name":"Leijun Zhou","hidden":false},{"_id":"6a742fb1c5e410d076869bf9","user":{"_id":"6a47795e7184ebf484046368","avatarUrl":"/avatars/a16ffead1d163ca92cae2f7ffa07116c.svg","isPro":false,"fullname":"liuzhihao","user":"zhihao43112","type":"user","name":"zhihao43112"},"name":"Zhihao Liu","status":"claimed_verified","statusLastChangedAt":"2026-08-06T08:45:05.544Z","hidden":false},{"_id":"6a742fb1c5e410d076869bfa","name":"Xiang Qu","hidden":false},{"_id":"6a742fb1c5e410d076869bfb","name":"Chenxu Liu","hidden":false},{"_id":"6a742fb1c5e410d076869bfc","name":"Yifei Liu","hidden":false},{"_id":"6a742fb1c5e410d076869bfd","name":"Yanke Yu","hidden":false},{"_id":"6a742fb1c5e410d076869bfe","name":"Jingzhe Xu","hidden":false},{"_id":"6a742fb1c5e410d076869bff","name":"Xuejun Wu","hidden":false},{"_id":"6a742fb1c5e410d076869c00","name":"Buyue Qian","hidden":false},{"_id":"6a742fb1c5e410d076869c01","name":"Xi Chen","hidden":false},{"_id":"6a742fb1c5e410d076869c02","user":{"_id":"642fef28a043f0ac7defa8a9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642fef28a043f0ac7defa8a9/RwOEkuj3fOnOA54tGR7Ea.png","isPro":false,"fullname":"Yaowei Zheng","user":"hiyouga","type":"user","name":"hiyouga"},"name":"Yaowei Zheng","status":"claimed_verified","statusLastChangedAt":"2026-08-06T08:45:05.538Z","hidden":false},{"_id":"6a742fb1c5e410d076869c03","name":"Junhao Hu","hidden":false}],"publishedAt":"2026-08-04T00:00:00.000Z","submittedOnDailyAt":"2026-08-06T00:00:00.000Z","title":"GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks","submittedOnDailyBy":{"_id":"642fef28a043f0ac7defa8a9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642fef28a043f0ac7defa8a9/RwOEkuj3fOnOA54tGR7Ea.png","isPro":false,"fullname":"Yaowei Zheng","user":"hiyouga","type":"user","name":"hiyouga"},"summary":"Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-evolution is difficult: existing benchmarks provide limited coverage of economically valuable task domains, do not always design training and test tasks such that test-time gains can be attributed to training experience, and remain vulnerable to data contamination. We present GDPevo, an evolution-native benchmark grounded in GDP-related enterprise workflows, together with the fully automated data pipeline that generates it. Its core mechanism, rule hybridization, decomposes each enterprise workflow into atomic business rules, distributes subsets of these rules across training tasks, and recombines them in held-out test tasks so that test-time gains are attributable. GDPevo spans CRM, ERP, finance, healthcare, legal, and data-centric workflows. Its V1 release contains 120 tasks in 12 groups, with five training and five held-out test tasks per group. Full automation enables the pipeline to expand the suite to 240 tasks in 24 groups (V2) within two days, providing a practical response to contamination. Using GDPevo, we evaluate four agents, each comprising a harness and a model, under four supervision types. Self-evolution consistently improves held-out accuracy by up to 16.44 percentage points. But the best evolved agents remain far below the fully informed oracle ceiling of 91.6%, indicating that the self-evolution ability of current agents remains far from fully realized. We publicly release the pipeline, benchmark, and full evaluation results at https://github.com/Prism-Shadow/GDPevo.","upvotes":19,"discussionId":"6a742fb1c5e410d076869c04","projectPage":"https://prism-shadow.github.io/GDPevo/","githubRepo":"https://github.com/Prism-Shadow/GDPevo","githubRepoAddedBy":"user","githubStars":44,"organization":{"_id":"694abfb70e71be39f26489b8","name":"PrismShadow","fullname":"Prism Shadow","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/642fef28a043f0ac7defa8a9/JpqbmwCufkE6PGKFDSiQX.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"642fef28a043f0ac7defa8a9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642fef28a043f0ac7defa8a9/RwOEkuj3fOnOA54tGR7Ea.png","isPro":false,"fullname":"Yaowei Zheng","user":"hiyouga","type":"user"},{"_id":"67a7cd4247fc008de2d8ce97","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67a7cd4247fc008de2d8ce97/pMYK0raKqPr_oGNV5nHSP.jpeg","isPro":false,"fullname":"LeijunZhou","user":"Saigyouji-Yuyuko1000","type":"user"},{"_id":"6775ab3df9a75d4e36c7d33a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/vkcfTmsbZq1NLL344Zecb.png","isPro":false,"fullname":"Xie","user":"YouhaiXie","type":"user"},{"_id":"66aa0ef97cda19fabe0f9c98","avatarUrl":"/avatars/e00c61ab0512dbf8f3ac371c72a40949.svg","isPro":false,"fullname":"Frank","user":"rky01","type":"user"},{"_id":"6809eb789111f049b281f0da","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/t9sYmKBAYqBTpLIfK-_J4.png","isPro":false,"fullname":"laodouuu","user":"douhengjin","type":"user"},{"_id":"6a7432372e0f9ebbddaf438f","avatarUrl":"/avatars/e3562402ebb4d9d4077142a25d62f7d4.svg","isPro":false,"fullname":"Zhengyuan Hu","user":"EdisonHu","type":"user"},{"_id":"681f312530b0f8095876351d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/Vuswaox2_-dZSuF4Hxzph.png","isPro":false,"fullname":"x","user":"j1n9zhe","type":"user"},{"_id":"6a47795e7184ebf484046368","avatarUrl":"/avatars/a16ffead1d163ca92cae2f7ffa07116c.svg","isPro":false,"fullname":"liuzhihao","user":"zhihao43112","type":"user"},{"_id":"66a1b7635f58258df73318a0","avatarUrl":"/avatars/f62f68ecfb7ae2d373459bb2ca7eeacc.svg","isPro":false,"fullname":"Junhao Hu","user":"DerekHJH","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6469d2d58dfd6ff79b68836a","avatarUrl":"/avatars/486e0c192559c59a8e863e44010ec27f.svg","isPro":false,"fullname":"Lewis","user":"CompileError","type":"user"},{"_id":"65701f501b048a9b25efbbce","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/_jGQsNSMbXxA6dR7Wx8yL.png","isPro":false,"fullname":"Yifei Liu","user":"NephrenCake","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"694abfb70e71be39f26489b8","name":"PrismShadow","fullname":"Prism Shadow","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/642fef28a043f0ac7defa8a9/JpqbmwCufkE6PGKFDSiQX.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.03764.md","query":{}}">
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks
Abstract
Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-evolution is difficult: existing benchmarks provide limited coverage of economically valuable task domains, do not always design training and test tasks such that test-time gains can be attributed to training experience, and remain vulnerable to data contamination. We present GDPevo, an evolution-native benchmark grounded in GDP-related enterprise workflows, together with the fully automated data pipeline that generates it. Its core mechanism, rule hybridization, decomposes each enterprise workflow into atomic business rules, distributes subsets of these rules across training tasks, and recombines them in held-out test tasks so that test-time gains are attributable. GDPevo spans CRM, ERP, finance, healthcare, legal, and data-centric workflows. Its V1 release contains 120 tasks in 12 groups, with five training and five held-out test tasks per group. Full automation enables the pipeline to expand the suite to 240 tasks in 24 groups (V2) within two days, providing a practical response to contamination. Using GDPevo, we evaluate four agents, each comprising a harness and a model, under four supervision types. Self-evolution consistently improves held-out accuracy by up to 16.44 percentage points. But the best evolved agents remain far below the fully informed oracle ceiling of 91.6%, indicating that the self-evolution ability of current agents remains far from fully realized. We publicly release the pipeline, benchmark, and full evaluation results at https://github.com/Prism-Shadow/GDPevo.
Community
We release GDPevo, the first benchmark for evaluating agent self-evolution on GDP-related enterprise tasks.
We also open-source the fully automated pipeline that generates evolution-native benchmark data, providing a practical response to contamination.
We found the best evolved agents remain far below a fully informed oracle ceiling of 91.6%, indicating that the self-evolution ability of current agents remains far from fully realized.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.03764 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.03764 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.03764 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.