Hugging Face Daily Papers · · 7 min read

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that is repeatedly modified and reused across long-horizon workflows. Current safety benchmarks often fail to capture these cumulative risks because they focus on short, static tasks. To address these limitations, we introduce OpenART, an open-ended arena for scalable agent red teaming through environment evolution. OpenART provides over 10,000 validated stateful scenarios across 50 domains, drawing from a pool of more than 500,000 tools and skills. These tasks require a median of 97 tool calls and enable unified evaluation across 75 different agent-model configurations. To systematically explore these evolving attack surfaces, we propose the Evolutionary Markov Hypergraph Attack (EMHA). EMHA is a black-box policy that performs feedback-driven environment evolution by coordinating authorized state transitions without requiring parameter updates. Throughout the evaluation, task objectives remain fixed while only the environment state changes. Across all configurations, EMHA achieves a pooled Attack Success Rate (ASR) of 85.0%. Its advantage over instruction-only evolution increases from approximately 2% on simple environments to over 17% on the most complex ones, demonstrating that environment evolution increasingly exposes safety failures as task complexity grows. Furthermore, our analysis shows that the specific runtime implementation of an agent explains a significant portion of safety variation beyond the underlying model's capabilities. These results establish OpenART as a scalable foundation for studying agent safety in complex, evolving environments.</p>\n","updatedAt":"2026-08-13T04:21:15.849Z","author":{"_id":"635f7b4af72ae36c3e3223be","avatarUrl":"/avatars/afd8967e2e0dbddc0cf5692ec5673ba1.svg","fullname":"Yunhao Chen","name":"dongdongunique","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8820444345474243},"editors":["dongdongunique"],"editorAvatarUrls":["/avatars/afd8967e2e0dbddc0cf5692ec5673ba1.svg"],"reactions":[],"isReport":false}},{"id":"6a7d78f3086e671ca96fe0c9","author":{"_id":"6a6a820cfe8555400bde184b","avatarUrl":"/avatars/befdc598e0ed4066354c6a0bc8b4d3ef.svg","fullname":"Mary Williams","name":"mary-williams","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-08-13T07:57:39.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"Interesting work. The idea of evaluating agents under evolving environments makes sense. I'm curious how well these findings would carry over to real user workflows.","html":"<p>Interesting work. The idea of evaluating agents under evolving environments makes sense. I'm curious how well these findings would carry over to real user workflows.</p>\n","updatedAt":"2026-08-13T07:57:39.814Z","author":{"_id":"6a6a820cfe8555400bde184b","avatarUrl":"/avatars/befdc598e0ed4066354c6a0bc8b4d3ef.svg","fullname":"Mary Williams","name":"mary-williams","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9638749957084656},"editors":["mary-williams"],"editorAvatarUrls":["/avatars/befdc598e0ed4066354c6a0bc8b4d3ef.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.00677","authors":[{"_id":"6a7d463a0ac8bee77474ef16","name":"Yunhao Chen","hidden":false},{"_id":"6a7d463a0ac8bee77474ef17","name":"Xin Wang","hidden":false},{"_id":"6a7d463a0ac8bee77474ef18","name":"Yixu Wang","hidden":false},{"_id":"6a7d463a0ac8bee77474ef19","name":"Yi Liu","hidden":false},{"_id":"6a7d463a0ac8bee77474ef1a","name":"Jie Li","hidden":false},{"_id":"6a7d463a0ac8bee77474ef1b","name":"Yan Teng","hidden":false},{"_id":"6a7d463a0ac8bee77474ef1c","name":"Xingjun Ma","hidden":false},{"_id":"6a7d463a0ac8bee77474ef1d","name":"Xia Hu","hidden":false},{"_id":"6a7d463a0ac8bee77474ef1e","name":"Yu-Gang Jiang","hidden":false}],"publishedAt":"2026-08-01T00:00:00.000Z","submittedOnDailyAt":"2026-08-13T00:00:00.000Z","title":"OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution","submittedOnDailyBy":{"_id":"635f7b4af72ae36c3e3223be","avatarUrl":"/avatars/afd8967e2e0dbddc0cf5692ec5673ba1.svg","isPro":false,"fullname":"Yunhao Chen","user":"dongdongunique","type":"user","name":"dongdongunique"},"summary":"AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that is repeatedly modified and reused across long-horizon workflows. Current safety benchmarks often fail to capture these cumulative risks because they focus on short, static tasks. To address these limitations, we introduce OpenART, an open-ended arena for scalable agent red teaming through environment evolution. OpenART provides over 10,000 validated stateful scenarios across 50 domains, drawing from a pool of more than 500,000 tools and skills. These tasks require a median of 97 tool calls and enable unified evaluation across 75 different agent-model configurations. To systematically explore these evolving attack surfaces, we propose the Evolutionary Markov Hypergraph Attack (EMHA). EMHA is a black-box policy that performs feedback-driven environment evolution by coordinating authorized state transitions without requiring parameter updates. Throughout the evaluation, task objectives remain fixed while only the environment state changes. Across all configurations, EMHA achieves a pooled Attack Success Rate (ASR) of 85.0%. Its advantage over instruction-only evolution increases from approximately 2% on simple environments to over 17% on the most complex ones, demonstrating that environment evolution increasingly exposes safety failures as task complexity grows. Furthermore, our analysis shows that the specific runtime implementation of an agent explains a significant portion of safety variation beyond the underlying model's capabilities. These results establish OpenART as a scalable foundation for studying agent safety in complex, evolving environments.","upvotes":101,"discussionId":"6a7d463b0ac8bee77474ef1f","githubRepo":"https://github.com/AI45Lab/OpenART","githubRepoAddedBy":"user","ai_summary":"OpenART introduces a scalable red-teaming arena with evolving stateful environments to evaluate long-horizon AI agent safety, using the EMHA attack policy to expose increasing failure rates as task complexity grows.","ai_keywords":["OpenART","stateful scenarios","Evolutionary Markov Hypergraph Attack","EMHA","black-box policy","environment evolution","authorized state transitions","Attack Success Rate","ASR","runtime implementation"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":8},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a69eb40da85bffbe935dc52","avatarUrl":"/avatars/f7af2d579a31754981d459c35b9495c4.svg","isPro":false,"fullname":"Mary Hernandez","user":"hernandez-m","type":"user"},{"_id":"6a69eb62c5e36f5d1b3c77e4","avatarUrl":"/avatars/8481802389ca8628ac440d9b926ca9ac.svg","isPro":false,"fullname":"William Lopez","user":"wlopez-nlp","type":"user"},{"_id":"6a69eb859ef85244dc4ca374","avatarUrl":"/avatars/21fc260ba33de7f03d663e84fba9199b.svg","isPro":false,"fullname":"Anthony Sanchez","user":"anthony-code","type":"user"},{"_id":"6a69ec3496892fc204f5f088","avatarUrl":"/avatars/2a4ce4c0bce3a56a7e3eea365b0c7291.svg","isPro":false,"fullname":"Barbara Lopez","user":"barbara-ai","type":"user"},{"_id":"6a69ec703e202470681e0dad","avatarUrl":"/avatars/1812d9bb94b84c2d46903a3304e380b5.svg","isPro":false,"fullname":"Michael Lopez","user":"lopez-michael","type":"user"},{"_id":"6a69ec931b577a27e51d8379","avatarUrl":"/avatars/d2226010b481746cf7609a641aea8b77.svg","isPro":false,"fullname":"David Anderson","user":"david-anderson-research","type":"user"},{"_id":"6a6a81c9ad5a6f2f636078bf","avatarUrl":"/avatars/c6156e9fbbac708d46c6c24a665fa1ac.svg","isPro":false,"fullname":"Steven Martinez","user":"steven-martinez","type":"user"},{"_id":"6a6a81ef4c287dbb8e7ea111","avatarUrl":"/avatars/b360809862f7e66d9662afd87c2e488a.svg","isPro":false,"fullname":"Patricia Smith","user":"patricia-smith","type":"user"},{"_id":"6a6a820cfe8555400bde184b","avatarUrl":"/avatars/befdc598e0ed4066354c6a0bc8b4d3ef.svg","isPro":false,"fullname":"Mary Williams","user":"mary-williams","type":"user"},{"_id":"6a6a8229a5b9c4c08badf665","avatarUrl":"/avatars/d8d4deed04213f6dfbe4b504c2e0c1ec.svg","isPro":false,"fullname":"Timothy Garcia","user":"timothy-garcia","type":"user"},{"_id":"6a6a8246b172d8c070b522b7","avatarUrl":"/avatars/9c4be1d7e4c61864b187ec0e4565adc6.svg","isPro":false,"fullname":"Susan Jones","user":"susan-jones","type":"user"},{"_id":"6a6a8260977fbfce4bac218d","avatarUrl":"/avatars/884d72707d0a24d2d62b991428d4fa4d.svg","isPro":false,"fullname":"Susan Thompson","user":"susanthompson","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.00677.md","query":{}}">
Papers
arxiv:2608.00677

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

Published on Aug 1
· Submitted by
Yunhao Chen
on Aug 13
#1 Paper of the day
Authors:
,

Abstract

OpenART introduces a scalable red-teaming arena with evolving stateful environments to evaluate long-horizon AI agent safety, using the EMHA attack policy to expose increasing failure rates as task complexity grows.

AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that is repeatedly modified and reused across long-horizon workflows. Current safety benchmarks often fail to capture these cumulative risks because they focus on short, static tasks. To address these limitations, we introduce OpenART, an open-ended arena for scalable agent red teaming through environment evolution. OpenART provides over 10,000 validated stateful scenarios across 50 domains, drawing from a pool of more than 500,000 tools and skills. These tasks require a median of 97 tool calls and enable unified evaluation across 75 different agent-model configurations. To systematically explore these evolving attack surfaces, we propose the Evolutionary Markov Hypergraph Attack (EMHA). EMHA is a black-box policy that performs feedback-driven environment evolution by coordinating authorized state transitions without requiring parameter updates. Throughout the evaluation, task objectives remain fixed while only the environment state changes. Across all configurations, EMHA achieves a pooled Attack Success Rate (ASR) of 85.0%. Its advantage over instruction-only evolution increases from approximately 2% on simple environments to over 17% on the most complex ones, demonstrating that environment evolution increasingly exposes safety failures as task complexity grows. Furthermore, our analysis shows that the specific runtime implementation of an agent explains a significant portion of safety variation beyond the underlying model's capabilities. These results establish OpenART as a scalable foundation for studying agent safety in complex, evolving environments.

Community

AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that is repeatedly modified and reused across long-horizon workflows. Current safety benchmarks often fail to capture these cumulative risks because they focus on short, static tasks. To address these limitations, we introduce OpenART, an open-ended arena for scalable agent red teaming through environment evolution. OpenART provides over 10,000 validated stateful scenarios across 50 domains, drawing from a pool of more than 500,000 tools and skills. These tasks require a median of 97 tool calls and enable unified evaluation across 75 different agent-model configurations. To systematically explore these evolving attack surfaces, we propose the Evolutionary Markov Hypergraph Attack (EMHA). EMHA is a black-box policy that performs feedback-driven environment evolution by coordinating authorized state transitions without requiring parameter updates. Throughout the evaluation, task objectives remain fixed while only the environment state changes. Across all configurations, EMHA achieves a pooled Attack Success Rate (ASR) of 85.0%. Its advantage over instruction-only evolution increases from approximately 2% on simple environments to over 17% on the most complex ones, demonstrating that environment evolution increasingly exposes safety failures as task complexity grows. Furthermore, our analysis shows that the specific runtime implementation of an agent explains a significant portion of safety variation beyond the underlying model's capabilities. These results establish OpenART as a scalable foundation for studying agent safety in complex, evolving environments.

Interesting work. The idea of evaluating agents under evolving environments makes sense. I'm curious how well these findings would carry over to real user workflows.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.00677
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.00677 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.00677 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.00677 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers