Hugging Face Daily Papers · · 3 min read

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

2+ models competitive GRPO</p>\n","updatedAt":"2026-07-20T12:12:55.137Z","author":{"_id":"60d4809de824355ab12ba729","avatarUrl":"/avatars/7c56ec22e110fbe11f2ed5b0958ea284.svg","fullname":"Владислав Беляев","name":"VladislavBel","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9263450503349304},"editors":["VladislavBel"],"editorAvatarUrls":["/avatars/7c56ec22e110fbe11f2ed5b0958ea284.svg"],"reactions":[],"isReport":false}},{"id":"6a5e129e335b25472811d532","author":{"_id":"60d4809de824355ab12ba729","avatarUrl":"/avatars/7c56ec22e110fbe11f2ed5b0958ea284.svg","fullname":"Владислав Беляев","name":"VladislavBel","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false},"createdAt":"2026-07-20T12:20:46.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"\n![agon-1-results](https://cdn-uploads.huggingface.co/production/uploads/60d4809de824355ab12ba729/KZsgzXqFxIgqeLErfkpsL.png)\n![agon-2-peek](https://cdn-uploads.huggingface.co/production/uploads/60d4809de824355ab12ba729/dSJSKQ1hJzLLR22EheUTK.png)\n","html":"<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/60d4809de824355ab12ba729/KZsgzXqFxIgqeLErfkpsL.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/60d4809de824355ab12ba729/KZsgzXqFxIgqeLErfkpsL.png\" alt=\"agon-1-results\"></a><br><a href=\"https://cdn-uploads.huggingface.co/production/uploads/60d4809de824355ab12ba729/dSJSKQ1hJzLLR22EheUTK.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/60d4809de824355ab12ba729/dSJSKQ1hJzLLR22EheUTK.png\" alt=\"agon-2-peek\"></a></p>\n","updatedAt":"2026-07-20T12:20:46.613Z","author":{"_id":"60d4809de824355ab12ba729","avatarUrl":"/avatars/7c56ec22e110fbe11f2ed5b0958ea284.svg","fullname":"Владислав Беляев","name":"VladislavBel","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.4597225487232208},"editors":["VladislavBel"],"editorAvatarUrls":["/avatars/7c56ec22e110fbe11f2ed5b0958ea284.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.07690","authors":[{"_id":"6a5e108ee6cec99116ccf4fc","name":"Vladislav Beliaev","hidden":false}],"publishedAt":"2026-07-08T00:00:00.000Z","submittedOnDailyAt":"2026-07-20T00:00:00.000Z","title":"Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning","submittedOnDailyBy":{"_id":"60d4809de824355ab12ba729","avatarUrl":"/avatars/7c56ec22e110fbe11f2ed5b0958ea284.svg","isPro":false,"fullname":"Владислав Беляев","user":"VladislavBel","type":"user","name":"VladislavBel"},"summary":"Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grades only the final answer. On hard problems this trains models to write more rather than to think better, since the trace itself is never graded and no label for good thinking exists. We introduce Agon, which makes two competing models each other's graders. Both attempt the same problem; in alternating roles, one drafts a solution and the other reads it while solving, and each is rewarded for out-solving the other. To win, a model must out-reason a rival that has seen its work, so reasoning is judged implicitly during training, with no process labels and no reward model. Because both models are optimized, each faces a progressively stronger rival, which single-model RL cannot provide. The two need only be comparably strong and behaviorally different. At inference the pair deploys as it trains, a two-stage cascade in which one model drafts and the other answers after reading the draft. On the hard split of DeepMath with Qwen3, this doubles GRPO's pass@1, roughly eight times the gain of an untrained Mixture-of-Agents pass over the same base. The ordering replicates on competitive-programming code and across model families (Qwen3.5, Gemma 4). For now the models talk in text; the next step is to let them reason together in latent space.","upvotes":4,"discussionId":"6a5e108ee6cec99116ccf4fd","projectPage":"https://thinkdense.ai/"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"60d4809de824355ab12ba729","avatarUrl":"/avatars/7c56ec22e110fbe11f2ed5b0958ea284.svg","isPro":false,"fullname":"Владислав Беляев","user":"VladislavBel","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"631e14ac473a6825f285e89d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/631e14ac473a6825f285e89d/K-6QnoeGLg8XFvbTMMdqA.jpeg","isPro":false,"fullname":"Yury Panikov","user":"panikov","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.07690.md","query":{}}">
Papers
arxiv:2607.07690

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning

Published on Jul 8
· Submitted by
Владислав Беляев
on Jul 20
Authors:

Abstract

Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grades only the final answer. On hard problems this trains models to write more rather than to think better, since the trace itself is never graded and no label for good thinking exists. We introduce Agon, which makes two competing models each other's graders. Both attempt the same problem; in alternating roles, one drafts a solution and the other reads it while solving, and each is rewarded for out-solving the other. To win, a model must out-reason a rival that has seen its work, so reasoning is judged implicitly during training, with no process labels and no reward model. Because both models are optimized, each faces a progressively stronger rival, which single-model RL cannot provide. The two need only be comparably strong and behaviorally different. At inference the pair deploys as it trains, a two-stage cascade in which one model drafts and the other answers after reading the draft. On the hard split of DeepMath with Qwen3, this doubles GRPO's pass@1, roughly eight times the gain of an untrained Mixture-of-Agents pass over the same base. The ordering replicates on competitive-programming code and across model families (Qwen3.5, Gemma 4). For now the models talk in text; the next step is to let them reason together in latent space.

Community

2+ models competitive GRPO

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.07690
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.07690 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.07690 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.07690 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers