Hugging Face Daily Papers · · 3 min read

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

A very interesting take on agent data synthesis: instead of just generating and validating executable tasks, CalibForge uses solver behavior to actively reshape tasks into a solver-relative “<strong>learnable zone</strong>.” The idea of calibrating task difficulty through multi-solver disagreement or strong-vs-weak contrast feels simple but powerful, and the downstream gains suggest that which tasks we train on matters as much as how many we generate.</p>\n<p>HF: <a href=\"https://huggingface.co/datasets/AweAI-Team/CalibForge\">https://huggingface.co/datasets/AweAI-Team/CalibForge</a><br>Model: <a href=\"https://huggingface.co/collections/AweAI-Team/calibforge\">https://huggingface.co/collections/AweAI-Team/calibforge</a><br>Github: <a href=\"https://github.com/AweAI-Team/CalibForge\" rel=\"nofollow\">https://github.com/AweAI-Team/CalibForge</a></p>\n","updatedAt":"2026-08-07T02:27:03.289Z","author":{"_id":"63f06116f1a47aaea5bd497b","avatarUrl":"/avatars/7d99ffa59c4579599e852a0ffb261268.svg","fullname":"Guoxin Chen","name":"GuoxinChen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":10,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9121016263961792},"editors":["GuoxinChen"],"editorAvatarUrls":["/avatars/7d99ffa59c4579599e852a0ffb261268.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.06352","authors":[{"_id":"6a7541cee1228e04b323819f","name":"Fanzhe Meng","hidden":false},{"_id":"6a7541cee1228e04b32381a0","name":"Guoxin Chen","hidden":false},{"_id":"6a7541cee1228e04b32381a1","name":"Jiale Zhao","hidden":false},{"_id":"6a7541cee1228e04b32381a2","name":"Shuang Sun","hidden":false},{"_id":"6a7541cee1228e04b32381a3","name":"Zhiyu Lin","hidden":false},{"_id":"6a7541cee1228e04b32381a4","name":"Wayne Xin Zhao","hidden":false},{"_id":"6a7541cee1228e04b32381a5","name":"Ruihua Song","hidden":false},{"_id":"6a7541cee1228e04b32381a6","name":"Ji-Rong Wen","hidden":false},{"_id":"6a7541cee1228e04b32381a7","name":"Kai Jia","hidden":false}],"publishedAt":"2026-08-06T00:00:00.000Z","submittedOnDailyAt":"2026-08-07T00:00:00.000Z","title":"CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks","submittedOnDailyBy":{"_id":"63f06116f1a47aaea5bd497b","avatarUrl":"/avatars/7d99ffa59c4579599e852a0ffb261268.svg","isPro":false,"fullname":"Guoxin Chen","user":"GuoxinChen","type":"user","name":"GuoxinChen"},"summary":"Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver behavior to revise candidate tasks through adversarial solver calibration. Multi-solver calibration targets disagreement within a heterogeneous solver pool, whereas contrastive solver calibration targets a designated strong-pass/weak-fail relation; both operationalize a solver-relative learnable zone anchored in demonstrated solvability. Using CalibForge, we construct 5,431 calibrated terminal tasks. Our ablations show that both strategies yield more effective supervision than authoring and validation alone or ordinary single-solver feedback. Models trained on the full collection achieve 32.58% and 47.57% on Terminal-Bench 2.0. The largest improvements over the corresponding base model reach 24.71 percentage points on Terminal-Bench 2.0, 27.68 points on SWE-bench Pro, and 30.04 points on Doc2Repo. Together, these results support solver-relative learnability as a practical target for constructing effective and transferable agent training data.","upvotes":11,"discussionId":"6a7541cee1228e04b32381a8","githubRepo":"https://github.com/AweAI-Team/CalibForge","githubRepoAddedBy":"user","githubStars":9,"organization":{"_id":"698becbc51046c8986e285cd","name":"AweAI-Team","fullname":"AweAI Team","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63f06116f1a47aaea5bd497b/nyrxZ_sO7l_2dhQ_NdU6P.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63f06116f1a47aaea5bd497b","avatarUrl":"/avatars/7d99ffa59c4579599e852a0ffb261268.svg","isPro":false,"fullname":"Guoxin Chen","user":"GuoxinChen","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"65c747f1bbc318a59eceb452","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65c747f1bbc318a59eceb452/W5ERLsLFmwhbt-blcNslJ.jpeg","isPro":false,"fullname":"Shuang Sun","user":"SNHE","type":"user"},{"_id":"674476e821e39628723f13ad","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/7iUE9VSdvuX979_vERXpI.png","isPro":false,"fullname":"mfzzzzzz","user":"mfzzzzzz","type":"user"},{"_id":"658e5664188f3934c6f44b2c","avatarUrl":"/avatars/2e45823c547b7c127e3f9eb36d21520d.svg","isPro":false,"fullname":"ZhangXinyu","user":"Zhongxieye","type":"user"},{"_id":"651a29d566e78720a78317ec","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/651a29d566e78720a78317ec/WKPcw6Ziqjl44pkrHtCVa.jpeg","isPro":false,"fullname":"Jie Chen","user":"survivi","type":"user"},{"_id":"697b44e90211501623740a0c","avatarUrl":"/avatars/df0d7f40071cb13603ae342b2017469c.svg","isPro":false,"fullname":"Awe-AI","user":"Awe-AI","type":"user"},{"_id":"673b00d77ffc522a34241c05","avatarUrl":"/avatars/f9b9b924b4523f2649cc6573aeca11c5.svg","isPro":false,"fullname":"zzz","user":"zyqiu","type":"user"},{"_id":"6901b520f8f20b9d7015a38d","avatarUrl":"/avatars/49e8d709818d1f0778758e39bd39f0a3.svg","isPro":false,"fullname":"zoushun","user":"shunzou1314","type":"user"},{"_id":"665ebae8bcbb98f60db0b4b1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/665ebae8bcbb98f60db0b4b1/YTKM4qTZXh_2SeU8U7BfB.webp","isPro":false,"fullname":"Jiale Zhao","user":"Heisenburger2000","type":"user"},{"_id":"661ab1f1fa3b144a381fa454","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/661ab1f1fa3b144a381fa454/IlpZBb9NCjo7ntFwMIH53.png","isPro":false,"fullname":"Urro","user":"urroxyz","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"698becbc51046c8986e285cd","name":"AweAI-Team","fullname":"AweAI Team","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63f06116f1a47aaea5bd497b/nyrxZ_sO7l_2dhQ_NdU6P.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.06352.md","query":{}}">
Papers
arxiv:2608.06352

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

Published on Aug 6
· Submitted by
Guoxin Chen
on Aug 7
Authors:
,

Abstract

Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver behavior to revise candidate tasks through adversarial solver calibration. Multi-solver calibration targets disagreement within a heterogeneous solver pool, whereas contrastive solver calibration targets a designated strong-pass/weak-fail relation; both operationalize a solver-relative learnable zone anchored in demonstrated solvability. Using CalibForge, we construct 5,431 calibrated terminal tasks. Our ablations show that both strategies yield more effective supervision than authoring and validation alone or ordinary single-solver feedback. Models trained on the full collection achieve 32.58% and 47.57% on Terminal-Bench 2.0. The largest improvements over the corresponding base model reach 24.71 percentage points on Terminal-Bench 2.0, 27.68 points on SWE-bench Pro, and 30.04 points on Doc2Repo. Together, these results support solver-relative learnability as a practical target for constructing effective and transferable agent training data.

Community

Paper submitter about 15 hours ago

A very interesting take on agent data synthesis: instead of just generating and validating executable tasks, CalibForge uses solver behavior to actively reshape tasks into a solver-relative “learnable zone.” The idea of calibrating task difficulty through multi-solver disagreement or strong-vs-weak contrast feels simple but powerful, and the downstream gains suggest that which tasks we train on matters as much as how many we generate.

HF: https://huggingface.co/datasets/AweAI-Team/CalibForge
Model: https://huggingface.co/collections/AweAI-Team/calibforge
Github: https://github.com/AweAI-Team/CalibForge

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.06352
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.06352 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.06352 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers