Hugging Face Daily Papers · · 3 min read

τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

hyper-tau-bench</p>\n","updatedAt":"2026-09-07T07:04:14.344Z","author":{"_id":"654fe67e45c0dccd570cf4bb","avatarUrl":"/avatars/5bf84b02b452acf9c13e9254fc820a5b.svg","fullname":"Ben Shi","name":"benshi34","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"hu","probability":0.4890359938144684},"editors":["benshi34"],"editorAvatarUrls":["/avatars/5bf84b02b452acf9c13e9254fc820a5b.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.04611","authors":[{"_id":"6a9e61a4de5ea82090db6662","name":"Quan Shi","hidden":false},{"_id":"6a9e61a4de5ea82090db6663","name":"Keshav Dhandhania","hidden":false},{"_id":"6a9e61a4de5ea82090db6664","name":"Karthik Narasimhan","hidden":false},{"_id":"6a9e61a4de5ea82090db6665","name":"Victor Barres","hidden":false}],"publishedAt":"2026-09-04T00:00:00.000Z","submittedOnDailyAt":"2026-09-07T00:00:00.000Z","title":"τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction","submittedOnDailyBy":{"_id":"654fe67e45c0dccd570cf4bb","avatarUrl":"/avatars/5bf84b02b452acf9c13e9254fc820a5b.svg","isPro":false,"fullname":"Ben Shi","user":"benshi34","type":"user","name":"benshi34"},"summary":"LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and operate internal systems. Notably, the work of building them is increasingly handed to coding agents, yet existing benchmarks say little about whether an AI system can deliver one under the conditions of a real client engagement. We introduce τ^τ-bench (pronounced hyper-tau-bench), a benchmark that makes agent construction the task. A developer agent is given the records a business actually keeps, a client who holds requirements, a production API that operations must run through, a codebase to inherit, and limits on serving cost and models: the same starting point a real engagement provides. From these it must deliver a complete customer-service agent, scored by deploying that agent against held-out simulated users. Across 53 tasks spanning four domains, the strongest configuration, Claude Opus 5 under Claude Code, passes just 23.9% of evaluation simulations. Meanwhile, an expert-authored reference ceiling scores 82.2%. The failures mirror ones human agent developers see: models issue shallow queries in place of deep comprehension of the records, communicate almost nothing to the client, and experiment too little with agent architecture and serving spend, shipping the first design that runs. We aim for τ^τ-bench to turn the work of cooperative agent building into a measurable target for coding agents.","upvotes":1,"discussionId":"6a9e61a5de5ea82090db6666","projectPage":"https://sierra-research.github.io/hyper-tau-bench/","githubRepo":"https://github.com/sierra-research/hyper-tau-bench","githubRepoAddedBy":"user","ai_summary":"The τ^τ-bench benchmark evaluates coding agents on building real-world customer-service agents from business records, client requirements, and production APIs, revealing substantial gaps versus expert performance.","ai_keywords":["τ^τ-bench","coding agents","LLM agents","customer-service agent","production API","simulated users","agent architecture","serving cost"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"663d2df4871fe37c2f7beb9f","name":"sierra-research","fullname":"Sierra Technologies, Inc.","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/663d2d93ec1aafe3d6532184/YwFidLHHPZOHtNQT-3E5H.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"651d2b485d3519c0b7595af7","avatarUrl":"/avatars/00ce2ecbc35e22a90f72b9015299aa29.svg","isPro":false,"fullname":"Sijia Liu","user":"sijial430","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"663d2df4871fe37c2f7beb9f","name":"sierra-research","fullname":"Sierra Technologies, Inc.","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/663d2d93ec1aafe3d6532184/YwFidLHHPZOHtNQT-3E5H.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.04611.md","query":{}}">
Papers
arxiv:2609.04611

τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction

Published on Sep 4
· Submitted by
Ben Shi
on Sep 7
Authors:
,

Abstract

The τ^τ-bench benchmark evaluates coding agents on building real-world customer-service agents from business records, client requirements, and production APIs, revealing substantial gaps versus expert performance.

LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and operate internal systems. Notably, the work of building them is increasingly handed to coding agents, yet existing benchmarks say little about whether an AI system can deliver one under the conditions of a real client engagement. We introduce τ^τ-bench (pronounced hyper-tau-bench), a benchmark that makes agent construction the task. A developer agent is given the records a business actually keeps, a client who holds requirements, a production API that operations must run through, a codebase to inherit, and limits on serving cost and models: the same starting point a real engagement provides. From these it must deliver a complete customer-service agent, scored by deploying that agent against held-out simulated users. Across 53 tasks spanning four domains, the strongest configuration, Claude Opus 5 under Claude Code, passes just 23.9% of evaluation simulations. Meanwhile, an expert-authored reference ceiling scores 82.2%. The failures mirror ones human agent developers see: models issue shallow queries in place of deep comprehension of the records, communicate almost nothing to the client, and experiment too little with agent architecture and serving spend, shipping the first design that runs. We aim for τ^τ-bench to turn the work of cooperative agent building into a measurable target for coding agents.

Community

Paper submitter about 4 hours ago

hyper-tau-bench

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.04611
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.04611 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.04611 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.04611 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers