Hugging Face Daily Papers · · 3 min read

TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<a href=\"https://cdn-uploads.huggingface.co/production/uploads/62dea2a3a8ccfacec7102dc3/h1EsJEcxErOIS8_zmo54t.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/62dea2a3a8ccfacec7102dc3/h1EsJEcxErOIS8_zmo54t.png\" alt=\"02-fig1-motivation\"></a><br><a href=\"https://cdn-uploads.huggingface.co/production/uploads/62dea2a3a8ccfacec7102dc3/UEnU7f5pcj44AXoA4mYF4.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/62dea2a3a8ccfacec7102dc3/UEnU7f5pcj44AXoA4mYF4.png\" alt=\"03-fig2-overview\"></a><br><a href=\"https://cdn-uploads.huggingface.co/production/uploads/62dea2a3a8ccfacec7102dc3/yblLNGUkyAaSb_lOHbPeJ.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/62dea2a3a8ccfacec7102dc3/yblLNGUkyAaSb_lOHbPeJ.png\" alt=\"04-fig4-keepmaps\"></a><br><a href=\"https://cdn-uploads.huggingface.co/production/uploads/62dea2a3a8ccfacec7102dc3/afxp7EiwxE4ozlgHRWzln.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/62dea2a3a8ccfacec7102dc3/afxp7EiwxE4ozlgHRWzln.png\" alt=\"05-fig5-efficiency\"></a></p>\n","updatedAt":"2026-09-14T14:22:55.046Z","author":{"_id":"62dea2a3a8ccfacec7102dc3","avatarUrl":"/avatars/2875ff8489a0b1d3378096a75af8285f.svg","fullname":"Yuhao Wang","name":"YHY-Test","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.47684600949287415},"editors":["YHY-Test"],"editorAvatarUrls":["/avatars/2875ff8489a0b1d3378096a75af8285f.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.10297","authors":[{"_id":"6aa802c15dd4cb9b4cc02201","name":"Yuhao Wang","hidden":false},{"_id":"6aa802c15dd4cb9b4cc02202","name":"Mu Qiao","hidden":false},{"_id":"6aa802c15dd4cb9b4cc02203","name":"Xindong Zhang","hidden":false},{"_id":"6aa802c15dd4cb9b4cc02204","name":"Yunzhi Zhuge","hidden":false},{"_id":"6aa802c15dd4cb9b4cc02205","name":"Lei Zhang","hidden":false},{"_id":"6aa802c15dd4cb9b4cc02206","name":"Huchuan Lu","hidden":false}],"publishedAt":"2026-09-09T00:00:00.000Z","submittedOnDailyAt":"2026-09-14T00:00:00.000Z","title":"TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents","submittedOnDailyBy":{"_id":"62dea2a3a8ccfacec7102dc3","avatarUrl":"/avatars/2875ff8489a0b1d3378096a75af8285f.svg","isPro":false,"fullname":"Yuhao Wang","user":"YHY-Test","type":"user","name":"YHY-Test"},"summary":"GUI agents accumulate high-resolution screenshots as the trajectory unfolds, increasing inference latency and memory usage. Training-free visual token pruning can reduce this cost, but cache reuse introduces a fundamental constraint. Once tokens are discarded, the corresponding visual evidence cannot be recovered without re-encoding. Pruning therefore becomes an irreversible admission decision that must remain useful for unknown future targets while preserving coverage of operable regions under tight budgets. To address these challenges, we propose \\method{}, a training-free framework for \\textbf{Trajectory-robust Admission and Coverage-aware Evidence ordering}. Specifically, we combine a query-independent layout-derived interaction prior with instruction relevance and feature novelty to rank visual evidence according to both potential future utility and diversity. Then, we reserve part of the budget for native visual tokens distributed across the screen, repairing missing spatial coverage without breaking the ordering. Together, these mechanisms produce a nested token order, allowing retained visual evidence to shrink monotonically across budgets while remaining reusable throughout the trajectory. Finally, our monotone KV contraction incrementally contracts retired frames into compact session state, avoiding repeated visual encoding or pruning. Extensive experiments across six GUI benchmarks and diverse models verify the effectiveness of our proposed under tight budgets. The source code will be released.","upvotes":0,"discussionId":"6aa802c15dd4cb9b4cc02207","githubRepo":"https://github.com/924973292/TRACE","githubRepoAddedBy":"user","ai_summary":"TRACE is a training-free framework that ranks visual evidence by future utility and diversity, reserves native tokens for spatial coverage, and contracts retired frames to reduce latency and memory in GUI agents.","ai_keywords":["visual token pruning","KV contraction","trajectory-robust admission","coverage-aware evidence ordering","GUI agents","interaction prior","nested token order"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[],"acceptLanguages":["en"],"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.10297.md","query":{}}">
Papers
arxiv:2609.10297

TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents

Published on Sep 9
· Submitted by
Yuhao Wang
on Sep 14
Authors:
,

Abstract

TRACE is a training-free framework that ranks visual evidence by future utility and diversity, reserves native tokens for spatial coverage, and contracts retired frames to reduce latency and memory in GUI agents.

GUI agents accumulate high-resolution screenshots as the trajectory unfolds, increasing inference latency and memory usage. Training-free visual token pruning can reduce this cost, but cache reuse introduces a fundamental constraint. Once tokens are discarded, the corresponding visual evidence cannot be recovered without re-encoding. Pruning therefore becomes an irreversible admission decision that must remain useful for unknown future targets while preserving coverage of operable regions under tight budgets. To address these challenges, we propose \method{}, a training-free framework for \textbf{Trajectory-robust Admission and Coverage-aware Evidence ordering}. Specifically, we combine a query-independent layout-derived interaction prior with instruction relevance and feature novelty to rank visual evidence according to both potential future utility and diversity. Then, we reserve part of the budget for native visual tokens distributed across the screen, repairing missing spatial coverage without breaking the ordering. Together, these mechanisms produce a nested token order, allowing retained visual evidence to shrink monotonically across budgets while remaining reusable throughout the trajectory. Finally, our monotone KV contraction incrementally contracts retired frames into compact session state, avoiding repeated visual encoding or pruning. Extensive experiments across six GUI benchmarks and diverse models verify the effectiveness of our proposed under tight budgets. The source code will be released.

Community

Paper submitter about 2 hours ago

02-fig1-motivation
03-fig2-overview
04-fig4-keepmaps
05-fig5-efficiency

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.10297
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.10297 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.10297 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.10297 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers