Hugging Face Daily Papers · · 4 min read

Bridging the Agent-World Gap: Text World Models for LLM-based Agents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We present a systematic overview of text world models for agent applications, providing insights into narrowing the agent–world gap.</p>\n","updatedAt":"2026-06-10T02:20:36.998Z","author":{"_id":"6645bdf6621ded608be9c37e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6645bdf6621ded608be9c37e/lHFUPBkqCyKc6KCaBUaCs.jpeg","fullname":"Yixia Li","name":"X1AOX1A","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8486330509185791},"editors":["X1AOX1A"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6645bdf6621ded608be9c37e/lHFUPBkqCyKc6KCaBUaCs.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2606.09032","authors":[{"_id":"6a27794b6dde1c5ef75bcf01","user":{"_id":"6645bdf6621ded608be9c37e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6645bdf6621ded608be9c37e/lHFUPBkqCyKc6KCaBUaCs.jpeg","isPro":false,"fullname":"Yixia Li","user":"X1AOX1A","type":"user","name":"X1AOX1A"},"name":"Yixia Li","status":"claimed_verified","statusLastChangedAt":"2026-06-09T12:42:28.856Z","hidden":false},{"_id":"6a27794b6dde1c5ef75bcf02","name":"Hongru Wang","hidden":false},{"_id":"6a27794b6dde1c5ef75bcf03","name":"Peng Lai","hidden":false},{"_id":"6a27794b6dde1c5ef75bcf04","name":"Zhiwen Ruan","hidden":false},{"_id":"6a27794b6dde1c5ef75bcf05","name":"He Zhu","hidden":false},{"_id":"6a27794b6dde1c5ef75bcf06","name":"Youxin Zhu","hidden":false},{"_id":"6a27794b6dde1c5ef75bcf07","name":"Ganlong Zhao","hidden":false},{"_id":"6a27794b6dde1c5ef75bcf08","name":"Minda Hu","hidden":false},{"_id":"6a27794b6dde1c5ef75bcf09","name":"Yun Chen","hidden":false},{"_id":"6a27794b6dde1c5ef75bcf0a","name":"Sibei Yang","hidden":false},{"_id":"6a27794b6dde1c5ef75bcf0b","name":"Peng Li","hidden":false},{"_id":"6a27794b6dde1c5ef75bcf0c","name":"Jeff Z. Pan","hidden":false},{"_id":"6a27794b6dde1c5ef75bcf0d","name":"Jia Pan","hidden":false},{"_id":"6a27794b6dde1c5ef75bcf0e","name":"Guanhua Chen","hidden":false},{"_id":"6a27794b6dde1c5ef75bcf0f","name":"Yang Liu","hidden":false},{"_id":"6a27794b6dde1c5ef75bcf10","name":"Guanbin Li","hidden":false}],"publishedAt":"2026-06-08T00:00:00.000Z","submittedOnDailyAt":"2026-06-10T00:00:00.000Z","title":"Bridging the Agent-World Gap: Text World Models for LLM-based Agents","submittedOnDailyBy":{"_id":"6645bdf6621ded608be9c37e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6645bdf6621ded608be9c37e/lHFUPBkqCyKc6KCaBUaCs.jpeg","isPro":false,"fullname":"Yixia Li","user":"X1AOX1A","type":"user","name":"X1AOX1A"},"summary":"Large language model (LLM)-based agents are increasingly used in interactive textual environments, from web navigation and code editing to tool use and long-horizon dialogue. Yet many remain largely reactive, mapping observations to actions without an explicit model of how these environments are structured and evolve. This motivates text world models (TWMs): transition models over textual states that, given a state and a candidate action, predict the resulting webpage, terminal output, API response, or user reply, thereby supporting planning, efficient learning, and principled evaluation. We systematically review text world models for LLM-based agents, organized around a formal framework and the agent lifecycle: (1) Foundations, defining text world models and characterizing them by state representation and grounding domain; (2) Construction, taxonomizing LLM-as-WM and code-as-WM paradigms and reviewing methods for building them; (3) Application, examining how world models support agents at training time through experience synthesis and at inference time through planning, verification, and adaptation; and (4) Evaluation, covering both evaluation of the world model itself and its use as an evaluation environment for agents. We aim to consolidate this rapidly developing area, clarify its design space, and highlight open challenges for future research.","upvotes":6,"discussionId":"6a27794b6dde1c5ef75bcf11","githubRepo":"https://github.com/sustech-nlp/awesome-text-world-models","githubRepoAddedBy":"user","ai_summary":"Text world models serve as transition models for LLM-based agents in interactive environments, enabling planning and efficient learning by predicting environmental changes from textual states and actions.","ai_keywords":["text world models","LLM-based agents","textual states","transition models","planning","experience synthesis","verification","adaptation"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":8,"organization":{"_id":"63d8ca1889a9683b46cf1175","name":"SUSTech-NLP","fullname":"NLP Group in SUSTech","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63d8c8e0da4f723392421bea/df08FZFYIJtKTS19qUCLA.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6645bdf6621ded608be9c37e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6645bdf6621ded608be9c37e/lHFUPBkqCyKc6KCaBUaCs.jpeg","isPro":false,"fullname":"Yixia Li","user":"X1AOX1A","type":"user"},{"_id":"6587e778dda02636b0ca8b05","avatarUrl":"/avatars/aef30eae9f7656140d0626aaa302495e.svg","isPro":false,"fullname":"YX","user":"ADUIDUIDUIi","type":"user"},{"_id":"63d8c8e0da4f723392421bea","avatarUrl":"/avatars/a11dea1c5d686a8e562882801152893f.svg","isPro":false,"fullname":"Guanhua","user":"eric18","type":"user"},{"_id":"6757ed528066b63cd7587f62","avatarUrl":"/avatars/2eb56fded80ee02f11fbbfea588da592.svg","isPro":false,"fullname":"laip","user":"lllp11","type":"user"},{"_id":"6878e75c25d1ed7d2b56a36f","avatarUrl":"/avatars/5056dd5cdaa442b260bc2fd85eea133a.svg","isPro":false,"fullname":"TIANYI","user":"BIMU233","type":"user"},{"_id":"6784e35486481d2df3b2b13e","avatarUrl":"/avatars/cfeda3791587fbc9ded6185b713b44c1.svg","isPro":false,"fullname":"Wanxing Wu","user":"wanxwu","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63d8ca1889a9683b46cf1175","name":"SUSTech-NLP","fullname":"NLP Group in SUSTech","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63d8c8e0da4f723392421bea/df08FZFYIJtKTS19qUCLA.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2606/2606.09032.md"}">
Papers
arxiv:2606.09032

Bridging the Agent-World Gap: Text World Models for LLM-based Agents

Published on Jun 8
· Submitted by
Yixia Li
on Jun 10
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

Text world models serve as transition models for LLM-based agents in interactive environments, enabling planning and efficient learning by predicting environmental changes from textual states and actions.

Large language model (LLM)-based agents are increasingly used in interactive textual environments, from web navigation and code editing to tool use and long-horizon dialogue. Yet many remain largely reactive, mapping observations to actions without an explicit model of how these environments are structured and evolve. This motivates text world models (TWMs): transition models over textual states that, given a state and a candidate action, predict the resulting webpage, terminal output, API response, or user reply, thereby supporting planning, efficient learning, and principled evaluation. We systematically review text world models for LLM-based agents, organized around a formal framework and the agent lifecycle: (1) Foundations, defining text world models and characterizing them by state representation and grounding domain; (2) Construction, taxonomizing LLM-as-WM and code-as-WM paradigms and reviewing methods for building them; (3) Application, examining how world models support agents at training time through experience synthesis and at inference time through planning, verification, and adaptation; and (4) Evaluation, covering both evaluation of the world model itself and its use as an evaluation environment for agents. We aim to consolidate this rapidly developing area, clarify its design space, and highlight open challenges for future research.

Community

Paper author Paper submitter about 15 hours ago

We present a systematic overview of text world models for agent applications, providing insights into narrowing the agent–world gap.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2606.09032
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2606.09032 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2606.09032 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2606.09032 in a Space README.md to link it from this page.

Collections including this paper 1

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers