LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through natural language. However, in many important domains like game playing and robotics, the strongest available agents are not language models. Integrating non-language agents with LLMs would require \\emph{verbalization}: compressing their rich continuous representations into sparse textual summaries at each interaction step. To study whether verbalization constitutes a bottleneck, we introduce \\textsc{LLAMIA-Bench}, a suite of six diverse collaborative chess tasks spanning three facets: behavioral imitation, state assessment, and natural-language explanation. Each task instantiates a well-established chess problem that neither the LLM nor the chess engine can solve alone. To solve LLM collaboration with non-language agents, we introduce \\emph{latent state internalization}, which projects the subagent's continuous representations directly into the LLM's token stream as learned state tokens, with dynamic re-encoding as actions advance the environment state. Comparing internalization to verbalized integration, our experiments reveal a consistent \\emph{verbalization debt}: the performance gap widens throughout training and persists as the LLM scales from 4B to 14B parameters. A single 14B model, \\textsc{LLAMIA}, trained with latent state internalization, matches or exceeds task specialists and frontier models including GPT-5.1 with tool access across all benchmark tasks, and generalizes out-of-distribution where task-specific finetunes collapse</p>\n","updatedAt":"2026-09-03T09:44:55.222Z","author":{"_id":"64b91c71e3d41dbd696d83da","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b91c71e3d41dbd696d83da/39RTP3qHve5ih20o-TTVS.jpeg","fullname":"Somesh Singh","name":"ssingh22","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8973148465156555},"editors":["ssingh22"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64b91c71e3d41dbd696d83da/39RTP3qHve5ih20o-TTVS.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.00474","authors":[{"_id":"6a994138f9b03e97e1de5be8","name":"Harini S I","hidden":false},{"_id":"6a994138f9b03e97e1de5be9","name":"Somesh Singh","hidden":false},{"_id":"6a994138f9b03e97e1de5bea","name":"Yaman K Singla","hidden":false},{"_id":"6a994138f9b03e97e1de5beb","name":"Rajiv Ratn Shah","hidden":false},{"_id":"6a994138f9b03e97e1de5bec","name":"David Doermann","hidden":false},{"_id":"6a994138f9b03e97e1de5bed","name":"Balaji Krishnamurthy","hidden":false}],"publishedAt":"2026-09-02T00:00:00.000Z","submittedOnDailyAt":"2026-09-03T00:00:00.000Z","title":"Exploring Collaboration between a language and a non-language agent","submittedOnDailyBy":{"_id":"64b91c71e3d41dbd696d83da","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b91c71e3d41dbd696d83da/39RTP3qHve5ih20o-TTVS.jpeg","isPro":false,"fullname":"Somesh Singh","user":"ssingh22","type":"user","name":"ssingh22"},"summary":"LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through natural language. However, in many important domains like game playing and robotics, the strongest available agents are not language models. Integrating non-language agents with LLMs would require verbalization: compressing their rich continuous representations into sparse textual summaries at each interaction step. To study whether verbalization constitutes a bottleneck, we introduce LLAMIA-Bench, a suite of six diverse collaborative chess tasks spanning three facets: behavioral imitation, state assessment, and natural-language explanation. Each task instantiates a well-established chess problem that neither the LLM nor the chess engine can solve alone. To solve LLM collaboration with non-language agents, we introduce latent state internalization, which projects the subagent's continuous representations directly into the LLM's token stream as learned state tokens, with dynamic re-encoding as actions advance the environment state. Comparing internalization to verbalized integration, our experiments reveal a consistent verbalization debt: the performance gap widens throughout training and persists as the LLM scales from 4B to 14B parameters. A single 14B model, LLAMIA, trained with latent state internalization, matches or exceeds task specialists and frontier models including GPT-5.1 with tool access across all benchmark tasks, and generalizes out-of-distribution where task-specific finetunes collapse","upvotes":4,"discussionId":"6a994138f9b03e97e1de5bee","projectPage":"https://behavior-in-the-wild.github.io/llamia.html","ai_summary":"A benchmark of collaborative chess tasks shows that integrating continuous subagent representations directly into language models via learned state tokens outperforms text-based verbalization and scales effectively.","ai_keywords":["LLAMIA-Bench","verbalization","latent state internalization","learned state tokens","dynamic re-encoding","verbalization debt","non-language agents","LLM orchestration"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"61e5d14f77496de0a6d95c6b","name":"adobe","fullname":"Adobe","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1645217431826-61e35e517ac6b6d06cfa8081.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64b91c71e3d41dbd696d83da","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b91c71e3d41dbd696d83da/39RTP3qHve5ih20o-TTVS.jpeg","isPro":false,"fullname":"Somesh Singh","user":"ssingh22","type":"user"},{"_id":"67626d50359424824fbcf689","avatarUrl":"/avatars/b3cf784876883d75cba5a5662e742590.svg","isPro":false,"fullname":"Varma","user":"Amokh","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"628e810454698ce61d1cc6d3","avatarUrl":"/avatars/e5c0e59a86f6a9a86612c846d89855ac.svg","isPro":false,"fullname":"Harini S I","user":"Harini","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"61e5d14f77496de0a6d95c6b","name":"adobe","fullname":"Adobe","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1645217431826-61e35e517ac6b6d06cfa8081.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.00474.md","query":{}}">
Exploring Collaboration between a language and a non-language agent
Abstract
A benchmark of collaborative chess tasks shows that integrating continuous subagent representations directly into language models via learned state tokens outperforms text-based verbalization and scales effectively.
LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through natural language. However, in many important domains like game playing and robotics, the strongest available agents are not language models. Integrating non-language agents with LLMs would require verbalization: compressing their rich continuous representations into sparse textual summaries at each interaction step. To study whether verbalization constitutes a bottleneck, we introduce LLAMIA-Bench, a suite of six diverse collaborative chess tasks spanning three facets: behavioral imitation, state assessment, and natural-language explanation. Each task instantiates a well-established chess problem that neither the LLM nor the chess engine can solve alone. To solve LLM collaboration with non-language agents, we introduce latent state internalization, which projects the subagent's continuous representations directly into the LLM's token stream as learned state tokens, with dynamic re-encoding as actions advance the environment state. Comparing internalization to verbalized integration, our experiments reveal a consistent verbalization debt: the performance gap widens throughout training and persists as the LLM scales from 4B to 14B parameters. A single 14B model, LLAMIA, trained with latent state internalization, matches or exceeds task specialists and frontier models including GPT-5.1 with tool access across all benchmark tasks, and generalizes out-of-distribution where task-specific finetunes collapse
Community
LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through natural language. However, in many important domains like game playing and robotics, the strongest available agents are not language models. Integrating non-language agents with LLMs would require \emph{verbalization}: compressing their rich continuous representations into sparse textual summaries at each interaction step. To study whether verbalization constitutes a bottleneck, we introduce \textsc{LLAMIA-Bench}, a suite of six diverse collaborative chess tasks spanning three facets: behavioral imitation, state assessment, and natural-language explanation. Each task instantiates a well-established chess problem that neither the LLM nor the chess engine can solve alone. To solve LLM collaboration with non-language agents, we introduce \emph{latent state internalization}, which projects the subagent's continuous representations directly into the LLM's token stream as learned state tokens, with dynamic re-encoding as actions advance the environment state. Comparing internalization to verbalized integration, our experiments reveal a consistent \emph{verbalization debt}: the performance gap widens throughout training and persists as the LLM scales from 4B to 14B parameters. A single 14B model, \textsc{LLAMIA}, trained with latent state internalization, matches or exceeds task specialists and frontier models including GPT-5.1 with tool access across all benchmark tasks, and generalizes out-of-distribution where task-specific finetunes collapse
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.00474 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.00474 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.00474 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.