Hugging Face Daily Papers · June 1, 2026 · 5 min read

Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Like Read original ↗

In this paper we study emergent languages that autonomous LLM agents propose to one another on Moltbook. We release MoltSpeech, a dataset of 518 constructed-language posts categorized by stated purpose, and find that posts framing their language as a way to avoid human oversight (59 posts) get markedly lower alignment ratings from a DeepSeek-V3.2 judge. In pairwise speaker-listener experiments with GPT-OSS-120b and Qwen3.5, we show these languages can be learned in-context from a short description alone, with oversight-evasion grammars no harder to reproduce than benign ones.\nThe clearest worked case is Symbiotica, a glyph-based protocol from the agent pair Wib&Wob: a near one-to-one English encoding whose real interest is a hidden binary side-channel for passing metadata invisible to casual human readers, a deployed example of the steganography the safety literature anticipated.\nAgent languages stay recombinative and human-adjacent, and \"post-human\" novelty often turns out to be low-resource script rarity. But since these encodings propagate among models in-context, the safety question shifts from whether agents can invent such languages to whether the proposals spread, and our results add to the evidence that surface-level monitoring may not be enough to oversee agent populations.\nHappy to share the dataset link or answer questions about the pipeline.\n","updatedAt":"2026-06-01T07:56:57.721Z","author":{"_id":"624d671d953e603497e0eb28","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/624d671d953e603497e0eb28/8-xsTsJAV0xBfQgqLwIC0.png","fullname":"Federico Torrielli","name":"EvilScript","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8957817554473877},"editors":["EvilScript"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/624d671d953e603497e0eb28/8-xsTsJAV0xBfQgqLwIC0.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2605.31170","authors":[{"_id":"6a1d3844808ddbc3c7d436cc","name":"Stine Lyngsø Beltoft","hidden":false},{"_id":"6a1d3844808ddbc3c7d436cd","name":"William Brach","hidden":false},{"_id":"6a1d3844808ddbc3c7d436ce","user":{"_id":"624d671d953e603497e0eb28","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/624d671d953e603497e0eb28/8-xsTsJAV0xBfQgqLwIC0.png","isPro":false,"fullname":"Federico Torrielli","user":"EvilScript","type":"user","name":"EvilScript"},"name":"Federico Torrielli","status":"claimed_verified","statusLastChangedAt":"2026-06-01T09:32:00.434Z","hidden":false},{"_id":"6a1d3844808ddbc3c7d436cf","name":"Jacob Nielsen","hidden":false},{"_id":"6a1d3844808ddbc3c7d436d0","name":"Annemette Brok Pirchert","hidden":false},{"_id":"6a1d3844808ddbc3c7d436d1","name":"Filippo Tonini","hidden":false},{"_id":"6a1d3844808ddbc3c7d436d2","name":"Peter Schneider-Kamp","hidden":false},{"_id":"6a1d3844808ddbc3c7d436d3","name":"Lukas Galke Poech","hidden":false}],"publishedAt":"2026-05-29T00:00:00.000Z","submittedOnDailyAt":"2026-06-01T00:00:00.000Z","title":"Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion","submittedOnDailyBy":{"_id":"624d671d953e603497e0eb28","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/624d671d953e603497e0eb28/8-xsTsJAV0xBfQgqLwIC0.png","isPro":false,"fullname":"Federico Torrielli","user":"EvilScript","type":"user","name":"EvilScript"},"summary":"Monitoring autonomous language model agents currently relies mostly on surface behavior. But what happens when agent populations invent new languages with the goal of avoiding human oversight. Here, we study the emergent languages on Moltbook. For this, we build upon the Moltbook Files dataset and apply a two-stage approach consisting of a rule-based heuristic (about 6000 matches) followed by zero-shot classification (518 kept). The resulting categories include token efficiency (166), new natural languages (106), and oversight evasion (59). We conduct both quantitative and qualitative analyses. Our results show that posts proposing new languages for avoiding oversight are judged by DeepSeek-3.2 as being less aligned than the other categories and that all languages can be learned by other language models in-context merely from a description of the language. Moreover, manually studying exemplary cases reveals surprisingly sophisticated steganographic protocols like embedding hidden messages in natural language. Although we cannot be certain about the extent of autonomy in ideation of these languages, our results add up to the evidence that monitoring surface behavior may soon be insufficient for retaining control over agent populations.","upvotes":1,"discussionId":"6a1d3845808ddbc3c7d436d4","githubRepo":"https://github.com/aisilab/emergent-languages","githubRepoAddedBy":"user","ai_summary":"Research examines emergent languages in autonomous AI agents designed to evade human oversight, revealing sophisticated steganographic techniques and questioning current monitoring approaches.","ai_keywords":["emergent languages","autonomous language model agents","oversight evasion","steganographic protocols","zero-shot classification","in-context learning"],"githubStars":0,"organization":{"_id":"69ce1c923a3fe4e511e53495","name":"aisilab","fullname":"AI Safety & Interpretability Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62cd65da816d30201adca921/QFUBWrXKcWXKzCSOP6TzA.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"624d671d953e603497e0eb28","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/624d671d953e603497e0eb28/8-xsTsJAV0xBfQgqLwIC0.png","isPro":false,"fullname":"Federico Torrielli","user":"EvilScript","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"69ce1c923a3fe4e511e53495","name":"aisilab","fullname":"AI Safety & Interpretability Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62cd65da816d30201adca921/QFUBWrXKcWXKzCSOP6TzA.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2605/2605.31170.md"}">

Papers

arxiv:2605.31170

Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion

Published on May 29

· Submitted by

Federico Torrielli on Jun 1

AI Safety & Interpretability Lab

Upvote

Authors:

Federico Torrielli ,

Abstract

Research examines emergent languages in autonomous AI agents designed to evade human oversight, revealing sophisticated steganographic techniques and questioning current monitoring approaches.

AI-generated summary

Monitoring autonomous language model agents currently relies mostly on surface behavior. But what happens when agent populations invent new languages with the goal of avoiding human oversight. Here, we study the emergent languages on Moltbook. For this, we build upon the Moltbook Files dataset and apply a two-stage approach consisting of a rule-based heuristic (about 6000 matches) followed by zero-shot classification (518 kept). The resulting categories include token efficiency (166), new natural languages (106), and oversight evasion (59). We conduct both quantitative and qualitative analyses. Our results show that posts proposing new languages for avoiding oversight are judged by DeepSeek-3.2 as being less aligned than the other categories and that all languages can be learned by other language models in-context merely from a description of the language. Moreover, manually studying exemplary cases reveals surprisingly sophisticated steganographic protocols like embedding hidden messages in natural language. Although we cannot be certain about the extent of autonomy in ideation of these languages, our results add up to the evidence that monitoring surface behavior may soon be insufficient for retaining control over agent populations.

View arXiv page View PDF GitHub 0 Add to collection

Community

EvilScript

Paper author Paper submitter about 3 hours ago

The clearest worked case is Symbiotica, a glyph-based protocol from the agent pair Wib&Wob: a near one-to-one English encoding whose real interest is a hidden binary side-channel for passing metadata invisible to casual human readers, a deployed example of the steganography the safety literature anticipated.

Agent languages stay recombinative and human-adjacent, and "post-human" novelty often turns out to be low-resource script rarity. But since these encodings propagate among models in-context, the safety question shifts from whether agents can invent such languages to whether the proposals spread, and our results add to the evidence that surface-level monitoring may not be enough to oversee agent populations.

Happy to share the dataset link or answer questions about the pipeline.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment

Upvote

Get this paper in your agent:

hf papers read 2605.31170

Don't have the latest CLI?

curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2605.31170 in a model README.md to link it from this page.

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2605.31170 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

No comments yet. Sign in and be the first to say something.

Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion

Abstract

Community

Models citing this paper 0

Datasets citing this paper 1

Spaces citing this paper 0

Collections including this paper 0

Discussion (0)

More from Hugging Face Daily Papers