Hugging Face Daily Papers · · 10 min read

Mental World Modeling

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

\n\t<a id=\"🧠-mental-world-modeling-project-page\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#🧠-mental-world-modeling-project-page\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\t🧠 Mental World Modeling <a href=\"https://mental-world.github.io/\" rel=\"nofollow\"><strong>(Project Page)</strong></a>\n\t</span>\n</h1>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/647773a1168cb428e00e9a8f/0yaxUkfp487xKag1yhhlo.jpeg\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/647773a1168cb428e00e9a8f/0yaxUkfp487xKag1yhhlo.jpeg\" alt=\"fig1-v4\"></a></p>\n<h3 class=\"relative group flex items-baseline\">\n\t<a id=\"from-simulating-physical-scenes-to-simulating-the-minds-that-act-within-them\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#from-simulating-physical-scenes-to-simulating-the-minds-that-act-within-them\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\tFrom simulating physical scenes to simulating the minds that act within them.\n\t</span>\n</h3>\n<p><strong>TL;DR: A world model can reconstruct the physical scene correctly and still predict the wrong human action.</strong></p>\n<p>Why? Because human decisions are shaped not only by objects, geometry, and physical dynamics, but also by hidden mental and social variables: what someone has observed, knows, believes, wants, intends, feels, and considers socially permissible.</p>\n<p>We introduce <strong>Mental World Modeling (MWM)</strong>, a general framework that makes these mental variables part of the world state itself—rather than treating them as post-hoc explanations.</p>\n<hr>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/647773a1168cb428e00e9a8f/XwFAeyo0KfcfBJ1sFRd-p.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/647773a1168cb428e00e9a8f/XwFAeyo0KfcfBJ1sFRd-p.png\" alt=\"intro\"></a></p>\n<p><em>The world’s next state is not only physical. It is also mental.</em></p>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"the-missing-half-of-a-world-model\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#the-missing-half-of-a-world-model\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\tThe missing half of a world model\n\t</span>\n</h2>\n<p>Most existing world models focus on the physical substrate of the world: What objects and agents are present? Where are they? How will the visible scene evolve?</p>\n<p>But two physically identical scenes can produce completely different actions when the people inside them hold different beliefs, goals, emotions, relationships, or social obligations.</p>\n<p>MWM therefore maintains a <strong>coupled physical–mental world state</strong>. For every target agent, it:</p>\n<ol>\n<li>represents both the physical environment and the agents’ latent mental states;</li>\n<li>renders a target-specific partial observation—what that person can actually see, hear, know, and infer;</li>\n<li>simulates how each candidate action changes both the physical world and the mental-social world.</li>\n</ol>\n<hr>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"mentis-an-inspectable-implementation\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#mentis-an-inspectable-implementation\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\tMENTIS: an inspectable implementation\n\t</span>\n</h2>\n<p>We instantiate the framework in <strong>MENTIS</strong>, a training-free and fully inspectable baseline that forces an LLM-based system to operate as a mental world model.<br>MENTIS decomposes decision prediction into explicit stages:</p>\n<p><strong>state parsing → target-observation generation → action decomposition → coupled transition simulation → branch-level evaluation → final decision</strong></p>\n<hr>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"what-do-the-experiment-results-show\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#what-do-the-experiment-results-show\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\tWhat do the experiment results show?\n\t</span>\n</h2>\n<p>The results reveal a consistent pattern:</p>\n<ul>\n<li>Full MWM achieves the strongest decision-prediction performance for every evaluated model.</li>\n<li>Removing the mental channel degrades all models.</li>\n<li>The gains are largest in interpersonal situations, where hidden beliefs, intentions, emotions, and norms determine the action.</li>\n<li>Explicit intermediate structure substantially improves multimodal reasoning.</li>\n<li>The largest remaining bottleneck is <strong>transition simulation</strong>: predicting how an action changes the coupled physical–mental world.</li>\n</ul>\n<hr>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"why-does-this-matter\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#why-does-this-matter\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\tWhy does this matter?\n\t</span>\n</h2>\n<ul>\n<li>A physically possible action is not necessarily the action a person will take.</li>\n<li>The same movement can represent help, pressure, deception, politeness, avoidance, or trust depending on the mental and social state surrounding it. This distinction matters for embodied assistants, human–AI collaboration, education, care, interactive agents, and any AI system expected to operate around people.</li>\n<li>MWM does <strong>not</strong> claim direct access to private consciousness. Mental states are treated as approximate, task-relevant hypotheses that must remain inspectable, revisable, and responsibly used.</li>\n</ul>\n","updatedAt":"2026-08-03T08:39:51.406Z","author":{"_id":"647773a1168cb428e00e9a8f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/647773a1168cb428e00e9a8f/NiRR3ScY6Plzjibfwy1hC.jpeg","fullname":"Hao Fei","name":"scofield7419","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8534244894981384},"editors":["scofield7419"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/647773a1168cb428e00e9a8f/NiRR3ScY6Plzjibfwy1hC.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.27201","authors":[{"_id":"6a704bce02c90f968f48a163","user":{"_id":"647773a1168cb428e00e9a8f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/647773a1168cb428e00e9a8f/NiRR3ScY6Plzjibfwy1hC.jpeg","isPro":false,"fullname":"Hao Fei","user":"scofield7419","type":"user","name":"scofield7419"},"name":"Hao Fei","status":"claimed_verified","statusLastChangedAt":"2026-08-03T08:45:04.537Z","hidden":false},{"_id":"6a704bce02c90f968f48a164","name":"Yiran Zhao","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/647773a1168cb428e00e9a8f/EjF2i-2RYgQ9ju1Cqpa5X.mp4","https://cdn-uploads.huggingface.co/production/uploads/647773a1168cb428e00e9a8f/mCV186_NQj7PyxQzLwB_L.jpeg"],"publishedAt":"2026-07-29T00:00:00.000Z","submittedOnDailyAt":"2026-08-03T00:00:00.000Z","title":"Mental World Modeling","submittedOnDailyBy":{"_id":"647773a1168cb428e00e9a8f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/647773a1168cb428e00e9a8f/NiRR3ScY6Plzjibfwy1hC.jpeg","isPro":false,"fullname":"Hao Fei","user":"scofield7419","type":"user","name":"scofield7419"},"summary":"World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and how will it evolve. Human behavior, however, is driven by hidden mental state (what a person believes, wants, intends, feels, and considers socially permissible), so a model that tracks the physical scene but not what each agent knows and believes about it predicts the wrong action for the right-looking scene. We formulate Mental World Modeling (MWM), a generic theoretical framework that makes mental variables core components of a world model rather than posthoc rationales: MWM aintains a coupled physical-mental world state, renders a target-specific partial observation, and simulates how candidate actions jointly update both components. We instantiate the framework in MENTIS, a training-free and fully inspectable baseline that decomposes the process into state parsing, target-observation generation, action decomposition, coupled physical and mental transition, and branch-level value evaluation. On a manually constructed, quality-controlled dataset of situated decision scenarios spanning text, image, and sounding-video stories, experiments with 8 modern LLM-based world models demonstrate that explicitly modeling the mental state is essential for predicting human decisions. Deeper analyses further expose the bottlenecks of current mental world modeling. We expect MWM as a next stage of world modeling, from simulating physical scenes to simulating the minds that act in them.","upvotes":22,"discussionId":"6a704bcf02c90f968f48a165","projectPage":"https://mental-world.github.io/","githubRepo":"https://github.com/mental-world/Mentis","githubRepoAddedBy":"user","githubStars":1,"organization":{"_id":"6a69155eb5990a5bdd7ab739","name":"mental-world-model","fullname":"Mental World Model","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/647773a1168cb428e00e9a8f/1_PJtEFw58elU8VR6yDX6.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"647773a1168cb428e00e9a8f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/647773a1168cb428e00e9a8f/NiRR3ScY6Plzjibfwy1hC.jpeg","isPro":false,"fullname":"Hao Fei","user":"scofield7419","type":"user"},{"_id":"67f4ef862da2b7c39b6d8cf5","avatarUrl":"/avatars/72e4f7426aa852687eecccb68f303a2f.svg","isPro":false,"fullname":"YiranZhao","user":"YiranZhao","type":"user"},{"_id":"661cf93620b47b0dad13ab89","avatarUrl":"/avatars/731d9f004875bfbc96661fc1eb901ac7.svg","isPro":false,"fullname":"Ma","user":"gh0stHunter","type":"user"},{"_id":"6a705e089539fed79d140d5f","avatarUrl":"/avatars/b79f6c82c2ec90745929b936a66c0888.svg","isPro":false,"fullname":"Wang","user":"Joyce62","type":"user"},{"_id":"6a705ea1ec309fdfa205dcef","avatarUrl":"/avatars/4821dc0d0cea7750d34181489a728312.svg","isPro":false,"fullname":"wenyuan huang","user":"hwy7777777777777777","type":"user"},{"_id":"676fb48f84c521ebd0e81b74","avatarUrl":"/avatars/1e6f2ca595cf812410e4d4062d52fdc6.svg","isPro":false,"fullname":"li changan","user":"lichangan00","type":"user"},{"_id":"6a706c7446440e48c1fe0330","avatarUrl":"/avatars/ae6c52577fffef3eb6c7ec4289cb9e24.svg","isPro":false,"fullname":"liangyulan","user":"liangyulan","type":"user"},{"_id":"65e1816b35191b15a3c8f52b","avatarUrl":"/avatars/92f96f4b112b20a1d69ebfdef6cc908d.svg","isPro":false,"fullname":"llll","user":"bujiangwude","type":"user"},{"_id":"695b79e0455cd4fc691de270","avatarUrl":"/avatars/8dcf9c9c2152659d5352d5b73e8f02aa.svg","isPro":false,"fullname":"Yuheng","user":"peanut724724","type":"user"},{"_id":"64f83118ce75bb0fb5fbf88e","avatarUrl":"/avatars/9e601e6fab5213f76b7f95bbeec6eee3.svg","isPro":false,"fullname":"shelly gauss","user":"shellys","type":"user"},{"_id":"665969d21ba271cba9a0236f","avatarUrl":"/avatars/c916445c516bef8ed57f24114b840d45.svg","isPro":false,"fullname":"gaoxuan","user":"lalala9527","type":"user"},{"_id":"65c4eb7cd1dcbd30d86febec","avatarUrl":"/avatars/001c8f02e8ce794b2c21883628b2da72.svg","isPro":false,"fullname":"free-bit","user":"free-bit","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a69155eb5990a5bdd7ab739","name":"mental-world-model","fullname":"Mental World Model","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/647773a1168cb428e00e9a8f/1_PJtEFw58elU8VR6yDX6.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.27201.md","query":{}}">
Papers
arxiv:2607.27201

Mental World Modeling

Published on Jul 29
· Submitted by
Hao Fei
on Aug 3
Authors:

Abstract

World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and how will it evolve. Human behavior, however, is driven by hidden mental state (what a person believes, wants, intends, feels, and considers socially permissible), so a model that tracks the physical scene but not what each agent knows and believes about it predicts the wrong action for the right-looking scene. We formulate Mental World Modeling (MWM), a generic theoretical framework that makes mental variables core components of a world model rather than posthoc rationales: MWM aintains a coupled physical-mental world state, renders a target-specific partial observation, and simulates how candidate actions jointly update both components. We instantiate the framework in MENTIS, a training-free and fully inspectable baseline that decomposes the process into state parsing, target-observation generation, action decomposition, coupled physical and mental transition, and branch-level value evaluation. On a manually constructed, quality-controlled dataset of situated decision scenarios spanning text, image, and sounding-video stories, experiments with 8 modern LLM-based world models demonstrate that explicitly modeling the mental state is essential for predicting human decisions. Deeper analyses further expose the bottlenecks of current mental world modeling. We expect MWM as a next stage of world modeling, from simulating physical scenes to simulating the minds that act in them.

Community

Paper author Paper submitter about 4 hours ago

🧠 Mental World Modeling (Project Page)

fig1-v4

From simulating physical scenes to simulating the minds that act within them.

TL;DR: A world model can reconstruct the physical scene correctly and still predict the wrong human action.

Why? Because human decisions are shaped not only by objects, geometry, and physical dynamics, but also by hidden mental and social variables: what someone has observed, knows, believes, wants, intends, feels, and considers socially permissible.

We introduce Mental World Modeling (MWM), a general framework that makes these mental variables part of the world state itself—rather than treating them as post-hoc explanations.


intro

The world’s next state is not only physical. It is also mental.

The missing half of a world model

Most existing world models focus on the physical substrate of the world: What objects and agents are present? Where are they? How will the visible scene evolve?

But two physically identical scenes can produce completely different actions when the people inside them hold different beliefs, goals, emotions, relationships, or social obligations.

MWM therefore maintains a coupled physical–mental world state. For every target agent, it:

  1. represents both the physical environment and the agents’ latent mental states;
  2. renders a target-specific partial observation—what that person can actually see, hear, know, and infer;
  3. simulates how each candidate action changes both the physical world and the mental-social world.

MENTIS: an inspectable implementation

We instantiate the framework in MENTIS, a training-free and fully inspectable baseline that forces an LLM-based system to operate as a mental world model.
MENTIS decomposes decision prediction into explicit stages:

state parsing → target-observation generation → action decomposition → coupled transition simulation → branch-level evaluation → final decision


What do the experiment results show?

The results reveal a consistent pattern:

  • Full MWM achieves the strongest decision-prediction performance for every evaluated model.
  • Removing the mental channel degrades all models.
  • The gains are largest in interpersonal situations, where hidden beliefs, intentions, emotions, and norms determine the action.
  • Explicit intermediate structure substantially improves multimodal reasoning.
  • The largest remaining bottleneck is transition simulation: predicting how an action changes the coupled physical–mental world.

Why does this matter?

  • A physically possible action is not necessarily the action a person will take.
  • The same movement can represent help, pressure, deception, politeness, avoidance, or trust depending on the mental and social state surrounding it. This distinction matters for embodied assistants, human–AI collaboration, education, care, interactive agents, and any AI system expected to operate around people.
  • MWM does not claim direct access to private consciousness. Mental states are treated as approximate, task-relevant hypotheses that must remain inspectable, revisable, and responsibly used.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.27201
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.27201 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.27201 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.27201 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers