In collaborative tasks with asymmetric information — maps with different landmarks, or a board game only one side knows — mutual understanding has to be built through the interaction, and gaze is one of the few observable traces of that process. We compare two such tasks: HCRC MapTask, where perspectivist labels mark a reference as aligned only when speaker and addressee interpretations match, and MUNDEX, German game explanations with retrospective understanding judgments from both sides. Both annotate gaze as discrete behavioral categories rather than eye-tracking coordinates, but with different category sets (up/down/off; partner/table/away), so we map them into a shared partner/task/away vocabulary and compute gaze features around each grounding-labeled unit.</p>\n<p>The associations point the same way in both: aligned references and \"understood\" judgments come with more task-directed gaze, less partner-directed gaze, lower gaze entropy, and fewer gaze transitions — clearest for whoever leads the task (givers, explainers). In same-speaker MapTask reference chains, speaker gaze entropy drops at the mention where a referent becomes aligned. Effects are small and prediction gains over role/condition controls are modest, so we read gaze as one contributing cue to grounding rather than a standalone signal. Because the representation only needs discrete gaze labels, it should port to other corpora — happy to discuss, especially whether shared category names pick out the same interactional function across quite different tasks.</p>\n","updatedAt":"2026-09-17T01:55:21.237Z","author":{"_id":"61bf8017c88f3fd22f654086","avatarUrl":"/avatars/ea9e762b3db755fe051751577353fbca.svg","fullname":"Nan Li","name":"chnln","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9100062847137451},"editors":["chnln"],"editorAvatarUrls":["/avatars/ea9e762b3db755fe051751577353fbca.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.18011","authors":[{"_id":"6aab34431d9cc4dec7962386","user":{"_id":"61bf8017c88f3fd22f654086","avatarUrl":"/avatars/ea9e762b3db755fe051751577353fbca.svg","isPro":false,"fullname":"Nan Li","user":"chnln","type":"user","name":"chnln"},"name":"Nan Li","status":"claimed_verified","statusLastChangedAt":"2026-09-17T08:45:04.437Z","hidden":false},{"_id":"6aab34431d9cc4dec7962387","name":"Albert Gatt","hidden":false},{"_id":"6aab34431d9cc4dec7962388","name":"Massimo Poesio","hidden":false}],"publishedAt":"2026-09-16T00:00:00.000Z","submittedOnDailyAt":"2026-09-17T00:00:00.000Z","title":"Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX","submittedOnDailyBy":{"_id":"61bf8017c88f3fd22f654086","avatarUrl":"/avatars/ea9e762b3db755fe051751577353fbca.svg","isPro":false,"fullname":"Nan Li","user":"chnln","type":"user","name":"chnln"},"summary":"In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction. We ask whether gaze provides evidence about grounding across two such tasks. Working from discrete behavioral annotations, we map HCRC MapTask (Anderson et al., 1991) and MUNDEX (Türk et al., 2023) into a shared partner/task/away vocabulary and compute gaze features around task-relevant dialogue units. In both corpora, aligned reference interpretations (MapTask) and UND (understood) judgments (MUNDEX) are associated with more task-directed gaze and with less partner-directed gaze, lower gaze entropy, and fewer gaze transitions. The associations are clearest for the participant leading the task: in giver-produced references, and in explainer judgments, which also co-vary with the explainee's gaze. In same-speaker MapTask reference chains, the speaker's gaze entropy is lower at the mention where a previously non-aligned referent becomes aligned. The best gaze feature groups improve modestly over controls under grouped cross-validation: temporal features in MapTask and raw proportions in MUNDEX. Because effects are small and several weaken when recurring participants rather than dialogues are the unit of inference, we treat gaze as one contributing cue to grounding, to be interpreted alongside task and dialogue context.","upvotes":22,"discussionId":"6aab34441d9cc4dec7962389","githubRepo":"https://github.com/chnln/gaze-as-grounding-evidence","githubRepoAddedBy":"user","githubStars":0,"organization":{"_id":"67110a86b798bde20c43ac84","name":"cs-nlp-uu","fullname":"NLP Group at Utrecht University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6305d542435ec751b72434b8/A9YelDlLilyEoGkqaksoL.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"61bf8017c88f3fd22f654086","avatarUrl":"/avatars/ea9e762b3db755fe051751577353fbca.svg","isPro":false,"fullname":"Nan Li","user":"chnln","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"6a6a925739a7f0911330c546","avatarUrl":"/avatars/21568a0081650e63921f1f70775f8bdd.svg","isPro":false,"fullname":"Timothy Lee","user":"timothylee","type":"user"},{"_id":"6a6a9449144c90a0d41692f0","avatarUrl":"/avatars/6af977a850fe2dfd55c11f7100d92251.svg","isPro":false,"fullname":"David Jackson","user":"lunarbeacon","type":"user"},{"_id":"6a6c7bb54c678c98fac02eb4","avatarUrl":"/avatars/cd68c63d1aecc36dbede25f2fa063602.svg","isPro":false,"fullname":"Linda Miller","user":"Indigo-Linda","type":"user"},{"_id":"6a6c83ab626b1eb992a0842b","avatarUrl":"/avatars/3f75008ac54daff76f8cf5ac20f4cd0c.svg","isPro":false,"fullname":"Michael Anderson","user":"EmberMichael","type":"user"},{"_id":"6a6c8887e7a7d1e67347457c","avatarUrl":"/avatars/822e32e5b1c5242266d8d4d0b4eeb00d.svg","isPro":false,"fullname":"Mary Perez","user":"ZenithTrail","type":"user"},{"_id":"6a6aa73fdc0af70390107374","avatarUrl":"/avatars/7751be34ed8f0b37c9b894e95ebf0b96.svg","isPro":false,"fullname":"Mary Davis","user":"kestrelRin","type":"user"},{"_id":"6a6dc9adb34a88441964b81d","avatarUrl":"/avatars/1b426a084d1aa8f762b3ab7ab6844afc.svg","isPro":false,"fullname":"Joshua Rodriguez","user":"Ember-Mika","type":"user"},{"_id":"6a701cdfb9714d274c3b746e","avatarUrl":"/avatars/57c2eaf1bed1f8ac512bdbf4db7e0c4b.svg","isPro":false,"fullname":"Richard Harris","user":"Echo-Lens","type":"user"},{"_id":"6a7d47fcd13e0b112aca30e5","avatarUrl":"/avatars/cba5363af27a4ad7738cbd5edf73b107.svg","isPro":false,"fullname":"Kevin Moore","user":"granitetrail","type":"user"},{"_id":"6a7e7ebc3f18e32b07f4bd99","avatarUrl":"/avatars/22b50f227da5e1cf44dc32bbba7a4b91.svg","isPro":false,"fullname":"EmberTrail","user":"EmberTrail299","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"67110a86b798bde20c43ac84","name":"cs-nlp-uu","fullname":"NLP Group at Utrecht University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6305d542435ec751b72434b8/A9YelDlLilyEoGkqaksoL.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.18011.md","query":{}}">
Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX
Published on Sep 16
· Submitted by Nan Li on Sep 17 Abstract
In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction. We ask whether gaze provides evidence about grounding across two such tasks. Working from discrete behavioral annotations, we map HCRC MapTask (Anderson et al., 1991) and MUNDEX (Türk et al., 2023) into a shared partner/task/away vocabulary and compute gaze features around task-relevant dialogue units. In both corpora, aligned reference interpretations (MapTask) and UND (understood) judgments (MUNDEX) are associated with more task-directed gaze and with less partner-directed gaze, lower gaze entropy, and fewer gaze transitions. The associations are clearest for the participant leading the task: in giver-produced references, and in explainer judgments, which also co-vary with the explainee's gaze. In same-speaker MapTask reference chains, the speaker's gaze entropy is lower at the mention where a previously non-aligned referent becomes aligned. The best gaze feature groups improve modestly over controls under grouped cross-validation: temporal features in MapTask and raw proportions in MUNDEX. Because effects are small and several weaken when recurring participants rather than dialogues are the unit of inference, we treat gaze as one contributing cue to grounding, to be interpreted alongside task and dialogue context.
Community
In collaborative tasks with asymmetric information — maps with different landmarks, or a board game only one side knows — mutual understanding has to be built through the interaction, and gaze is one of the few observable traces of that process. We compare two such tasks: HCRC MapTask, where perspectivist labels mark a reference as aligned only when speaker and addressee interpretations match, and MUNDEX, German game explanations with retrospective understanding judgments from both sides. Both annotate gaze as discrete behavioral categories rather than eye-tracking coordinates, but with different category sets (up/down/off; partner/table/away), so we map them into a shared partner/task/away vocabulary and compute gaze features around each grounding-labeled unit.
The associations point the same way in both: aligned references and "understood" judgments come with more task-directed gaze, less partner-directed gaze, lower gaze entropy, and fewer gaze transitions — clearest for whoever leads the task (givers, explainers). In same-speaker MapTask reference chains, speaker gaze entropy drops at the mention where a referent becomes aligned. Effects are small and prediction gains over role/condition controls are modest, so we read gaze as one contributing cue to grounding rather than a standalone signal. Because the representation only needs discrete gaze labels, it should port to other corpora — happy to discuss, especially whether shared category names pick out the same interactional function across quite different tasks.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.18011 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.18011 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.