Hugging Face Daily Papers · · 3 min read

360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

🚀 #ECCV2026 𝐈𝐧𝐭𝐫𝐨𝐝𝐮𝐜𝐢𝐧𝐠 𝟑𝟔𝟎𝐂𝐢𝐭𝐲𝐀𝐫𝐞𝐧𝐚 🌆🧭</p>\n<p>\"Can AI truly understand and navigate a real city?\"</p>\n<p>We introduce 360CityArena, a realistic urban navigation benchmark built from 360° videos of Akihabara, Tokyo. 🇯🇵<br>🏙️ Our arena:<br>✅ 602 real-world 360° videos*<br>✅ 85 streets*<br>✅ 175 navigation &amp; spatial reasoning tasks</p>\n","updatedAt":"2026-08-12T07:42:39.548Z","author":{"_id":"6527b37c0ae663e384eb1b85","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6527b37c0ae663e384eb1b85/zKWa8h6YU4BWfcitpM5Pl.png","fullname":"Atsuyuki Miyai","name":"AtsuMiyai","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":7,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5529422163963318},"editors":["AtsuMiyai"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6527b37c0ae663e384eb1b85/zKWa8h6YU4BWfcitpM5Pl.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.08814","authors":[{"_id":"6a7c23b61653ef87c6af1db3","name":"Kenta Watanabe","hidden":false},{"_id":"6a7c23b61653ef87c6af1db4","name":"Atsuyuki Miyai","hidden":false},{"_id":"6a7c23b61653ef87c6af1db5","name":"Mizuki Takenawa","hidden":false},{"_id":"6a7c23b61653ef87c6af1db6","name":"Kiyoharu Aizawa","hidden":false},{"_id":"6a7c23b61653ef87c6af1db7","name":"Toshihiko Yamasaki","hidden":false}],"publishedAt":"2026-08-09T00:00:00.000Z","submittedOnDailyAt":"2026-08-12T00:00:00.000Z","title":"360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents","submittedOnDailyBy":{"_id":"6527b37c0ae663e384eb1b85","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6527b37c0ae663e384eb1b85/zKWa8h6YU4BWfcitpM5Pl.png","isPro":false,"fullname":"Atsuyuki Miyai","user":"AtsuMiyai","type":"user","name":"AtsuMiyai"},"summary":"We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Existing outdoor benchmarks either lack sufficient photorealism or complexity, resulting in a considerable gap from real-world urban environments. 360CityArena is built on a realistic reconstruction of the Akihabara district in Tokyo, Japan, using 602 360-degree video segments covering 85 streets, and consists of 175 meticulously human-crafted tasks. It encompasses three task categories: Environment Understanding, Path Reasoning, and Spatial Reasoning, covering fundamental abilities required for urban exploration, such as localization, landmark search, path planning, and relational spatial reasoning, thereby enabling comprehensive evaluation in realistic urban scenes. Our evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level (human: 77.3% vs. Gemini 2.5 Flash: 17.1%), revealing substantial challenges that remain in city-scale embodied navigation and reasoning. 360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning.","upvotes":4,"discussionId":"6a7c23b71653ef87c6af1db8","projectPage":"https://360mm-team.github.io/360CityArena/","githubRepo":"https://github.com/360MM-Team/360CityArena","githubRepoAddedBy":"user","ai_summary":"A new photorealistic urban benchmark reveals large performance gaps for embodied agents in city-scale navigation and spatial reasoning.","ai_keywords":["embodied agents","360-degree videos","photorealistic environment","LMM-based agents","spatial reasoning","path planning","landmark search","city-scale embodied navigation"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":10,"organization":{"_id":"6730c40cfab94f649285ded2","name":"hal-utokyo","fullname":"Hal Lab UTokyo","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65d86a59bdb95b4bbc744e10/FOjSKBb6nGV-BBRp9fQU5.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6527b37c0ae663e384eb1b85","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6527b37c0ae663e384eb1b85/zKWa8h6YU4BWfcitpM5Pl.png","isPro":false,"fullname":"Atsuyuki Miyai","user":"AtsuMiyai","type":"user"},{"_id":"69bcbd47dfac620f178cc18e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/82dzAMoOedlV8XLDpNhMe.png","isPro":false,"fullname":"가은 문","user":"harperjsxz68","type":"user"},{"_id":"6615d9cb51fbbd38cc65f3eb","avatarUrl":"/avatars/63f0626c4728ab67a3095446a1476787.svg","isPro":false,"fullname":"So Kuroki","user":"Kuroki1931","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6730c40cfab94f649285ded2","name":"hal-utokyo","fullname":"Hal Lab UTokyo","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65d86a59bdb95b4bbc744e10/FOjSKBb6nGV-BBRp9fQU5.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.08814.md","query":{}}">
Papers
arxiv:2608.08814

360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents

Published on Aug 9
· Submitted by
Atsuyuki Miyai
on Aug 12
Authors:
,

Abstract

A new photorealistic urban benchmark reveals large performance gaps for embodied agents in city-scale navigation and spatial reasoning.

We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Existing outdoor benchmarks either lack sufficient photorealism or complexity, resulting in a considerable gap from real-world urban environments. 360CityArena is built on a realistic reconstruction of the Akihabara district in Tokyo, Japan, using 602 360-degree video segments covering 85 streets, and consists of 175 meticulously human-crafted tasks. It encompasses three task categories: Environment Understanding, Path Reasoning, and Spatial Reasoning, covering fundamental abilities required for urban exploration, such as localization, landmark search, path planning, and relational spatial reasoning, thereby enabling comprehensive evaluation in realistic urban scenes. Our evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level (human: 77.3% vs. Gemini 2.5 Flash: 17.1%), revealing substantial challenges that remain in city-scale embodied navigation and reasoning. 360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning.

Community

Paper submitter about 12 hours ago

🚀 #ECCV2026 𝐈𝐧𝐭𝐫𝐨𝐝𝐮𝐜𝐢𝐧𝐠 𝟑𝟔𝟎𝐂𝐢𝐭𝐲𝐀𝐫𝐞𝐧𝐚 🌆🧭

"Can AI truly understand and navigate a real city?"

We introduce 360CityArena, a realistic urban navigation benchmark built from 360° videos of Akihabara, Tokyo. 🇯🇵
🏙️ Our arena:
✅ 602 real-world 360° videos*
✅ 85 streets*
✅ 175 navigation & spatial reasoning tasks

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.08814
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.08814 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.08814 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.08814 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers