Hugging Face Daily Papers · · 4 min read

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<strong>A benchmark for health reasoning over real-world wearable data.</strong></p>\n<p>WearableQA comprises <strong>4,084 ten-option multiple-choice questions</strong> built from the wearable time series, blood biomarkers, and demographics of <strong>200 real users</strong>, each with up to about 500 days of daily measurements. Unlike benchmarks built on synthetic or idealized signals, it preserves authentic wearable distributions — device noise, missing days, and inter-individual variability included.</p>\n","updatedAt":"2026-09-10T05:02:18.736Z","author":{"_id":"62d3ab1f95806c44fe062189","avatarUrl":"/avatars/0b318269d62d2d9d65b3788cfbe586a4.svg","fullname":"Ji Soo Lee","name":"simplecloud","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8876590132713318},"editors":["simplecloud"],"editorAvatarUrls":["/avatars/0b318269d62d2d9d65b3788cfbe586a4.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.05405","authors":[{"_id":"6aa04ef3d0174964227beb34","name":"Ji Soo Lee","hidden":false},{"_id":"6aa04ef3d0174964227beb35","name":"Xilun Chen","hidden":false},{"_id":"6aa04ef3d0174964227beb36","name":"Pierce Chuang","hidden":false},{"_id":"6aa04ef3d0174964227beb37","name":"Ashish Shenoy","hidden":false},{"_id":"6aa04ef3d0174964227beb38","name":"Jason Wei","hidden":false},{"_id":"6aa04ef3d0174964227beb39","name":"Dohwan Ko","hidden":false},{"_id":"6aa04ef3d0174964227beb3a","name":"Hyunwoo J. Kim","hidden":false},{"_id":"6aa04ef3d0174964227beb3b","name":"Benoit Corda","hidden":false}],"publishedAt":"2026-09-04T00:00:00.000Z","submittedOnDailyAt":"2026-09-10T00:00:00.000Z","title":"WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data","submittedOnDailyBy":{"_id":"62d3ab1f95806c44fe062189","avatarUrl":"/avatars/0b318269d62d2d9d65b3788cfbe586a4.svg","isPro":true,"fullname":"Ji Soo Lee","user":"simplecloud","type":"user","name":"simplecloud"},"summary":"Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, each with up to 500 days of daily measurements. WearableQA preserves authentic wearable distributions that include device noise and inter-individual variability. To evaluate distinct reasoning capabilities, we introduce 16 question types organized along two complementary axes: data versus health reasoning, which distinguishes computation over longitudinal measurements from physiological interpretation; and single- versus cross-signal reasoning, which separates reasoning about individual signals from the integration of multiple signals. To construct reliable questions at scale, we adopt a dual-grounding framework that combines literature-grounded physiological findings with statistically validated population-grounded physiological patterns. This enables the capture of meaningful relationships observed in real-world wearable data. Evaluation of 14 proprietary and open-source LLMs demonstrates that WearableQA effectively differentiates model capabilities, with performance ranging from 19.6% to 72.9% against a 10% chance baseline. Moreover, WearableQA remains far from solved: most models achieve accuracies below 60%. Overall, WearableQA provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.","upvotes":24,"discussionId":"6aa04ef4d0174964227beb3c","githubRepo":"https://github.com/facebookresearch/WearableQA","githubRepoAddedBy":"user","ai_summary":"WearableQA is a benchmark of multiple-choice questions derived from real longitudinal wearable data that evaluates large language model reasoning across data and health dimensions.","ai_keywords":["LLMs","cross-signal reasoning","dual-grounding framework"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":6,"organization":{"_id":"6a78cbce04f58391d2d7c93b","name":"meta","fullname":"Meta","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Yq_gbM7i78tR3nfEvcwSA.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"62d3ab1f95806c44fe062189","avatarUrl":"/avatars/0b318269d62d2d9d65b3788cfbe586a4.svg","isPro":true,"fullname":"Ji Soo Lee","user":"simplecloud","type":"user"},{"_id":"64b4c344ee7a5f18251d11ac","avatarUrl":"/avatars/8fabaa335d84388524519e8758978c6d.svg","isPro":false,"fullname":"Jinyoung Kim","user":"jinyoungkim","type":"user"},{"_id":"685128512da8b6bcf5c24660","avatarUrl":"/avatars/216fe6a2719c0b744bcc25684c1cacb5.svg","isPro":false,"fullname":"jinwoo seo","user":"sjwoo0612","type":"user"},{"_id":"64c08117f6d4c1e5cf8ed697","avatarUrl":"/avatars/a8c21d4f477913b0987cd832c4ffafca.svg","isPro":false,"fullname":"Seunghun Lee","user":"Lemoni","type":"user"},{"_id":"665ef6e0a319eefe295d66a0","avatarUrl":"/avatars/a5f406b15c86c17d806b37627528f06a.svg","isPro":false,"fullname":"SeungminYun","user":"Seungmin12","type":"user"},{"_id":"6aa0ffb9e2c43beda015fb13","avatarUrl":"/avatars/718688cbe6a1693272a499c44600ce32.svg","isPro":false,"fullname":"Sojin Lee","user":"sojinleeme","type":"user"},{"_id":"64549b4cc13cdb83f1100b94","avatarUrl":"/avatars/4e809bed509e6ce2dd5532730af81a12.svg","isPro":false,"fullname":"Minseong Bae","user":"KyleBae1017","type":"user"},{"_id":"69d39c2704bf200050b58d79","avatarUrl":"/avatars/5b2f6605c37565228b67f62f04aa82f2.svg","isPro":false,"fullname":"Kyujin Lee","user":"KyujinL","type":"user"},{"_id":"662b7fee3bdfe519489403d0","avatarUrl":"/avatars/5545ebff1cde05a3cbd76efe3ae4c07c.svg","isPro":true,"fullname":"minseok joo","user":"zoomkinseok","type":"user"},{"_id":"6667c2c4934b48c3c79ed535","avatarUrl":"/avatars/94e2b747deabf6f206a50eb9727412da.svg","isPro":false,"fullname":"sehyung","user":"sehyungkim","type":"user"},{"_id":"671b68eaf50df4c3c72ea53b","avatarUrl":"/avatars/c33d8baffab4e4d715c5a6aa3a85b55b.svg","isPro":false,"fullname":"Jaewonchu","user":"allonsy07","type":"user"},{"_id":"69968f51e4f5126d0e1f4e66","avatarUrl":"/avatars/d6a57b95038dd511d76d560350d73af1.svg","isPro":false,"fullname":"Yejun Ju","user":"dpwns99","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a78cbce04f58391d2d7c93b","name":"meta","fullname":"Meta","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Yq_gbM7i78tR3nfEvcwSA.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.05405.md","query":{}}">
Papers
arxiv:2609.05405

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

Published on Sep 4
· Submitted by
Ji Soo Lee
on Sep 10
Authors:
,

Abstract

WearableQA is a benchmark of multiple-choice questions derived from real longitudinal wearable data that evaluates large language model reasoning across data and health dimensions.

Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, each with up to 500 days of daily measurements. WearableQA preserves authentic wearable distributions that include device noise and inter-individual variability. To evaluate distinct reasoning capabilities, we introduce 16 question types organized along two complementary axes: data versus health reasoning, which distinguishes computation over longitudinal measurements from physiological interpretation; and single- versus cross-signal reasoning, which separates reasoning about individual signals from the integration of multiple signals. To construct reliable questions at scale, we adopt a dual-grounding framework that combines literature-grounded physiological findings with statistically validated population-grounded physiological patterns. This enables the capture of meaningful relationships observed in real-world wearable data. Evaluation of 14 proprietary and open-source LLMs demonstrates that WearableQA effectively differentiates model capabilities, with performance ranging from 19.6% to 72.9% against a 10% chance baseline. Moreover, WearableQA remains far from solved: most models achieve accuracies below 60%. Overall, WearableQA provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.

Community

Paper submitter about 3 hours ago

A benchmark for health reasoning over real-world wearable data.

WearableQA comprises 4,084 ten-option multiple-choice questions built from the wearable time series, blood biomarkers, and demographics of 200 real users, each with up to about 500 days of daily measurements. Unlike benchmarks built on synthetic or idealized signals, it preserves authentic wearable distributions — device noise, missing days, and inter-individual variability included.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.05405
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.05405 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.05405 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers