WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
- 类型:arxiv
- 标识:2609.05405
- 链接:https://arxiv.org/abs/2609.05405
- 主分类:evaluation
- 形态:benchmark
- TLDR:Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, each with up to 500 days of daily measurements. WearableQA preserves authentic wearable distributions that include device noise and inter-individual variability. To evaluate distinct reaso
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-10-agent-rag-longcontext-candidates.json