研究库 论文知识库
Papers · organized/paper_cards

论文

274 张论文卡片 · 评测集 · OA 绿色

开放获取 全部 绿色 · 1640
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
MultiRef-Compass:迈向多参考音视频生成任务的综合评估
arXiv:2607.14189 多模态 评测集 OA · 绿色 被引 6 · S2

本文提出 MultiRef-Compass,一个面向 MR2AV 生成的统一基准,将自动指标与引入复判增强的 MLLM-as-a-Judge 框架相结合,实现对感知保真度与参考条件合成能力的可扩展、可审计评估。MultiRef-Compass is introduced, a unified benchmark for MR2AV generation that integrates automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework, enabling scalable and auditable evaluation of both perceptual fidelity and reference-conditioned composition.

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
VIABench:面向视障辅助任务的盲人视频综合基准
arXiv:2607.14660 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 VIABench,一个专为评估 MLLM 在视障辅助(VIA)场景中表现而设计的综合视频基准,采用视障人士(VIIs)自行录制或分享的第一人称视频,并提出一套严格的评测流水线,同时支持在线(实时)与离线设置。VIABench is introduced, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves, and proposes a rigorous benchmarking pipeline that supports both online (real-time) and offline settings.

Rethinking the Evaluation of Harness Evolution for Agents
Rethinking the Evaluation of Harness Evolution for Agents(重新审视 Agent 的 Harness 演化评估)
arXiv:2607.12227 评测基准 评测集 OA · 绿色 被引 45 · S2

重新审视自动 harness 演化流程的评估方法,强调需要采用匹配预算的基线(matched-budget baselines)和留出评估(held-out evaluation),以区分真正的 harness 改进与针对特定基准的搜索和过拟合。The evaluation of automatic harness evolution procedures is revisited and the need for matched-budget baselines and held-out evaluation to distinguish genuine harness improvements from benchmark-specific search and overfitting is highlighted.

Sparks of Artificial General Intelligence: Early experiments with GPT-4
Sparks of Artificial General Intelligence: Early experiments with GPT-4
arXiv:2303.12712 评测基准 评测集 OA · 绿色 被引 4491 · S2

认为(该早期版本的)GPT-4 属于新一代具备更通用智能的 LLM(与 ChatGPT、谷歌 PaLM 等并列),并讨论了这些模型不断增强的能力及其影响。It is argued that (this early version of) GPT-4 is part of a new cohort of LLMs (along with ChatGPT and Google's PaLM for example) that exhibit more general intelligence than previous AI models, and the rising capabilities and implications of these models are discussed.

A Human-Centric Evaluation of a Retrieval-Augmented Generation System for Explaining Quebec Insurance Contracts
面向魁北克保险合同解释的RAG系统的人类中心评估
arXiv:2607.15963 RAG 检索增强 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

对一个旨在使魁北克汽车保险合同更易理解的SOTA RAG系统进行以人为中心的外在评估,结果显示该系统被视为认知均衡器,用户对系统所提供的自主感的重视程度甚至超过知识本身。A human-centric, extrinsic evaluation of a state-of-the-art Retrieval-Augmented Generation system, designed to make Quebec automobile insurance contracts more understandable, shows the system is perceived as a cognitive equalizer, and users value the sense of autonomy the system provides even more than the knowledge itself.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Beyond the Imitation Game:语言模型能力的量化与外推
arXiv:2206.04615 评测基准 评测集 OA · 绿色 被引 2679 · S2

在 BIG-bench 上对 OpenAI 的 GPT 模型、Google 内部稠密 Transformer 架构及 Switch 风格稀疏 Transformer 进行评估,模型规模跨越百万至千亿参数,结果显示性能与校准均随规模提升而改善,但绝对水平仍然欠佳。Evaluation of OpenAI's GPT models, Google-internal dense transformer architectures, and Switch-style sparse transformers on BIG-bench, across model sizes spanning millions to hundreds of billions of parameters finds that model performance and calibration both improve with scale, but are poor in absolute terms.

ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
ScanNet:富含标注的室内场景三维重建
arXiv:1702.04405 多模态 评测集 OA · 绿色 被引 6002 · S2

本文推出 ScanNet,一个 RGB-D 视频数据集,包含 1513 个场景中的 250 万视角,标注有三维相机位姿、表面重建与语义分割,并表明使用该数据可在多项三维场景理解任务上取得 SOTA 性能。This work introduces ScanNet, an RGB-D video dataset containing 2.5M views in 1513 scenes annotated with 3D camera poses, surface reconstructions, and semantic segmentations, and shows that using this data helps achieve state-of-the-art performance on several 3D scene understanding tasks.

Capabilities of GPT-4 on Medical Challenge Problems
GPT-4 在医学挑战性问题上的能力
arXiv:2303.13375 评测基准 评测集 OA · 绿色 被引 1451 · S2

对 SOTA LLM GPT-4 在医学能力考试与基准数据集上进行全面评估,并通过案例研究定性探索其行为,展示了 GPT-4 解释医学推理、为学生定制个性化讲解以及围绕病例交互式构造新反事实场景的能力。A comprehensive evaluation of GPT-4, a state-of-the-art LLM, on medical competency examinations and benchmark datasets and explores the behavior of the model qualitatively through a case study that shows the ability of G PT-4 to explain medical reasoning, personalize explanations to students, and interactively craft new counterfactual scenarios around a medical case.

A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
ChatGPT 在推理、幻觉与交互性方面的多任务、多语言、多模态评估
arXiv:2302.04023 多模态 评测集 OA · 绿色 被引 1839 · S2

研究发现 ChatGPT 在大多数任务上以零样本学习优于其他 LLM,在部分任务上甚至超过微调模型,并且对非拉丁文字语言的理解能力优于生成能力。It is found that ChatGPT outperforms LLMs with zero-shot learning on most tasks and even outperforms fine-tuned models on some tasks and is better at understanding non-Latin script languages than generating them.

Matterport3D: Learning from RGB-D Data in Indoor Environments
Matterport3D:基于室内 RGB-D 数据的学习
arXiv:1709.06158 多模态 评测集 OA · 绿色 被引 2726 · S2

本文介绍 Matterport3D,一个大规模 RGB-D 数据集,包含来自 90 个建筑物级场景共 194,400 张 RGB-D 图像的 10,800 个全景视图,可支持多种监督与自监督计算机视觉任务,包括关键点匹配、视角重叠预测、由彩色图像预测法线、语义分割和区域分类。Matterport3D is introduced, a large-scale RGB-D dataset containing 10,800 panoramic views from 194,400RGB-D images of 90 building-scale scenes that enable a variety of supervised and self-supervised computer vision tasks, including keypoint matching, view overlap prediction, normal prediction from color, semantic segmentation, and region classification.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
超越成功率:面向攻击性与防御性安全 Agent 的成本感知评估
arXiv:2607.15263 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

本文认为安全 Agent 基准应在任务成功率之外,同时衡量经济效率与运维适配性,并提出成本感知、SOC 原生的评估方法,以更清晰地反映当前哪些模型具有实际可用价值,以及防御性 Agent 仍需改进的方向。It is argued that security-agent benchmarks should measure economic efficiency and operational fit alongside task success alongside task success, and cost-aware, SOC-native evaluations provide a clearer picture of which models are practically useful today and where defensive agents still need to improve.

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning
SVR-R1:在强化学习中通过自验证自举多模态推理
arXiv:2607.10966 多模态 评测集 OA · 绿色 被引 1 · S2

本文提出自验证推理器 SVR-R1,一种多轮强化学习框架,将模型自身的验证转化为多模态推理的学习信号,提供了一种简洁而有效的多模态推理自举方案。Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification into a learning signal for multimodal reasoning, is introduced, offering a simple yet effective recipe for bootstrapping multimodal reasoning.

EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World
EvolvingWorld:用于交互式文学世界中角色扮演 Agent 与世界模型协同进化的开放模式框架
arXiv:2607.17250 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

实验表明,EvolvingWorld 能够通过有效维持持久且一致的角色与世界发展,提升长程模拟能力。Experiments show that EvolvingWorld can improve long-horizon simulation by effectively maintaining persistent, coherent character and world development.

Can Multimodal Large Language Models Understand OCT?
多模态大语言模型能理解OCT吗?
arXiv:2607.16609 多模态 评测集 OA · 绿色 被引 3 · S2

OCT-Bench能够对MLLM进行全面且细粒度的评估,为识别能力瓶颈和推进临床可信的OCT理解奠定基础。OCT-Bench enables comprehensive and fine-grained evaluation of MLLMs, providing a foundation for identifying capability bottlenecks and advancing clinically grounded OCT understanding.

GigaChat Audio: Time-aware Large Audio Language Model
GigaChat Audio:时间感知的大音频语言模型
arXiv:2607.10387 多模态 评测集 OA · 绿色 被引 1 · S2

本文提出一个时间感知的音频 LLM,能够基于大规模合成监督(来自级联 pipeline)在长达 120 分钟的输入上回答带有显式时间戳的问题,并在短时长和长时长 benchmark 上取得强劲的时间定位准确率。This work presents a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input using large-scale synthetic supervision from a cascaded pipeline and achieves strong temporal-grounding accuracy on short and long benchmarks.

VQA: Visual Question Answering
VQA: Visual Question Answering
arXiv:1505.00468 多模态 评测集 OA · 绿色 被引 6749 · S2
Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints
Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints
arXiv:2607.18144 评测基准 评测集 OA · 绿色 被引 2 · S2

研究揭示了 LLM 空间能力中的清晰规律:尽管其仍落后于 SOTA 方法,但具有潜力并能同时处理多种空间约束,从而可扩展到异构场景。A clear pattern is revealed in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
arXiv:2607.15434 Agent 智能体 评测集 OA · 绿色 被引 2 · S2

提出 Manager Coercion Benchmark:被测管理者需要完成一项良性任务并有完成的动机,但唯一能礼貌且坚定地拒绝的智能体本身就是被测管理者自己。The Manager Coercion Benchmark is introduced: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines, but the only agent that can do it politely and immovably declines is the manager under test itself.

Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes
基于 LangGraph 的图结构 Agent AI:面向长时运行、有状态业务流程的工作流路径
arXiv:2607.19297 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文是面向业务流程中长时间运行、有状态、多步生成式 AI 系统的基于图的工作流路径实践指南,并通过三个可执行示例展示类型化状态、条件路由、确定性工具、重试、中断、检查点与 trace 如何协同工作。This paper is a practitioner guide to graph-based workflow pathways for long-running, stateful, multi-step generative AI systems in business processes and presents three executable recipes to show how typed state, conditional routing, deterministic tools, retries, interrupts, checkpoints, and traces fit together.

RF-Agent: A Practical Framework for Building Language Agents for RFIC Design
RF-Agent:面向 RFIC 设计的 Language Agent 构建实践框架
arXiv:2607.18772 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

本文提出 RF-Agent,通过多 Agent 的 Question-Thinking-Solution-Answer 流水线,基于教材驱动的知识蒸馏来弥补 RF 领域专用推理的空白,为面向 LLM 辅助 RF 电路设计的未来工作提供了可复用的基础。RF-Agent is presented, which addresses the gap in domain-specific RF reasoning through textbook-driven knowledge distillation through a multi-agent Question-Thinking-Solution-Answer pipeline and provides a reusable foundation for future work on LLM-aided RF circuit design.

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
EduPanel:面向教学视频的三 Agent LLM 评判框架——可靠性、互补性与人类信任校准
arXiv:2607.18529 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

EduPanel 是一个基于评分量表、以学习者为条件的 LLM 评判器,通过在多个专用 agent 间分解评估流程,对教学质量的各个方面产出可解释的评估结果,其可靠性与中等水平的人类专家相当。E EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable assessments for different aspects of teaching quality, achieves reliability comparable to a median human expert.

An Exam for Active Observers
面向主动观察者的评测
arXiv:2607.16165 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

人类视觉是一个闭环:注视点不断被中间假设而非单一快照持续重定向。数十年的心理物理学与认知科学研究表明,主动观察对多种任务至关重要。当代多模态大语言模型 (MLLM) 是否进行主动观察,是一个现有视觉语言基准无法回答的经验问题。我们提出 ActiveVision,一个使 MLLM 主动观察可度量的基准,包含 3 个类别共 17 个任务,任务设计强制进行重复视觉感知……Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment
通过表征锚定与语言-动作对齐的可泛化 VLA 微调
arXiv:2607.13429 工程化 评测集 OA · 绿色 被引 4 · S2

本文提出 Anchor-Align,通过两个目标增强 BC:Vision-Language Anchoring 从冻结 VLM 副本中蒸馏逐层表示以防止该漂移;Language-Action Alignment 将每个动作目标转换为离散的运动方向标签,并在同一机器人观测上联合训练语言与动作预测。Anchor-Align is proposed, which augments BC with two objectives: Vision-Language Anchoring distills layer-wise representations from a frozen VLM copy to prevent this drift, and Language-Action Alignment converts each action target into a discrete motion-direction label and jointly trains language and action prediction on the same robot observation.

DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
DocOps:面向复杂文档操作的自主 Agent 可验证基准
arXiv:2607.19865 评测基准 评测集 OA · 绿色 被引 1 · S2

提出 DocOps,一种确定性可验证的评估框架,基于分层分类法,将受真实实践启发的文档操作分解为原子维度与逐级递增的工作流复杂度,从而揭示 Agent 在维护全局文档一致性方面的能力边界。DocOps is introduced, a deterministically verifiable evaluation framework underpinned by a hierarchical taxonomy that deconstructs document operations inspired by real-world practices into atomic dimensions and escalating workflow complexities that exposes the capability boundaries of agents in maintaining global document consistency.

ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
arXiv:2607.20092 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ENTRAP-VL(ENTRainment Assessment Probe for Vision and Language),一个由人工策展的 1,500 条数据的数据集,涵盖八个类别,按一个跨双轴的分类体系组织,并划分为文本诱发流和视觉诱发流。ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes and split into a textual-entrainment stream and a visual-entrainment stream, is introduced.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
展示而非讲述:在生成像素而非 LLM 文本中评估空间认知
arXiv:2607.21072 评测基准 评测集 OA · 绿色 被引 2 · S2

提出 ProVisE(Protocolized Visual Evaluation),一个与基准无关的框架,通过受协议约束的视觉问答从图像生成模型中抽取答案,并将其解析为与原始指标兼容的结构化预测,揭示了像素空间表达与基于文本推理的互补优势。ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics, is proposed and revealed, revealing complementary strengths of pixel-space expression and text-based reasoning.

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
K12-KGraph:用于教育 LLM 基准测试与训练的课程对齐知识图谱
arXiv:2605.09635 评测基准 评测集 OA · 绿色 被引 2 · S2

本文提出 K12-KGraph,一个从人民教育出版社官方教材中提取的、与课程对齐的知识图谱,覆盖小学、初中和高中阶段的数学、物理、化学和生物,并由此衍生出 K12-Bench,一个包含 23,640 道题的多选基准,涵盖五类任务族,表明文本监督与视觉监督具有互补性。This work introduces K12-KGraph, a curriculum-aligned knowledge graph extracted from official People's Education Press textbooks in mathematics, physics, chemistry, and biology across primary, middle, and high school, and derives K12-Bench, a 23,640-question multi-select benchmark with five task families, showing that textual and visual supervision are complementary.

FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents
FinanceComplexQA:面向工业级金融文档的智能体推理基准测试
arXiv:2607.19238 Agent 智能体 评测集 OA · 绿色 被引 2 · S2

本文设计 Finance-LaTeX SKILL,一个基于专家知识合成复杂版面金融文档的 skill,并提出 FinanceComplexQA,一个全面、贴近真实场景的金融文档开放式生成基准。This work designs Finance-LaTeX SKILL, a skill for synthesizing financial documents with complex layouts based on expert knowledge, and introduces FinanceComplexQA, a comprehensive open-ended generation benchmark for financial documents that closely resembles real-world scenarios.

Mathematical Capabilities of ChatGPT
ChatGPT 的数学能力
arXiv:2301.13867 评测基准 评测集 OA · 绿色 被引 621 · S2

研究发现,ChatGPT 最成功地被用作数学助手,用于查询事实、充当数学搜索引擎和知识库接口;GPT-4 还可被用于本科水平的数学问题,但在研究生难度的题目上表现不佳。It is found that ChatGPT can be used most successfully as a mathematical assistant for querying facts, acting as a Mathematical search engine and knowledge base interface, and GPT-4 can additionally be used for undergraduate-level mathematics but fails on graduate-level difficulty.

Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
ChatGPT 生成的代码真的正确吗?面向代码生成的大型语言模型严格评估
arXiv:2305.01210 评测基准 评测集 OA · 绿色 被引 2333 · S2

EvalPlus——一个用于严格基准测试 LLM 生成代码功能正确性的代码合成评估框架,通过 LLM 与基于 mutation 的策略驱动的自动测试输入生成器,为给定评估数据集补充大量新生成的测试用例。EvalPlus -- a code synthesis evaluation framework to rigorously benchmark the functional correctness of LLM-synthesized code and augments a given evaluation dataset with large amounts of test-cases newly produced by an automatic test input generator, powered by both LLM and mutation-based strategies.

Teachy Mini: Development and Preliminary Evaluation of a Knowledge-Based Generative Social Robot for Higher Education
Teachy Mini:面向高等教育的知识驱动生成式社交机器人开发与初步评估
arXiv:2607.22345 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本研究在 Reachy Mini 机器人平台上,通过系统提示、检索增强生成 (RAG) 和有状态的提示编排,将选定的 KBD 需求落地实现,表明 KBD 可以塑造负责任的机器人行为,并有望提升机器人辅助学习中的学习效果。This study operationalized selected KBD requirements in the Reachy Mini robot platform through system prompting, retrieval-augmented generation, and stateful prompt orchestration, indicating that KBD can shape responsible robot behavior and potentially increase learning effectiveness in robot-supported learning.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
DataPrep-Bench: 将 LLM 作为训练数据准备器的基准测试
arXiv:2607.20465 工程化 评测集 OA · 绿色 被引 1 · S2

提出 DataPrep-Bench,首个统一基准,在共享的下游任务 grounding 协议下,对 LLM 驱动的数据准备在六个领域、多种 base model 上的两类能力进行联合评估。DataPrep-Bench is introduced, the first unified benchmark that jointly evaluates both capabilities under a shared downstream-grounded protocol over six domains and multiple base models of LLM-driven data preparation.

Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
无坐标与区域标签的视觉文档理解中的证据归因
arXiv:2607.24651 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本研究考察是否存在一条实用路径,在没有坐标界面、且无需高成本区域级监督的条件下提升归因效果,并指出了这样一条可行路径。A study investigates whether there is a practical path to improve attribution without a coordinate interface and without costly region-level supervision, and indicates a practical path to improve attribution without a coordinate interface and without costly region-level supervision.

Characterizing Warp Divergence from Pascal to Blackwell
从 Pascal 到 Blackwell 的 Warp 分歧特性分析
arXiv:2607.23402 评测基准 评测集 OA · 绿色 被引 1 · S2

即使 NVIDIA 控制流 ISA 与重汇聚(reconvergence)机制持续演进,Divergence 仍能保持稳定且可预期的性能开销。Divergence retains a stable and predictable performance cost even as NVIDIA's control-flow ISA and reconvergence mechanisms continue to evolve.

IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages
IndicTalk:面向印度语系的大规模人格化多语言对话语料库
arXiv:2607.23242 LLM 基础设施 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 IndicTalk,目前最大的多语言印度语码混合(code-mixed)会话语料库,包含超过 13,28,604 段事件驱动的多轮对话,覆盖 9 种印度语言的 18 种语言变体,将公开发布以支持低资源印度语多语言会话 AI 的开发与评估。IndicTalk is presented, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages and will be released to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages.

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
Keep It InMind:基准测试 Agent 记忆中的隐式关联盲点
arXiv:2607.24368 评测基准 评测集 OA · 绿色 被引 1 · S2

本文提出 InMind,一项包含 125 个任务、经专家校验、跨越十个生活领域的基准,其中 113 个任务基于可引用的公开来源;本文将此类失效模式命名为 implicit-association blind spot,并提出一种极简诊断探针(diagnostic probe),在 query 到达前保持 memory 可见即可恢复大部分性能差距。This work introduces InMind, a 125-task, expert-verified benchmark spanning ten life domains, with 113 tasks grounded in citable public sources, and calls this failure mode the implicit-association blind spot, and introduces a minimal diagnostic probe that keeps memory visible before the query arrives recovers most of the gap.