研究库 论文知识库
Papers · organized/paper_cards

论文

1753 张论文卡片

开放获取 全部 绿色 · 1640
Improving Health Literacy through Lay Summarization of Radiological Reports: An Evaluation of BioNER and Retrieval-Augmented Generation
通过放射学报告的大众化摘要提升健康素养:BioNER 与检索增强生成的评估
arXiv:2609.02396 RAG 检索增强 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本研究探讨了相较于标准 LLM 生成方式,RAG 与命名实体识别在多大程度上能提升自动生成通俗摘要的质量、事实一致性及可读性。This study investigates the extent to which Retrieval-Augmented Generation and Named Entity Recognition improve the quality, factual consistency, and readability of automatically generated lay summaries compared with standard LLM-based generation.

NE-R1: Enhancing Named Entity Recognition Model via Reinforcement Learning
NE-R1:通过强化学习增强命名实体识别模型
arXiv:2609.02366 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 NE-R1,一种面向自适应检索增强 NER 的新框架,在多个基准上达到 SOTA 性能,域内评估平均 F1 提升 2.52%,零样本跨域评估平均 F1 提升 1.18%。This paper proposes NE-R1, a novel framework for adaptive retrieval-augmented NER, which achieves state-of-the-art performance on various benchmarks, with an average F1 score gain of 2.52% in in-domain evaluation and 1.18% in zero-shot cross-domain evaluation.

Exploring Collaboration between a language and a non-language agent
探索语言智能体与非语言智能体之间的协作
arXiv:2609.00474 Agent 智能体 应用落地 OA · 绿色 被引 1 · S2

为解决 LLM 与非语言 Agent 的协作问题,本文提出 latent state internalization,将子 Agent 的连续表示直接投射到 LLM 的 token 流中作为习得的状态 token,并随着动作推进环境状态而进行动态重编码。To solve LLM collaboration with non-language agents, latent state internalization is introduced, which projects the subagent's continuous representations directly into the LLM's token stream as learned state tokens, with dynamic re-encoding as actions advance the environment state.

NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference
NeoMME:用于高效微调与推理的单塔多模态原生多语言基础编码器
arXiv:2609.01657 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 NeoMME,一个 260M 与 800M 参数的多模态多语言双向编码器系列,可在单个双向 Transformer encoder 中处理多语言文本与原始图像 patch。This work introduces NeoMME, a family of 260M and 800M-parameter Multimodal and Multilingual bidirectional Encoders that process multilingual text and raw image patches in a single bidirectional Transformer encoder.

FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos
FoldingAgent:从演示视频中推断参数化折纸过程
arXiv:2609.00377 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 FoldingAgent,一个从折纸演示视频中推断显式参数化折纸程序的 Agent 框架,利用预训练 Vision-Language Model 的推理能力,并配备可模拟几何变换、验证物理合理性、检索与比较视觉内容以及评估自身预测的专用工具集。FoldingAgent is presented, an agentic framework for inferring explicit parametric folding programs directly from origami demonstration videos that leverages the reasoning power of a pre-trained Vision-Language Model equipped with a suite of specialized tools that enable the agent to simulate geometric transitions, verify physical plausibility, retrieve and compare visual content, and evaluate its own predictions.

Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents
无知还是无能?为 LLM 智能体构建知识门控的可验证任务
arXiv:2608.30322 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种知识门控的任务构建协议,将任务指令与一个包含私有约定、参考表与效用算子的紧凑工件分离,并证明被保留的任务能够改善 post-training。A knowledge-gated task-construction protocol is introduced that separates a task instruction from a compact artefact containing private conventions, reference tables, and utility operators, and it is shown that the retained tasks improve post-training.

5. GraphRAG / LLMs+Graphs 综合研究
arXiv:2606.11560 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本教程综合了推动这些汇聚方向的算法、系统与设计原则,为数据科学与数据挖掘研究者提供统一视角,涵盖将 LLM、图数据管理、图挖掘、图 ML 与 agentic 计算融合到下一代 graph-native AI 系统中。This tutorial synthesizes the algorithms, systems, and design principles driving these converging directions, offering data science and data mining researchers a unified perspective on integrating LLMs, graph data management, graph mining, graph ML, and agentic computation into next-generation graph-native AI systems.

Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering
超越视觉相似性:面向知识库视觉问答的实体对齐检索
arXiv:2608.21450 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 KBMR,首个面向 KB-VQA 的基于 MLLM 的 embedding retriever,并引入一个基于 MLLM 的语义判别器以生成连续的实体一致性权重,应对维基百科规模检索中的噪声监督挑战。KBMR is proposed, the first MLLM-based embedding retriever tailored for KB-VQA, and an MLLM-based semantic discriminator that generates continuous entity-consistency weights is introduced to tackle the challenge of noisy supervision in Wikipedia-scale retrieval.

Debias-SparseGPT: Bias-Aware Pruning for Large Language Models
Debias-SparseGPT:面向 LLM 的偏置感知剪枝
arXiv:2609.02496 工程化 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Debias-SparseGPT,一种 post-training 剪枝方法,通过在人口统计对比输入上定义的二阶项引入表示去偏好,在保持模型困惑度与零样本准确率的前提下,一致地降低剪枝带来的偏差,效果优于 SparseGPT。Debias-SparseGPT is introduced, a post-training pruning method incorporating representational debiasing using a second-order term defined over demographically contrasting inputs that consistently reduces pruning-induced bias compared to SparseGPT while preserving model perplexity and zero-shot accuracy.

Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens
Sparse Readout Prism:用特征而非 token 解释 Logit-Lens 分数
arXiv:2609.01936 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Sparse Readout Prism (SRP) 仅使用 readout 的权重对其进行分解,将任意 token logit 或 logit 差表示为来自稀疏 readout 特征贡献之和,揭示了 readout 特征作为 lens 解读新单元的价值,暴露出 token 身份可能掩盖的结构,并支持跨 token、上下文、层与 lens 的比较。Sarse Readout Prism (SRP), which decomposes the readout using only its weights and expresses any token logit or logit difference as a sum of contributions from sparse readout features, reveals readout features as a new unit of analysis for lens readings, exposing structure that token identities can obscure and enabling comparisons across tokens, contexts, layers, and lenses.

WHALE: A Simple Recipe for Joint Harness-Weight Optimization
WHALE:联合 Harness-权重优化的简洁方案
arXiv:2609.00196 评测基准 方法 OA · 绿色 被引 8 · S2

本文提出 Weight-Harness Alternating LEarning (WHALE),一种简单的方法,交替进行两个阶段:先在当前 harness 下更新模型,再在更新后的模型下通过在线拒绝采样微调与 Meta-Harness 搜索更优的 harness。Weight-Harness Alternating LEarning (WHALE), a simple recipe that alternates two phases: updating the model under the current harness, then searching for a better harness under the updated model with online rejection-sampling fine-tuning and Meta-Harness, is proposed.

Small Language Models as Judges for Rubric-Based Reinforcement Learning
小语言模型作为基于评分标准的强化学习评判器
arXiv:2608.30005 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文研究较小的语言模型能否作为高效且可靠的 rubric 评分器,并比较了从小模型中提取逐项判断的三种方式:生成式判定、Yes/No Logprob 边际以及探针判别器。This work studies whether smaller language models can serve as efficient and reliable rubric-based judges, and compares three ways of extracting criterion-level judgments from small models: Generative verdicts, Yes/No Logprob margins, and Probe judges.

Portfolio Risk Bounds without Cross-Asset Return Covariances: Distributional Fields from Language-Model Representations
无需跨资产收益协方差的投资组合风险边界:来自语言模型表征的分布场
arXiv:2608.29692 安全与风险 方法 被引 2 · S2

投资组合风险评估通常依赖可靠的跨资产收益协方差估计,而在短、高维面板中难以获得。我们表明公司层面的分布型特征可提供投资组合风险的单边证书。在"从特征到系统性暴露、从暴露到收益"的既定联系下,多公司 Wasserstein-2 离散度对系统性投资组合方差给出紧上界,并对标准化收益给出相应边界。加权成对松弛产生一个目标函数……Portfolio risk assessment ordinarily relies on reliable estimates of cross-asset return covariances, which are difficult to obtain in short, high-dimensional panels. We show that firm-level distribution-valued characteristics can instead provide one-sided certificates of portfolio risk. Under maintained links from characteristics to systematic exposures and from exposures to returns, multi-firm Wasserstein-2 dispersion yields a sharp upper bound on systematic portfolio variance and a corresponding bound for standardized returns. A weighted pairwise relaxation produces an objective that is conv

Wasserstein-Barycentric Interaction Fields for Spatial Factor Models: Evidence from Language-Model Representations
空间因子模型的 Wasserstein-重心交互场:来自语言模型表征的证据
arXiv:2608.29669 RAG 检索增强 方法 被引 1 · S2

空间资产定价模型将企业间交互结构视为已知,并利用语言模型表征从企业的信息环境中推断该结构;语言模型表征充当资本市场中潜在企业间信息结构的测量工具。Spatial asset-pricing models take the structure of inter-firm interaction as given and infer that structure from firms' information environments using language-model representations, which serve as a measurement instrument for latent inter-firm information structure in capital markets.

An Empirical Study on Zero-Data Bootstrapping for Conversational Recommender Systems
面向对话推荐系统的零数据引导的实证研究
arXiv:2504.15476 多模态 方法 OA · 绿色 被引 7 · S2

结果表明:领域驱动的合成数据一致优于零样本提示与朴素合成基线;主动选择相比随机采样提升了数据效率;元数据与协同过滤信号各自提升选择质量;在低资源场景下,合成数据可优于稀缺的真实对话,并进一步对真实对话形成补充。The results show that domain-grounded synthetic data consistently outperforms zero-shot prompting and naive synthetic baselines; active selection improves data efficiency over random sampling; metadata and collaborative filtering signals each improve selection quality; and, in low-resource settings, synthetic data can outperform scarce real dialogues while further complementing them.

Extending concurrent separation logic to the hardware level to verify the xv6 OS kernel on RISC-V with AI agents
将并发分离逻辑扩展到硬件层面,借助 AI Agent 在 RISC-V 上验证 xv6 OS 内核
arXiv:2609.04043 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

证明了一个应用层定理:若用户在 UART 控制台输入 echo hello world,系统唯一能产生的输出即为 hello world;这证明了基于 LLM 的 Agent 能够对如此底层的细节进行推理。An application-level theorem is proved: if the user types echo hello world as input on the UART console, the only output the system can produce is hello world, which proves LLM-based agents are capable of reasoning about such low-level details.

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States
Puffin-World:以原生 3D 世界状态扩展统一多模态模型
arXiv:2609.04196 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Puffin-World,一种统一的多模态架构,集成物理理解、空间仿真与 3D 世界生成重建,无需依赖外部离线模块,可支持需要多任务协同的交错式闭环应用。Puffin-World is proposed, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules and enables interleaved closed-loop applications requiring synergy across multiple tasks.

4️⃣ arXiv · Securing the Agent: Vendor-Neutral, Multitenant Enterprise Retrieval and Tool Use(⭐⭐⭐⭐ 高优先级)
4️⃣ arXiv · 守护 Agent:厂商中立的多租户企业级检索与工具调用(⭐⭐⭐⭐ 高优先级)
arXiv:2605.05287 Agent 智能体 观点 OA · 绿色 被引 1 · S2

本文提出一种分层隔离架构,结合策略感知的 ingestion、retrieval-time gating 与共享推理,并通过服务端 agentic 编排加以执行,在为多租户隔离提供天然强制点的同时,允许客户端框架保留对 agent 组合与延迟敏感操作的控制权。A layered isolation architecture combining policy-aware ingestion, retrieval-time gating, and shared inference, enforced through server-side agentic orchestration is introduced, creating natural enforcement points for multitenant isolation while allowing client-side frameworks to retain control over agent composition and latency-sensitive operations.

Compile by Training: Turning Natural-Language Specifications into Local Neural Functions
Compile by Training:将自然语言规约转化为本地神经函数
arXiv:2609.04199 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

许多重复出现的文本函数易于描述却难以用规则实现;而为每个输入调用大型远程模型会带来重复开销、延迟与对服务方的依赖。我们提出 compile by training,将自然语言规约转化为可复用的神经函数。在编译时,教师模型生成任务专属样本,用于为精简解释器训练一个小型适配器。生成的函数可在没有教师模型的情况下运行,并能像普通软件一样被存储、版本化管理与组合。在 FuzzyBench-Hard 这一子集上……Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and dependency on a provider. We present compile by training, which turns a natural-language specification into a reusable neural function. At compile time, teacher models generate task-specific examples that are used to train a small adapter for a compact interpreter. The resulting function runs without the teachers and can be stored, versioned, and composed like ordinary software. On FuzzyBench-Hard, a subset on which

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding
超越检索:面向流式视频理解的渐进式潜在记忆演化
arXiv:2609.04131 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出 LatentStream,一种渐进式 latent working memory 框架,将流式记忆从"存储-检索"转变为"检索-内化",在现有在线和离线视频 benchmark 上取得新的 SOTA 结果。This work introduces LatentStream, a progressive latent working memory framework that shifts streaming memory from store-and-retrieve to retrieve-and-internalize, and achieves new state-of-the-art results on existing online and offline video benchmarks.

PACE: Towards Surfacing Hidden Conflicts in User Requests
PACE:揭示用户请求中的隐性冲突
arXiv:2609.03293 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 PACE 数据集,用于评估模型能否识别以自我中心知识或事件形式表达的潜在约束(这些约束使看似合理的用户请求变得不当),以及 PaceMaker 多 Agent 框架,其中专门 Agent 通过查询重构、多跳图遍历与冲突感知过滤进行协调,以检索上下文决定性证据。PACE is introduced, a dataset for evaluating whether models can identify latent constraints, expressed as egocentric knowledge or events, that render seemingly reasonable user requests inappropriate, and PaceMaker, a multi-agent framework in which specialized agents coordinate across query reformulation, multi-hop graph traversal, and conflict-aware filtering to retrieve contextually decisive evidence.

The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation
缺失的时间纽带:面向剧本驱动音视频生成的时间上下文路由
arXiv:2609.02367 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Temporal Context Routing,将脚本时序映射到视频与音频生成的共享时间轴上,并将每个 prompt 的引导路由到两种模态中的对应位置,同时保持与 baseline 相当的视觉质量与音视频同步性。Temporal Context Routing is introduced, which maps the script timing onto the shared temporal axis of video and audio generation and routes each prompt's guidance to the corresponding positions in both modalities, while maintaining visual quality and audio-visual synchronization comparable to those of the baselines.

Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration
改变置信度,而非改变预测:面向后验校准的预测保持型修复
arXiv:2609.01072 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 CORD,第一个 post-fit adapter,通过从 calibrator 拟合中移除 preservation constraint 来修复完整校准概率向量,从而实现精确的预测保持,并将原始决策的精确恢复交由后续输出修复完成。CORD is proposed, the first post-fit adapter to impose exact prediction preservation by repairing the full calibrated probability vector by removing the preservation constraint from calibrator fitting and assigns exact recovery of the original decision to subsequent output repair.

Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training
知道何时不复用:自主 LLM 后训练中的条件经验迁移
arXiv:2608.26730 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

将条件经验迁移进行形式化,并提出 Boundary-Calibrated Intervention Transfer,一种在权重变化的训练之前即授权经验复用的方法,在相同预算下取得比所评估替代方案更高的最终模型质量。Conditional experience transfer is formulated as conditional experience transfer and Boundary-Calibrated Intervention Transfer is introduced, a method that authorizes experience reuse before weight-changing training and attains higher equal-budget final-model quality than the evaluated alternatives.

Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction
Scal3R:学习高效多相对位姿查询以实现可扩展的在线 3D 重建
arXiv:2609.04201 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该方法将在线重建重构为多参考相对位姿查询,在 Virtual KITTI、Sintel、TUM-Dynamic、ScanNet 和 7-Scenes 上取得 SOTA 性能,并在单 GPU 上 8 小时内收敛。This approach reformulates online reconstruction as multi-reference relative pose querying, which achieves state-of-the-art performance across Virtual KITTI, Sintel, TUM-Dynamic, ScanNet, and 7-Scenes and reaches convergence in 8 hours on a single GPU.

Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
选择、压缩、再投入:长视频 MLLM 中视觉 token 分配的控制变量研究
arXiv:2609.03820 多模态 方法 被引 0 · S2

过程中发现 AKS baseline 自身存在实现 bug,且两个 harness 在相同预算下运行相同已发布规则存在 0.74 分的差距,这表明此类对比应在同一受控 harness 内进行,而非跨论文比较。Along the way, an implementation bug in the own AKS baseline and a 0.74 point gap between two harnesses running the same published rules at the same budget show why these comparisons need to happen inside one controlled harness rather than across papers.

Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance
优化中的渗流动力学:方差级联与离散尺度不变性
arXiv:2609.02373 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

通过将随机梯度流建模为渗流过程来研究随机梯度下降的动态,其中嵌套的架构对称性迫使子网络以离散块的形式合并,而非通过单边附着。The dynamics of Stochastic Gradient Descent is studied by modeling the stochastic gradient flow as a percolation process, in which nested architectural symmetries force subnetworks to merge in discrete blocks rather than by single-edge attachment.

Using Grounded Theory for Agent Behavior Analysis at Scale
大规模 Agent 行为分析中的扎根理论应用
arXiv:2608.30391 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出 AutoTraceGT(Automated Trace analysis through Grounded Theory),首个在 Agent 轨迹上自动化 grounded theory 的多 Agent pipeline,并指出 Grounded Theory 为研究 Agent 实际行为的 ML 研究者和 Agent 开发者提供了可扩展的分析工具。This work proposes AutoTraceGT (Automated Trace analysis through Grounded Theory), the first multi-agent pipeline that automates grounded theory on agent trajectories and suggests Grounded Theory offers a scalable analytic tool for ML researchers and agent developers studying what agents actually do.

4️⃣ arXiv · LoRAFusion(⭐⭐⭐ 值得追踪)
4️⃣ arXiv · LoRAFusion(⭐⭐⭐ 值得追踪)
arXiv:2510.00206 工程化 方法 被引 7 · S2

本文提出 LoRAFusion,一种面向 LLM 的高效 LoRA 微调系统,可消除不必要的内存访问,在不付出重算或同步代价的前提下保持 compute-bound GEMM 的性能,并引入面向多任务微调的自适应批处理算法。LoRAFusion is introduced, an efficient LoRA fine-tuning system for LLMs that eliminates unnecessary memory accesses and preserves the performance of compute-bound GEMMs without incurring the cost of recomputation or synchronization and introduces an adaptive batching algorithm for multi-job fine-tuning.

QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation
QCell:重组与对齐细胞查询用于重叠实例分割
arXiv:2608.29253 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 QCell,一种新颖的基于查询的模型,用于在显微镜场景中去重叠细胞实例,在多个 benchmark 上优于 SOTA 方法,在 ISBI2014 上取得 +2.2 AP 和 +2.7 AJI。QCell is presented, a novel query-based model that de-overlaps cell instances in microscopy scenes and outperforms state-of-the-art methods across multiple benchmarks, achieving +2.2 AP and +2.7 AJI on ISBI2014.

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training
DRACO:面向长视野 agent 训练的动态 rubric 细粒度信用分配
arXiv:2609.04094 Agent 智能体 方法 OA · 绿色 被引 1 · S2

该工作提出 DRACO:Distributing Rubric-based Advantage for Credit Optimization,在训练期间动态生成 rubric 以追踪 policy 的演化能力,对已完成的轨迹一次性评分,并将该判断重新分配到负责标注 rubric 的步骤上,以在 GRPO 中产生差异化的 per-step advantage。This work proposes DRACO: Distributing Rubric-based Advantage for Credit Optimization, which generates rubrics dynamically during training to track the policy's evolving capability, scores those rubrics once per completed trajectory, and redistributes that judgment over the steps responsible for annotated rubrics to produce differentiated per-step advantages in GRPO.

Last Translation Benchmark
终极翻译基准(Last Translation Benchmark)
arXiv:2609.04173 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 The Last Translation Benchmark,一组由人工编写并经同行评审的样本,可打破领先的机器翻译模型,并提出新评估方法:每个样本附带手工编写的验证规则,描述该样本上的具体失败案例,从而支持可靠且可操作的未来评估。The Last Translation Benchmark is introduced, a collection of human-authored and peer-reviewed examples that break leading machine translation models and a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and actionable future evaluation.

VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement
VeriPhy:用于世界模型评估与精进的 agent 物理推理
arXiv:2609.03153 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 VeriPhy,一种可审计的物理验证系统,其中纯文本 planner 在观察任何帧之前将 prompt 编译为类型化的物理义务与静态验证的执行计划。VeriPhy, an auditable physical-verification system in which a text-only planner compiles the prompt into typed physical obligations and a statically validated execution plan before any frame is observed, is presented.

A Common Measure of Communication for Speech Brain-Computer Interfaces
面向语音脑机接口的通用通信度量指标
arXiv:2609.02887 多模态 方法 被引 3 · S2

推导出 open-vocabulary mutual information (OVMI),一种衡量 decoder 相对于用户可能希望传达词汇的参考分布所传达信息的信息论量度,为语音 BCI 社区提供了一种原则性方法以比较异构系统、改进词汇设计并衡量领域进展。Deriving open-vocabulary mutual information (OVMI), an information-theoretic quantity that measures the information conveyed by a decoder relative to a reference distribution over the words a user may wish to communicate, provides the speech BCI community with a principled way to compare heterogeneous systems, improve vocabulary design, and measure progress in the field.

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space
锁于入口,开于内里:RLVR 收窄解空间的位置
arXiv:2608.29188 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

表面 prompting 未能恢复多样性,而针对入口的干预则成功:使用早期 checkpoint 的 late-layer parameter interpolation 在不损失 pass@1 的情况下将解的覆盖度提高了 37%。While surface prompting fails to recover diversity, entrance-targeted interventions succeed: late-layer parameter interpolation with early checkpoints increases solution coverage by 37% at no loss in pass@1 and late-layer parameter interpolation with early checkpoints increases solution coverage by 37% at no loss in pass@1.

RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning
RoboTok:面向人类演示检索与灵巧操作学习的大规模互联网数据引擎
arXiv:2609.03199 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本工作提出 RoboTok,一个可扩展的数据引擎:利用人类操作视频作为查询,从互联网检索与操作相关的演示以训练灵巧机器人策略,并从以演员为中心的参考系下估计的 3D 手部轨迹中学习一个潜在运动空间This work introduces RoboTok, a scalable data engine that uses a query human manipulation video to retrieve manipulation-relevant internet demonstrations for training dexterous robot policies and learns a latent motion space from 3D hand trajectories expressed in estimated actor-centered reference frames.