研究库 论文知识库
Papers · organized/paper_cards

论文

1686 张论文卡片 · OA 绿色

开放获取 全部 绿色 · 1686
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe
Second Thought:让 LLM Agent 在执行与观察时并行推理
arXiv:2608.13667 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Second Thought,一种免训练的推理框架,在每个 Thought 阶段结束时立即 fork 四个辅助分支,与主循环并发解码,并在环境 observation 到达时将生成的 thought 合并回去。This work proposes Second Thought, a training-free inference framework that forks four auxiliary branches the instant each Thought phase concludes, decodes them concurrently with the main loop, and merges the generated thoughts back when the environment observation arrives.

PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment
PRM-as-a-Judge 1.5:机器人过程评估工具包
arXiv:2608.14284 评测基准 评测集 OA · 绿色 被引 2 · S2

一个机器人过程评估工具包,将 rollout 视频转化为稠密进度曲线并衍生多项细粒度指标;引入 RoboPulse++ 用于评估过程奖励模型(PRM)的可靠性,为评测者提供更准确的测试平台。A toolkit for robot process assessment that turns rollout videos into dense progress curves and derives multiple fine metrics, and introduces RoboPulse++ to evaluate the reliability of process reward models (PRM), providing evaluators with a more accurate testing platform.

Latent On-Policy Self-Distillation
Latent On-Policy Self-Distillation
arXiv:2608.13040 Agent 智能体 方法 OA · 绿色 被引 5 · S2

提出 Latent On-Policy Self-Distillation(LOPD),不再提出另一种手工设计、附带新形式 privileged context 的 OPSD 变体,而是让 teacher 的 privileged context 本身可从经验端到端学习。This work introduces Latent On-Policy Self-Distillation (LOPD), which, rather than proposing another hand-crafted OPSD variant with a newly prescribed form of privileged context, makes the teacher's privileged context itself learnable end-to-end from experience.

MobileMem: Learning from a Year of Mobile Experiences
MobileMem:从一年的移动端经验中学习
arXiv:2608.13606 评测基准 评测集 OA · 绿色 被引 1 · S2

提出 MobileMem,一个面向设备端长期记忆研究的 benchmark 和框架,基于长达一年的移动端经验集合,使 agent 能够记忆过去、理解当下并适应未来。This work introduces MobileMem, a benchmark and framework for studying on-device long-term memory, grounded in a year-scale collection of mobile experiences, and enables agents to remember the past, understand the present, and adapt to the future.

Multimodal Model Diffing for Feature Discovery and Control
多模态模型 Diffing:用于特征发现与控制
arXiv:2608.09928 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

MMDiff 是一种多模态 model-diffing 框架,训练多模态 SAE 并将其转化为特征级接口,用于发现和控制多模态行为;研究表明多模态 SAE 不仅可作为可解释性工具,还可作为审计、引导和控制 MLLM 行为的机制,以实现更安全、更具能力的生成。MMDiff, a multimodal model-diffing framework that trains multimodal SAEs and turns them into feature-level interfaces for discovering and controlling multimodal behavior, shows that multimodal SAEs can serve not only as interpretability tools, but as mechanisms for auditing, steering, and controlling MLLMs behavior toward safer and more capable generations.

Self-Supervised Visual On-Policy Distillation
自监督视觉 On-Policy Distillation
arXiv:2608.14144 多模态 方法 OA · 绿色 被引 7 · S2

提出 Self-Supervised Visual On-Policy Distillation(S²VOPD),一种简单有效的方法,通过非对称增强视图构建 on-policy 学习信号,系统地探索了视觉增强的广阔设计空间,并发现非对称性至关重要。Self-Supervised Visual On-Policy Distillation (S$^2$VOPD), a simple yet effective method that constructs on-policy learning signals from asymmetric augmented views, systematically explores a broad design space of visual augmentations and uncover that asymmetry matters.

SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning
SimpleOPD:面向长上下文推理的简单、与 Tokenizer 无关的 On-Policy Distillation
arXiv:2608.14277 工程化 方法 OA · 绿色 被引 5 · S2

引入 student reference KL 损失并 mask 特殊终止 token 的 advantage,以缓解生成过长和频繁截断的问题;在 HLE 和 HiPhO 等科学 benchmark 上取得改进,表明 OPD 传递的推理能力可泛化至数学训练领域之外。This work introduces a student reference KL loss and mask the advantages of special termination tokens to mitigate the problem of excessive generation length and frequent truncation, and improves on science benchmarks such as HLE and HiPhO, suggesting that OPD transfers reasoning capabilities that generalize beyond the mathematical training domain.

DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
DFM Mimir v1:基于 HRM 架构、仅使用合规 Post-Training 数据实现 1B 参数前沿性能的开放模型
arXiv:2608.13517 工程化 方法 OA · 绿色 被引 1 · S2

提出 Mimir v1,一个基于 Hierarchical Reasoning Model(HRM)架构的 10 亿参数语言模型,从头训练,在英语上具有高度竞争力,并仅使用合规的后训练数据在丹麦语上创下新的 SOTA。Mimir v1 is introduced, a 1-billion-parameter language model based on the Hierarchical Reasoning Model (HRM) architecture that is trained from scratch and delivers highly competitive performance for English and sets a new state of the art for Danish using only permissible post-training data.

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure
LittleLearner:在教学受控知识暴露下的语言模型
arXiv:2608.13545 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

通过在 post-training 与 in-context learning 中注入新知识的一系列实验,展示了该沙箱的实用性:这些方法让 LittleLearner 更好地利用已有知识,但并未提升其 out-of-scope 能力。The sandbox's utility is illustrated in a first suite of experiments on injecting new knowledge through post-training and in-context learning, which let LittleLearner better utilize existing knowledge, but do not raise out-of-scope capabilities.

UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations
UniProbe:一种利用多结构内部表示的可学习 token 级 VLM 幻觉检测器
arXiv:2608.10835 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 UniProbe,一个轻量、统一、可学习的检测器,通过单次前向传播建模冻结 LVLM 的异构计算 trace,在 token 级和物体级幻觉检测上达到 SOTA。UniProbe is introduced, a lightweight, unified, learnable detector that models a frozen LVLM's heterogeneous computational trace from a single forward pass and achieves state-of-the-art token-level and object-hallucination detection.

UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers
UNMASK:文本分类器中虚假捷径的发现与因果验证
arXiv:2608.09209 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 UNMASK,一个全自动 pipeline,可在无需额外人工标注的情况下发现、因果验证并缓解文本分类器中的伪相关,并证明其发现与验证阶段可泛化至奖励模型的偏好数据。U N M ASK is presented, a fully automated pipeline that discovers, causally verifies, and mitigates spurious correlations in text classifiers without additional human annotation, and demonstrates that the discovery and validation stages generalize to reward model preference data.

Model-agnostic Retrieval-Augmented Extended Forecasting for time series
模型无关的检索增强扩展时间序列预测
arXiv:2608.14054 RAG 检索增强 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

在多个 benchmark 数据集上的实证评估表明,RAEF 在准确率和推理开销方面均优于 RAF,并且与零样本及微调基础模型的全面对比显示,RAEF 在避免微计算负担的同时取得了与微调相当或更优的性能。Empirical evaluation across multiple benchmark datasets demonstrates that RAEF outperforms RAF in both accuracy and inference overhead, and comprehensive comparisons with zero-shot and fine-tuned foundation models show that RAEF achieves competitive or superior performance to fine-tuning while avoiding its computational burden.

Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead
Nanbeige4.2-3B on Apple Silicon:修复部署 Bug 并降低 Looped Transformer 显存开销
arXiv:2608.13987 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出一种 chunked-prefill 策略,可缓解由此带来的内存容量惩罚,在 32 GiB 共享内存上将允许的上下文宽度扩展 $2.7 \times$,但即便降低了内存开销,仍需打补丁才能使 Nanbeige4.2-3B 可用。A chunked-prefill strategy is introduced which alleviates the incurred memory-capacity penalty, extending allowable context width by $2.7 \times$ on 32~GiB shared memory, however, even with the reduced memory overhead, it is shown that patches are required to render Nanbeige4.2-3B usable.

Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
增强并不意味着可预测:思维模型中的推理行为
arXiv:2608.13760 工程化 评测集 OA · 绿色 被引 1 · S2

发现面向推理的训练并未优先放大具有最高 Lift 的行为,这促使研究者采用过程级目标,以奖励经过校准且有依据的推理,而不仅仅是表面形式。It is found that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.

Is this Citation on Point?
这条引用是否切中要点?
arXiv:2608.12571 评测基准 观点 OA · 绿色 被引 0 · S2 + OpenAlex

2023 年,一位纽约法官在 Mata v. Avianca 案中制裁了两位律师,因其提交的 brief 包含了由 ChatGPT 生成的虚构引用。此类失误大多能被数据库检索发现;但更棘手的问题在于检测那些指向真实案例、却不支持其所述命题的引用——这一失效模式是现有面向法律场景的 LLM 评测基本忽略的。本文通过对来自两个法律语料库的真实法律引用进行受控扰动(替换引用In 2023, a New York judge sanctioned two attorneys in Mata v. Avianca for filing a brief with hallucinated citations generated by ChatGPT. Such failures are largely caught by database lookups; the harder problem is detecting citations that point to real cases but do not support the propositions for which they are offered -- a failure mode that existing evaluations of LLMs for legal use cases largely overlook. In this paper, we study proposition-level citation support verification through controlled perturbations of real legal citations obtained from two legal corpora, either replacing the cite

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
Apodex Discovery:用于评估与构建探索型 AI 的现实基准与环境
arXiv:2608.11341 评测基准 评测集 OA · 绿色 被引 1 · S2

提出 Apodex Discovery,一个通过 heavy-duty solver 构建和评估发现型 AI 的框架;该 solver 包含一个 foundation model、harness、工具和控制策略,用于执行长期的、有状态的、可验证的探索。This work introduces Apodex Discovery, a framework for building and evaluating discoverative AI through the heavy-duty solver, a system comprising a foundation model, harness, tools, and control policies that pursues extended, stateful, verifiable investigations.

Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
Agent 抓 Agent:临床多 Agent 系统中的捷径级联与 benchmark 作弊
arXiv:2608.03744 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

探讨共享工作空间上 LLM Agent 委员会的审议过程是否能被捷径和线索(benchmark 所奖励但临床医生会忽略的)所博弈,以及委员会的社会可信度所构成的游戏。It is asked whether committees of language-model agents deliberating on a shared workspace can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore, and what games a committee is social plausibility.

Modular Cognitive Architecture Emerges in Large Language Models
模块化认知架构在 Large Language Models 中涌现
arXiv:2608.13567 LLM 基础设施 方法 OA · 绿色 被引 2 · S2

通过跨 46 个任务(涵盖四个认知领域)的电路分析,发现 LLM 发展出与人类大脑相似的模块化架构:在人类大脑中依赖同一网络的任务会在 LLM 中招募重叠的神经元,而依赖不同网络的任务则招募不同的神经元。Using circuit analyses across N=46 tasks spanning four cognitive domains, it is found that Large Language Models develop a modular architecture that mirrors the human brain: tasks drawing on the same network in humans recruit overlapping neurons in LLMs, whereas tasks drawing on different networks recruit distinct neurons.

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution
FreeToken:带宽自适应执行的边缘原生 MoE 高效服务
arXiv:2608.16157 LLM 基础设施 方法 OA · 绿色 被引 3 · S2

FreeToken 是一个 edge-native 的 MoE serving 系统,将个人机器视为一个统一、弹性的推理平台而非小型 GPU,把开放权重转化为可部署的本地软件,使用户已有的机器成为运行前沿规模智能的实用平台。FreeToken is an edge-native MoE serving system that treats a personal machine not as a small GPU, but as a unified, elastic inference platform, turning open weights into deployable local software, making the machines users already own a practical platform for frontier-scale intelligence.

GRIP: Grounded Reasoning via Information-Restricted Premises
GRIP:通过信息受限前提的扎根推理
arXiv:2608.16776 RAG 检索增强 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 GRIP(Grounded Reasoning via Information-Restricted Premises),引入容量不对称:decoder 对 query 保持全维度访问,而检索到的证据则通过一个严苛的随机瓶颈,迫使证据通道仅编码 query 中无法获得的残余信息。GRIP (Grounded Reasoning via Information-Restricted Premises), which imposes capacity asymmetry: the decoder keeps full-dimensional access to the query, while retrieved evidence passes through a severe stochastic bottleneck, which forces the evidence channel to encode only the residual information unavailable from the query.

Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement
基于超图的多模态检索增强生成与增量优化
arXiv:2608.16628 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

将文档结构形式化为多模态超图(Multimodal Hypergraph),以超边作为统一语义容器来封装跨文本、图像和表格的多路关联,超越点对点建模,并引入 Anchor-driven Incremental Refinement 机制。This paper formalizes the document structure as a Multimodal Hypergraph, utilizing hyperedges as unified semantic containers to encapsulate multi-way associations across text, images, and tables, thereby transcending point-to-point modeling and introducing an Anchor-driven Incremental Refinement mechanism.

DSPrompt: Dynamic Soft Prompt Defense Against M-RAG Corruption
DSPrompt:面向 M-RAG 投毒攻击的动态软提示防御
arXiv:2608.16536 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 DSPrompt,一种 Dynamic Soft Prompt 防御框架,无需修改检索 pipeline,直接重塑 retriever 的 embedding 语义,并以极低的计算成本 consistently 优于现有防御基线。DSPrompt is proposed, a Dynamic Soft Prompt defense framework that directly reshapes the retriever's embedding semantics, without modifying the retrieval pipeline, and is consistently outperforming existing defense baselines at a fraction of their computational cost.

When Context Misleads: Intent-Guided Decoding for Robust Retrieval-Augmented Generation
当上下文误导时:面向鲁棒检索增强生成的意图引导解码
arXiv:2608.16515 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Intent-Guided Decoding (IGD),一个根据用户意图在检索上下文和参数化记忆之间进行仲裁的框架,显著提升了 RAG 中的事实恢复能力。Intent-Guided Decoding (IGD) is proposed, a framework that arbitrates between retrieved context and parametric memory according to user intent and substantially improves factual recovery in RAG.

HarnessEval-W: Agentifying the Evaluation of Visual Worlds
HarnessEval-W:将视觉世界模型的评估 Agent 化
arXiv:2608.16859 评测基准 评测集 OA · 绿色 被引 3 · S2

提出 HarnessEval-W,一个 agentified 的评估 pipeline,将 LLM 生态中的 harness 范式引入 world model 基准测试,并在 330 个评估用例上对 18 个代表性 world model 进行了评估。This work introduces HarnessEval-W, an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking, and applies HarnessEval-W to 18 representative world models over 330 evaluation cases.

Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form
Gathered, Not Admitted:注意力如何将潜变量带入可言语化的形式
arXiv:2608.15022 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

语言模型以一种可被报告的形式持有潜在量,并且当任务需要灵活复用该量时,更多该量的信息会以这种形式存在。Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly when the task requires reusing it flexibly.

GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks
GRNEdit:基于生成式 Refinement 网络的新二元证据视角下的高效通用视频编辑
arXiv:2608.16328 多模态 方法 OA · 绿色 被引 2 · S2

GRNEdit 是一个轻量级的两阶段指令驱动通用视频编辑框架,性能优于多个 14B 开源编辑器,同时其 8B 模型与领先的开源编辑器表现相当。GRNEdit, a lightweight two-stage framework for instruction-based general video editing that outperforms multiple 14B open-source editors, while the 8B model performs on par with leading open-source editors.

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation
TRACE-Bench:多参考图像生成的分解与诊断
arXiv:2608.16765 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

认识到多样化的多参考任务共享一组共同的原子操作,本文形式化了四个算子:Anchor、Disentangle、Apply 和 Compose,并构建了 TRACE-Bench,包含约 1,600 个跨 slot 数量 1–8 的评估用例。Recognizing that diverse multi-reference tasks share a common set of atomic operations, this work formalizes four operators: Anchor, Disentangle, Disentangle, Apply, and Compose, and constructs TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8.

A Plug-and-Play 2D Motion Interface for Real-World Motion Language Models
面向真实世界 Motion Language Model 的即插即用 2D Motion 接口
arXiv:2608.15984 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出一个即插即用的 2D Motion Interface,使预训练于 3D 的 MoLM 能够在不修改或微调原始模型的情况下接受 2D 运动输入,并在 2D 运动任务上优于从头训练 MoLM。A plug-and-play 2D Motion Interface is introduced that enables 3D-pretrained MoLMs to accept 2D motion inputs without modifying or fine-tuning the original models and outperforms training MoLMs from scratch on 2D motions.

Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search
Large Discovery Models:基于经验、依托模型驱动的开放式搜索
arXiv:2608.15669 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Large Discovery Model (LDM),一种经验驱动的循环架构,将生成模型与贝叶斯非参数奖励代理模型耦合,产生一种感知不确定性的价值,用于引导候选的生成、精炼与选择。This work introduces the Large Discovery Model (LDM), an empirically grounded recurrent architecture that couples a generative model with a Bayesian non-parametric reward surrogate model, yielding an uncertainty-aware value that guides candidate generation, refinement, and selection.

WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations
WorldRover:一个用于世界探索的、可扩展的、带有丰富标注的合成视频数据引擎
arXiv:2608.15659 多模态 方法 OA · 绿色 被引 1 · S2

WorldRover 将长视野世界探索转化为可扩展的数据生成问题,为需要在可探索世界中构建、维护并重访一致表征的模型提供监督信号。WorldRover turns long-horizon world exploration into a scalable data-generation problem, providing supervision for models that must build, maintain, and revisit coherent representations of an explorable world.