研究库 论文知识库
Papers · organized/paper_cards

论文

1094 张论文卡片 · 方法 · OA 绿色

开放获取 全部 绿色 · 1640
The Compaction Cliff in Long-Running AI Agent Memory
长时运行 AI Agent 记忆中的压缩悬崖
arXiv:2608.22752 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本文提出 Knowledge Triage 框架,对 agent 知识库的每一行按类型分类,并为每种类型配置独立的保留策略,同时开源发布 AgentArtifactCorpus 数据集、分类器及参考实现。Knowledge Triage, a framework that classifies each line of an agent's knowledge base by type and routes each type through its own retention policy, is addressed, and AgentArtifactCorpus, the classifier, and the reference implementation are released.

8. TrustMargin:RAG 答案级仲裁框架
arXiv:2606.08397 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 TRUSTMARGIN,一种免训练、即插即用的仲裁层,利用模型自身的似然对两个候选进行打分,在不微调、无需外部评判或额外生成的情况下,在直接回答与 RAG 之间进行选择。TRUSTMARGIN is proposed, a training-free, plug-and-play arbitration layer that scores the two existing candidates with the model's own likelihoods and selects between Direct and RAG without fine-tuning, external judges, or additional generation.

Towards a Densing Law for User Representation Learning at Billion-Scale Capacity
迈向十亿级容量用户表示学习的稠密定律
arXiv:2608.23392 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 User Behavioral Densing Law,为大规模用户表示学习中的 tokenization 配置选择提供实用指导,并开发了 ALGN——一种自适应变长 tokenization 方法,可改善容量分配。The proposed User Behavioral Densing Law is proposed, providing practical guidance for tokenization configuration selection in large-scale user representation learning and ALGN, an adaptive variable-length tokenization method that improves capacity allocation, is developed.

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks
LongWoF-Bench:评估 EvoMap Gene 的可验证长工作流任务基准
arXiv:2608.23200 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

EvoMap 的结果表明,经过验证的执行经验可以被保留并共享为可复用的外部资源,使模型能够提升长工作流完成度,而无需反复承担经验探索的全部成本。EvoMap results show that verified execution experience can be retained and shared as a reusable external resource, enabling models to improve long-workflow completion without repeatedly paying the full cost of experience discovery.

AutoResearch: Insight In, Hallucination Out
AutoResearch:洞察输入,幻觉输出
arXiv:2608.17906 Agent 智能体 方法 OA · 绿色 被引 1 · S2

介绍 AutoResearch,一个连接 Idea Generation 与 Idea Execution 的两阶段系统,分别解决研究思路如何形成与如何通过实验可靠验证的问题,展示「实验前先夯实洞见、接受前先夯实结论」的研究流程。AutoResearch is introduced, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation to demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance.

EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment
EXPL-FR:通过视觉-语言对齐解释人脸识别模型
arXiv:2608.21486 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

覆盖 4 个 FR backbone 与 2 个 VLM 编码器;EXPL-FR 无需访问模型架构,支持身份级、单图及差异式解释,并在三种监督设置(人工标注、VLM 伪标签、完全 prompt 驱动的审计)下针对真实核验行为进行属性级审计基准测试。This work covers four FR backbones and two VLM encoders, EXPL-FR needs no architecture access, and supports identity-level, per-image, and differential explanations, and benchmark attribute-level auditing under three supervision settings, human labels, VLM pseudo-labels, and the authors' fully prompt-driven audit, against real verification behavior.

Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav
直接语料交互中的证据盲区:基于 AtlasNav 的持久化导航
arXiv:2608.24764 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

AtlasNav 减少了 Evidence Blindness,更早实现完整证据,在 PhantomWiki 上对语料结构与规模变化保持鲁棒,并在异构企业数据上取得领先性能。AtlasNav reduces Evidence Blindness, reduces Evidence Blindness, realizes complete evidence earlier, remains robust to corpus-structure and scale shifts on PhantomWiki, and achieves leading performance on heterogeneous enterprise data.

Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training
Game2World Engine:解锁真实游戏视频用于世界模型训练
arXiv:2608.24680 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 GameCleaner,一个无需 mask 的游戏 UI 移除模型,结合多模态语义理解与视频编辑能力,整体 VideoReward 较在带 UI 数据上训练的模型提升 6.83%。GameCleaner is proposed, a mask-free gameplay UI removal model that combines multimodal semantic understanding with video editing capabilities and improves overall VideoReward by 6.83% over those trained on UI-overlaid data.

Length-Adaptive Decoding for Masked Diffusion Machine Translation
掩码扩散机器翻译的长度自适应解码
arXiv:2608.22274 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 Entropy-Valley (EV):一种无需训练的画布长度选择器,通过 all-mask 前向的预测平均熵对候选目标画布打分,并挑选出 backbone 最「准备好」填充的画布。Entropy-Valley (EV), a training-free length selector that scores candidate target canvases by mean predictive entropy from all-mask forward passes and selects the canvas the backbone is most prepared to fill, is introduced.

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs
以标注作为 Rollout:面向视频 MLLMs 的高效可扩展强化学习
arXiv:2608.20492 多模态 方法 OA · 绿色 被引 1 · S2

本文研究视频 MLLM 的 RL 后训练样本效率与可扩展性,并提出 OraRL——一种随模型规模与数据规模共同 scaling 的解耦 advantage estimator,在 0.8B 到 9B backbone 上均超越其基线,并在 100k prompts 规模下超越 GRPO。The sample efficiency and scalability of RL post-training for video MLLMs and introduces OraRL, a decoupled advantage estimator that scales with model size and data, surpassing its backbone from 0.8B to 9B and GRPO up to 100k prompts.

CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild
CyberFactory:基于真实实例扩展网络安全能力
arXiv:2608.23181 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 CyberFactory,一个统一的开源框架,贯通 PoC 生成、漏洞修补与 CyberQA 三大任务中的数据构建、轨迹合成与模型训练。CyberFactory is introduced, a unified open-source framework that connects data construction, trajectory synthesis, and model training across proof-of-concept (PoC) generation, vulnerability patching, and cybersecurity question answering (CyberQA).

DREAM Technical Report
DREAM 技术报告
arXiv:2608.09408 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 DREAM,一种自主优化控制架构:在不替换现有流水线的前提下叠加感知可感知、可编排、可审计的策略层,支持将 agentic meta-control 作为工业推荐的一种可行范式。This work presents DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them, supporting agentic meta-control as a viable paradigm for industrial recommendation.

When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows
当 "Must" 变成 "Maybe":LLM Agent 工作流中的约束弱化
arXiv:2608.24569 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本文指出 LLM agent 中信息抽取与动作之间的 state-transmission 失效,并展示 handoff 变换如何在保留状态内容的同时削弱其对下游动作的约束。This work identifies a state-transmission failure between information extraction and action in large language model agents, and shows how handoff transformations can retain state content while weakening its constraints on downstream action.

MoTE: Mixture of Task Experts for Multi-Task Video Understanding
MoTE:面向多任务视频理解的 Task Expert 混合模型
arXiv:2608.24763 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MoTE(Mixture of Task Experts),一种将大语言模型前馈网络转化为任务特定专家同时保持多模态 backbone 共享的 decoder 架构,并在五个 COIN 基准上使用显式任务路由进行评估。This work proposes MoTE (Mixture of Task Experts), a decoder architecture that converts large language model feed-forward networks into task-specific experts while keeping the multimodal backbone shared and evaluates it on five COIN benchmarks using explicit task routes.

AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace
AgentRoom:基于 CRDT 共享工作空间的并发多 Agent 编程
arXiv:2608.23740 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

AgentRoom 是一种面向并发编码 Agent 的实时协同编辑协议,通过在 CRDT 合并的共享文件系统上将文件级 claim、status 和 broadcast 暴露为 MCP 工具,且运行间的差异小于 CLI-stable 模型。AgentRoom is a realtime collaborative editing protocol for concurrent coding agents that exposes file-level claim, status, and broadcast as MCP tools on a CRDT-merged shared filesystem and has less run-to-run variation than CLI-stable models.

7. Triton Attention Kernel 学术分析 (arXiv 2511.11581)
Triton Attention Kernel 学术分析 (arXiv 2511.11581)
arXiv:2511.11581 LLM 基础设施 方法 OA · 绿色 被引 4 · S2

本工作开发了一个 SOTA 的 paged attention kernel,完全基于领域特定即时编译语言 Triton 构建,在 NVIDIA 与 AMD GPU 上均达到 SOTA 性能。This work develops a state-of-the-art paged attention kernel that builds exclusively on the domain-specific just-in-time compiled language Triton to achieve state-of-the-art performance on both NVIDIA and AMD GPUs.

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
开放世界多 Agent 环境中的自主数学发现
arXiv:2608.23691 Agent 智能体 方法 OA · 绿色 被引 2 · S2

The Station 被评估为一个开放世界多 agent 环境,不同模型家族的 AI agent 在其中无需中心协调器或脚本化流程即可共同追求同一研究目标,并提供发现产生过程的透明记录。The Station is evaluated, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline, providing a transparent record of how discoveries emerged.

SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation
SecOPD:通过 On-Policy Distillation 缓解自适应 Prompt 注入
arXiv:2608.21500 Agent 智能体 方法 OA · 绿色 被引 4 · S2

本文提出 Secure On-Policy Distillation (SecOPD),提供 token 级反馈以指导防御性微调,并能泛化到训练中完全未见过的领域。This paper proposes Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning, and generalizes to domains completely unseen in training.

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
GigaBrain-0.7:以三系统架构将具身基础模型扩展至涌现能力
arXiv:2608.15875 多模态 方法 OA · 绿色 被引 6 · S2

本文提出 GigaBrain-0.7,一种跨多种机器人 embodiment 泛化能力显著增强的 embodied foundation model,并引入一阶段对齐训练,联合优化 vision-language 理解和多 embodiment 动作生成。This work presents GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation.

Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation
以 Rubric 作为视觉修复上下文以实现自演化的 UI-to-Code 生成
arXiv:2608.24138 多模态 方法 OA · 绿色 被引 1 · S2

评估表明,RubSE 在终轮和最佳轮设置下均显著优于朴素 self-evolution,refinement 轨迹更稳定,且轨迹级性能上限更高。Evaluations demonstrate that RubSE substantially outperforms na\"ive self-evolution in final-round and best-round settings, achieving more stable refinement trajectories and a higher trajectory-level performance ceiling.

The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents
交接代价:LLM Agent 中非原生轨迹的延续
arXiv:2608.24358 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文变化 handoff 方向、时机与接口,对比保留仓库状态下的全轨迹传输、压缩与轨迹移除,发现偏好接口随方向反转:减少 LC-model 轨迹信息可提升 escalation 质量,而移除 HC-model 轨迹则会降低 downshift 质量。This work varies handoff direction, timing, and interface, comparing full-trajectory transfer, compaction, and trajectory removal while preserving the repository state, and finds that the preferred interface also reverses with direction: reducing LC-model trajectory information improves escalation quality, whereas removing the HC-model trajectory reduces downshift quality.

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
面向持久化故事与交互式世界的长时音视频生成
arXiv:2608.23383 多模态 方法 OA · 绿色 被引 2 · S2

结果表明,记忆、几何控制以及 rollout-aware 训练为生成连贯故事和持续演化的交互式世界提供了实用基础。Results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds.

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data
Next-chunk reasoning RL 真的优于 SFT 吗?——在 no-CoT 数据下重新审视训练策略
arXiv:2608.23256 工程化 方法 OA · 绿色 被引 1 · S2

Mixed SFT 是一种单阶段监督微调,联合在 no-CoT 和 long-CoT 数据上训练,相比 next-chunk reasoning RL 取得了明显更高的 RLVR 后性能上限,同时训练算力开销减少超过 60 倍。Mixed SFT, a single supervised fine-tuning stage that jointly trains on no-CoT and long-CoT data, achieves a clearly higher post-RLVR performance ceiling than next-chunk reasoning RL while requiring over 60 times less training compute.

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling
FIRM-Video:先核查再评分,实现可靠的 text-to-video 奖励建模
arXiv:2608.21839 多模态 方法 OA · 绿色 被引 1 · S2

本文提出 FHM-Video,一种基于 check-before-score 原则的、由 checklist 驱动的统一数据构建框架,在 FIRM-Video-Bench 上取得最佳综合 MAE,并在三种视频生成器的 Best-of-8 采样中始终获得最高的 VBench Total、Quality 和 Semantic Score。FHM-Video, a unified checklist-driven data construction framework based on a check-before-score principle, is introduced, which achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.

Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction
Real-TurnTurk:用于话轮预测的多模态土耳其语语料库
arXiv:2608.22071 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文发布一个土耳其语多模态对话数据集,包含未脚本化的双人交互,并提供同步的前向视频、可归属到每位说话人的独立音频通道以及时间对齐的转写文本。A multimodal Turkish conversational dataset of unscripted dyadic interactions, comprising synchronized front-facing video, per-speaker audio channels that allow overlapping speech to be attributed to individual speakers, and time-aligned transcriptions is introduced.

7. LLM 压缩:联合剪枝 + 混合精度 PTQ
arXiv:2606.07819 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本工作提出一种新颖的混合精度 PTQ 策略,直接最小化整个模型的全局误差传播,而非孤立地处理逐层误差,并开发了一种新颖的联合优化方法,在统一搜索空间中同时学习结构化剪枝决策与混合精度量化策略。This work proposes a novel mixed-precision PTQ strategy that directly minimizes global error propagation across the entire model, rather than isolating layer-wise errors, and develops a novel joint optimization approach that simultaneously learns structural pruning decisions and mixed-precision quantization policies within a unified search space.

Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models
Stream4D:面向流式自回归扩散视频模型的 4D 一致性
arXiv:2608.19556 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文用前馈 4D 重建奖励替代静态 critic,显式建模场景动态,使连贯运动获得高一致性奖励,并加入对自然 scene-flow 幅值进行奖励同时抑制抖动与非刚性伪影的运动先验。This work replaces the static critic with a feed-forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards, and adds a motion prior that rewards natural scene-flow magnitude while penalizing jitter and non-rigid artifacts.

Multi-Granularity Context-Enhanced RAG over Multimodal Knowledge Graphs
多模态知识图谱上的多粒度上下文增强 RAG
arXiv:2608.25986 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种构建 Context-Enhanced MMKG (CEMMKG) 的新框架,能够有效利用上下文信息提升基于 MMKG 的 RAG 性能,并在不同基于 MMKG 的 RAG 方法上的有效性验证了其广泛适用性。A novel framework for constructing a Context-Enhanced MMKG (CEMMKG) is proposed, effective in leveraging contextual information to improve MMKG-based RAG performance and its effectiveness across different MMKG-based RAG methods demonstrates its broad applicability.

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
LibriBrain100:面向大规模神经语音解码的百小时广深 MEG 数据集
arXiv:2608.25204 多模态 方法 OA · 绿色 被引 4 · S2

本文发布 LibriBrain100,一个面向语音解码的大规模 MEG 数据集,从设计上保证可复现、标准化评估,并展示了广泛多被试数据的价值:对预训练模型进行有监督微调可大幅弥补单被试数据不足。LibriBrain100 is introduced, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation and the value of broad multi-subject data is demonstrated: supervised finetuning of a pre-trained model can substantially compensate for limited per-subject data.

Prefix Sliding for efficient test-time scaling
Prefix Sliding:面向高效 test-time scaling 的前缀滑动方法
arXiv:2608.26070 LLM 基础设施 方法 OA · 绿色 被引 2 · S2

提出 Prefix Sliding,在推理过程中丢弃不属于前缀或最近几千 token 窗口的 token,从而实现高效的长时程 test-time scaling。Prefix Sliding is proposed, which discards tokens during reasoning that are not part of the prefix or the window of the last few thousand tokens, allowing for efficient long-horizon test-time scaling.

6. QBugLM:量子软件调试多智能体框架
arXiv:2606.07314 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本工作提出 QBugLM,一个多 Agent 框架,可自动化量子软件调试流水线,覆盖基于分类法的缺陷注入、基于 LLM 的检测与修复,直至基于仿真的验证,框架无关地支持 OpenQASM 3.0 程序。This work proposes QBugLM, a multi-agent framework that automates the quantum software debugging pipeline, from taxonomy-driven bug injection to LLM-based detection and repair, and finally to simulation-based validation, for framework-agnostic OpenQASM 3.0 programs.

TTPO: Test-Time Policy Optimization
TTPO:测试时策略优化。
arXiv:2608.27448 工程化 方法 OA · 绿色 被引 1 · S2

提出 Test-Time Policy Optimization,一种非对称目标,通过 OPSD 蒸馏一致性 rollout,并使用 Grouped RL 惩罚不一致的 rollout;进一步通过 token 级选择精炼两个分支:蒸馏降低已收敛位置的权重,而 RL 仅惩罚置信的错误。Test-Time Policy Optimization is proposed, an asymmetric objective that distills agreeing rollouts via OPSD and penalizes disagreeing rollouts with Grouped RL and Token-level selection further refines both branches: distillation down-weights already-converged positions, while RL penalizes only confident errors.

PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
PILOT in the Loop:面向长时序 Agent 的实时自我改进。
arXiv:2608.26530 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本文提出 PILOT,一种通过两种耦合机制实现实时自我改进的 supervisor-worker 框架:(1) live steering 允许独立的 supervisor 在执行期间重定向或中止当前 worker;(2) live self-evolution 将执行中发现的过程与失败模式提炼为可复用的 skills 与记忆。PILOT is presented, a supervisor-worker harness for live self-improvement through two coupled mechanisms: (1) live steering lets a separate supervisor redirect or abort the active worker during execution; and (2) live self-evolution distils procedures and failure modes revealed during execution into reusable skills and memory.

Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning
Thinking on Shots:基于 Agentic 推理的一致性多镜头视频编辑。
arXiv:2608.26809 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个 Agentic 视频编辑框架,利用 LLM 与 VLM 的协同实现 shot 级视频解耦与精确指令解析,并构建 MMLVE-Bench,一个聚焦 MMLVE 的数据集,具有复杂的真实世界时空动态、高密度异构指令以及稀疏随机的实体分布。This work introduces an agentic editing framework that leverages the synergy of Large Language Models (LLMs) and Vision-Language Models (VLMs) to achieve shot-level video decoupling and precise instruction parsing, and constructs MMLVE-Bench, which is an MMLVE-focused dataset characterized by complex real-world spatiotemporal dynamics, high-density heterogeneous instructions, and sparse, random entity distributions.

Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning
Aphanta:诊断任务对齐的图像编辑中间态以服务多模态推理。
arXiv:2608.26993 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果将图像编辑定位为一种专门的视觉工作空间而非通用推理机制,并将 Aphanta 确立为可复用的协议,用于度量任务-表征对齐、编辑器实现及下游 pipeline 实用性。The results position image editing as a specialized visual workspace rather than a universal reasoning mechanism, and establish Aphanta as a reusable protocol for measuring task--representation alignment, editor realization, and downstream pipeline utility.

Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
将 Agentic 游戏开发作为可扩展世界模型的可验证轨迹数据引擎
arXiv:2608.25518 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出 Reinforcement Learning with Human-Engine Verification(RLHEV),一种结合稠密引擎信号与开发过程中隐式人类接受反馈的后训练范式,用于支持强化学习后训练。This work proposes Reinforcement Learning with Human-Engine Verification (RLHEV), a post-training paradigm that combines dense engine signals with implicit human acceptance feedback from the development process to support RL post-training.