研究库 论文知识库
Papers · organized/paper_cards

论文

1640 张论文卡片 · OA 绿色

开放获取 全部 绿色 · 1640
AutoResearch: Insight In, Hallucination Out
AutoResearch:洞察输入,幻觉输出
arXiv:2608.17906 Agent 智能体 方法 OA · 绿色 被引 1 · S2

介绍 AutoResearch,一个连接 Idea Generation 与 Idea Execution 的两阶段系统,分别解决研究思路如何形成与如何通过实验可靠验证的问题,展示「实验前先夯实洞见、接受前先夯实结论」的研究流程。AutoResearch is introduced, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation to demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance.

Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs
量化感知修复:恢复压缩后 4-Bit LLM 的实用方案
arXiv:2608.20953 LLM 基础设施 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

目标是提供无需数周超参搜索即可部署的方案,直接从原始未压缩模型蒸馏 4-bit 学生模型,并以开源权重形式发布为 Hypernova-60B。The aim is a recipe deployable without a multi-week hyper-parameter search, which distills the 4-bit student directly from the original, uncompressed model, and is released open-weight as Hypernova-60B.

EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment
EXPL-FR:通过视觉-语言对齐解释人脸识别模型
arXiv:2608.21486 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

覆盖 4 个 FR backbone 与 2 个 VLM 编码器;EXPL-FR 无需访问模型架构,支持身份级、单图及差异式解释,并在三种监督设置(人工标注、VLM 伪标签、完全 prompt 驱动的审计)下针对真实核验行为进行属性级审计基准测试。This work covers four FR backbones and two VLM encoders, EXPL-FR needs no architecture access, and supports identity-level, per-image, and differential explanations, and benchmark attribute-level auditing under three supervision settings, human labels, VLM pseudo-labels, and the authors' fully prompt-driven audit, against real verification behavior.

ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts
ClawProBench:基于 trace 感知、运行时覆盖与冻结式工作场景 holdout 的 AI Agent 评测
arXiv:2608.22510 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 ClawProBench:基于 OpenClaw(具备 workspace 工具及浏览、记忆、消息、调度、skill、subagent 等原生能力的实时 agent 运行时)实例化的 trace-aware、runtime-native agent 评估基准。ClawProBench is presented, a trace-aware benchmark for runtime-native agent evaluation instantiated on OpenClaw, a live agent runtime with workspace tools and native surfaces for browsing, memory, messaging, scheduling, skills, and subagents.

PinSieve: Production Selective VLM Serving and a Governed Memory Flywheel for Enterprise Content-Quality Triage
PinSieve:生产级选择性 VLM 服务与企业内容质量分诊的可治理记忆飞轮
arXiv:2608.24040 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 PinSieve——大规模内容质量流水线中的生产级案例:一个选择性 vision-language-model Serving Agent,仅处理轻量上游模型无法覆盖的 grey-zone 切片,在线暴露标量路由评分,并保留受控的人工升级通道。This work presents PinSieve, a production case study in a large-scale content-quality pipeline, a selective vision-language-model Serving Agent that operates only on the grey-zone slice left unresolved by lightweight upstream models, exposes a scalar routing score online, and preserves controlled human escalation.

Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav
直接语料交互中的证据盲区:基于 AtlasNav 的持久化导航
arXiv:2608.24764 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

AtlasNav 减少了 Evidence Blindness,更早实现完整证据,在 PhantomWiki 上对语料结构与规模变化保持鲁棒,并在异构企业数据上取得领先性能。AtlasNav reduces Evidence Blindness, reduces Evidence Blindness, realizes complete evidence earlier, remains robust to corpus-structure and scale shifts on PhantomWiki, and achieves leading performance on heterogeneous enterprise data.

Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training
Game2World Engine:解锁真实游戏视频用于世界模型训练
arXiv:2608.24680 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 GameCleaner,一个无需 mask 的游戏 UI 移除模型,结合多模态语义理解与视频编辑能力,整体 VideoReward 较在带 UI 数据上训练的模型提升 6.83%。GameCleaner is proposed, a mask-free gameplay UI removal model that combines multimodal semantic understanding with video editing capabilities and improves overall VideoReward by 6.83% over those trained on UI-overlaid data.

7️⃣ ByteHouse · 字节跳动云原生数据仓库架构深度解析(arXiv)⭐⭐⭐⭐ 系统复现
arXiv:2602.08226 数据与向量库 应用落地 Open MIND OA · 绿色 被引 0 · S2 + OpenAlex
From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms
从看见到行动:智能眼镜作为第一人称智能平台
arXiv:2608.24877 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本综述首次以统一框架系统研究智能眼镜,形式化第一人称数据流与受限任务效用,并提出覆盖采集、反应式感知、上下文辅助、持续状态、受控行动与具身耦合的 L0–L5 框架。This survey is the first to systematically study smart glasses through a unified framework, formalizing first-person data flow and constrained task utility, and introducing an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling.

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
LAION-BVD:一个用于多模态预训练的千万小时级开放视频数据集
arXiv:2608.24845 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 LAION-BVD,面向多模态学习的大规模开放视频数据集,包含从 CommonCrawl 收集的 1.3B 条平台特定视频 URL,并通过抽取场景切换帧,将视频帧作为图文数据的替代来源加以探索。This work presents LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl, and explores video frames as an alternative source of image-text data by extracting scene-changing frames.

Length-Adaptive Decoding for Masked Diffusion Machine Translation
掩码扩散机器翻译的长度自适应解码
arXiv:2608.22274 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 Entropy-Valley (EV):一种无需训练的画布长度选择器,通过 all-mask 前向的预测平均熵对候选目标画布打分,并挑选出 backbone 最「准备好」填充的画布。Entropy-Valley (EV), a training-free length selector that scores candidate target canvases by mean predictive entropy from all-mask forward passes and selects the canvas the backbone is most prepared to fill, is introduced.

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs
以标注作为 Rollout:面向视频 MLLMs 的高效可扩展强化学习
arXiv:2608.20492 多模态 方法 OA · 绿色 被引 1 · S2

本文研究视频 MLLM 的 RL 后训练样本效率与可扩展性,并提出 OraRL——一种随模型规模与数据规模共同 scaling 的解耦 advantage estimator,在 0.8B 到 9B backbone 上均超越其基线,并在 100k prompts 规模下超越 GRPO。The sample efficiency and scalability of RL post-training for video MLLMs and introduces OraRL, a decoupled advantage estimator that scales with model size and data, surpassing its backbone from 0.8B to 9B and GRPO up to 100k prompts.

CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild
CyberFactory:基于真实实例扩展网络安全能力
arXiv:2608.23181 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 CyberFactory,一个统一的开源框架,贯通 PoC 生成、漏洞修补与 CyberQA 三大任务中的数据构建、轨迹合成与模型训练。CyberFactory is introduced, a unified open-source framework that connects data construction, trajectory synthesis, and model training across proof-of-concept (PoC) generation, vulnerability patching, and cybersecurity question answering (CyberQA).

DREAM Technical Report
DREAM 技术报告
arXiv:2608.09408 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 DREAM,一种自主优化控制架构:在不替换现有流水线的前提下叠加感知可感知、可编排、可审计的策略层,支持将 agentic meta-control 作为工业推荐的一种可行范式。This work presents DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them, supporting agentic meta-control as a viable paradigm for industrial recommendation.

When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows
当 "Must" 变成 "Maybe":LLM Agent 工作流中的约束弱化
arXiv:2608.24569 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本文指出 LLM agent 中信息抽取与动作之间的 state-transmission 失效,并展示 handoff 变换如何在保留状态内容的同时削弱其对下游动作的约束。This work identifies a state-transmission failure between information extraction and action in large language model agents, and shows how handoff transformations can retain state content while weakening its constraints on downstream action.

MoTE: Mixture of Task Experts for Multi-Task Video Understanding
MoTE:面向多任务视频理解的 Task Expert 混合模型
arXiv:2608.24763 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MoTE(Mixture of Task Experts),一种将大语言模型前馈网络转化为任务特定专家同时保持多模态 backbone 共享的 decoder 架构,并在五个 COIN 基准上使用显式任务路由进行评估。This work proposes MoTE (Mixture of Task Experts), a decoder architecture that converts large language model feed-forward networks into task-specific experts while keeping the multimodal backbone shared and evaluates it on five COIN benchmarks using explicit task routes.

AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace
AgentRoom:基于 CRDT 共享工作空间的并发多 Agent 编程
arXiv:2608.23740 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

AgentRoom 是一种面向并发编码 Agent 的实时协同编辑协议,通过在 CRDT 合并的共享文件系统上将文件级 claim、status 和 broadcast 暴露为 MCP 工具,且运行间的差异小于 CLI-stable 模型。AgentRoom is a realtime collaborative editing protocol for concurrent coding agents that exposes file-level claim, status, and broadcast as MCP tools on a CRDT-merged shared filesystem and has less run-to-run variation than CLI-stable models.

Automata from Agent Traces: Failure and Next-Step Prediction
来自 Agent 轨迹的自动机:失败与下一步预测
arXiv:2608.23670 Agent 智能体 应用落地 OA · 绿色 被引 1 · S2

行为拓扑更多由部署 harness 决定而非 LLM 本身,为安全审计和运行时监控提供一种与模型无关的结构化 primitive,并同时满足两类预测目标。Behavioral topology is shaped more by the deployment harness than by the LLM, providing a model-agnostic structural primitive for safety auditing and runtime monitoring, and addresses both prediction goals.

7. Triton Attention Kernel 学术分析 (arXiv 2511.11581)
Triton Attention Kernel 学术分析 (arXiv 2511.11581)
arXiv:2511.11581 LLM 基础设施 方法 OA · 绿色 被引 4 · S2

本工作开发了一个 SOTA 的 paged attention kernel,完全基于领域特定即时编译语言 Triton 构建,在 NVIDIA 与 AMD GPU 上均达到 SOTA 性能。This work develops a state-of-the-art paged attention kernel that builds exclusively on the domain-specific just-in-time compiled language Triton to achieve state-of-the-art performance on both NVIDIA and AMD GPUs.

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
开放世界多 Agent 环境中的自主数学发现
arXiv:2608.23691 Agent 智能体 方法 OA · 绿色 被引 2 · S2

The Station 被评估为一个开放世界多 agent 环境,不同模型家族的 AI agent 在其中无需中心协调器或脚本化流程即可共同追求同一研究目标,并提供发现产生过程的透明记录。The Station is evaluated, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline, providing a transparent record of how discoveries emerged.

SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation
SecOPD:通过 On-Policy Distillation 缓解自适应 Prompt 注入
arXiv:2608.21500 Agent 智能体 方法 OA · 绿色 被引 4 · S2

本文提出 Secure On-Policy Distillation (SecOPD),提供 token 级反馈以指导防御性微调,并能泛化到训练中完全未见过的领域。This paper proposes Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning, and generalizes to domains completely unseen in training.

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
GigaBrain-0.7:以三系统架构将具身基础模型扩展至涌现能力
arXiv:2608.15875 多模态 方法 OA · 绿色 被引 6 · S2

本文提出 GigaBrain-0.7,一种跨多种机器人 embodiment 泛化能力显著增强的 embodied foundation model,并引入一阶段对齐训练,联合优化 vision-language 理解和多 embodiment 动作生成。This work presents GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation.

Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation
以 Rubric 作为视觉修复上下文以实现自演化的 UI-to-Code 生成
arXiv:2608.24138 多模态 方法 OA · 绿色 被引 1 · S2

评估表明,RubSE 在终轮和最佳轮设置下均显著优于朴素 self-evolution,refinement 轨迹更稳定,且轨迹级性能上限更高。Evaluations demonstrate that RubSE substantially outperforms na\"ive self-evolution in final-round and best-round settings, achieving more stable refinement trajectories and a higher trajectory-level performance ceiling.

The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents
交接代价:LLM Agent 中非原生轨迹的延续
arXiv:2608.24358 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文变化 handoff 方向、时机与接口,对比保留仓库状态下的全轨迹传输、压缩与轨迹移除,发现偏好接口随方向反转:减少 LC-model 轨迹信息可提升 escalation 质量,而移除 HC-model 轨迹则会降低 downshift 质量。This work varies handoff direction, timing, and interface, comparing full-trajectory transfer, compaction, and trajectory removal while preserving the repository state, and finds that the preferred interface also reverses with direction: reducing LC-model trajectory information improves escalation quality, whereas removing the HC-model trajectory reduces downshift quality.

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
面向持久化故事与交互式世界的长时音视频生成
arXiv:2608.23383 多模态 方法 OA · 绿色 被引 2 · S2

结果表明,记忆、几何控制以及 rollout-aware 训练为生成连贯故事和持续演化的交互式世界提供了实用基础。Results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds.

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data
Next-chunk reasoning RL 真的优于 SFT 吗?——在 no-CoT 数据下重新审视训练策略
arXiv:2608.23256 工程化 方法 OA · 绿色 被引 1 · S2

Mixed SFT 是一种单阶段监督微调,联合在 no-CoT 和 long-CoT 数据上训练,相比 next-chunk reasoning RL 取得了明显更高的 RLVR 后性能上限,同时训练算力开销减少超过 60 倍。Mixed SFT, a single supervised fine-tuning stage that jointly trains on no-CoT and long-CoT data, achieves a clearly higher post-RLVR performance ceiling than next-chunk reasoning RL while requiring over 60 times less training compute.

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling
FIRM-Video:先核查再评分,实现可靠的 text-to-video 奖励建模
arXiv:2608.21839 多模态 方法 OA · 绿色 被引 1 · S2

本文提出 FHM-Video,一种基于 check-before-score 原则的、由 checklist 驱动的统一数据构建框架,在 FIRM-Video-Bench 上取得最佳综合 MAE,并在三种视频生成器的 Best-of-8 采样中始终获得最高的 VBench Total、Quality 和 Semantic Score。FHM-Video, a unified checklist-driven data construction framework based on a check-before-score principle, is introduced, which achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.

Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction
Real-TurnTurk:用于话轮预测的多模态土耳其语语料库
arXiv:2608.22071 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文发布一个土耳其语多模态对话数据集,包含未脚本化的双人交互,并提供同步的前向视频、可归属到每位说话人的独立音频通道以及时间对齐的转写文本。A multimodal Turkish conversational dataset of unscripted dyadic interactions, comprising synchronized front-facing video, per-speaker audio channels that allow overlapping speech to be attributed to individual speakers, and time-aligned transcriptions is introduced.

7. LLM 压缩:联合剪枝 + 混合精度 PTQ
arXiv:2606.07819 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本工作提出一种新颖的混合精度 PTQ 策略,直接最小化整个模型的全局误差传播,而非孤立地处理逐层误差,并开发了一种新颖的联合优化方法,在统一搜索空间中同时学习结构化剪枝决策与混合精度量化策略。This work proposes a novel mixed-precision PTQ strategy that directly minimizes global error propagation across the entire model, rather than isolating layer-wise errors, and develops a novel joint optimization approach that simultaneously learns structural pruning decisions and mixed-precision quantization policies within a unified search space.

Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models
Stream4D:面向流式自回归扩散视频模型的 4D 一致性
arXiv:2608.19556 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文用前馈 4D 重建奖励替代静态 critic,显式建模场景动态,使连贯运动获得高一致性奖励,并加入对自然 scene-flow 幅值进行奖励同时抑制抖动与非刚性伪影的运动先验。This work replaces the static critic with a feed-forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards, and adds a motion prior that rewards natural scene-flow magnitude while penalizing jitter and non-rigid artifacts.

PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans
PlanSightRAG:面向土木标准图自动化问答与合规审查的视觉优先多模态 RAG
arXiv:2608.26091 RAG 检索增强 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

PlanSightRAG 是一种 Visual-First 多模态 RAG 框架,直接对图纸图像建立索引并进行推理,集成了 ColNomic-3B 多向量检索、Agentic Planner-Retriever-Auditor-Synthesizer,并以 MaxSim 热力图作为证据链。A Visual-First Multimodal Retrieval-Augmented Generation (RAG) framework called PlanSightRAG, which indexes and reasons directly over plan imagery, integrates a ColNomic-3B multi-vector retrieval, an agentic Planner-Retriever-Auditor-Synthesizer, and MaxSim heatmaps as an evidence trail.

Multi-Granularity Context-Enhanced RAG over Multimodal Knowledge Graphs
多模态知识图谱上的多粒度上下文增强 RAG
arXiv:2608.25986 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种构建 Context-Enhanced MMKG (CEMMKG) 的新框架,能够有效利用上下文信息提升基于 MMKG 的 RAG 性能,并在不同基于 MMKG 的 RAG 方法上的有效性验证了其广泛适用性。A novel framework for constructing a Context-Enhanced MMKG (CEMMKG) is proposed, effective in leveraging contextual information to improve MMKG-based RAG performance and its effectiveness across different MMKG-based RAG methods demonstrates its broad applicability.

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
LibriBrain100:面向大规模神经语音解码的百小时广深 MEG 数据集
arXiv:2608.25204 多模态 方法 OA · 绿色 被引 4 · S2

本文发布 LibriBrain100,一个面向语音解码的大规模 MEG 数据集,从设计上保证可复现、标准化评估,并展示了广泛多被试数据的价值:对预训练模型进行有监督微调可大幅弥补单被试数据不足。LibriBrain100 is introduced, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation and the value of broad multi-subject data is demonstrated: supervised finetuning of a pre-trained model can substantially compensate for limited per-subject data.

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?
SWE Refactor Bench:编码 Agent 能否完成长时序全仓库技术栈迁移?
arXiv:2608.23564 Agent 智能体 评测集 OA · 绿色 被引 3 · S2

本文发布 SWE Refactor Bench,一个包含 20 项全仓库迁移的基准,涵盖 4 类技术债务,作为开发面向可靠全仓库迁移的编码 Agent 的严格测试平台。SWE Refactor Bench is introduced, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt, and SWE Refactor Bench is positioned as a rigorous testbed for developing coding agents for reliable whole-repository migrations.

Skill Issue: Are Skills Language-Invariant in LLMs?
Skill Issue:LLM 中的 Skill 是否具备语言无关性?
arXiv:2608.25832 评测基准 评测集 OA · 绿色 被引 1 · S2

本文通过多语言 self-play,正交于知识与综合基准性能对跨语言技能不一致性进行量化,表明技能差异是开发真正多语言模型过程中可衡量且主要的障碍。This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark performance via multilingual self-play, and shows that skill discrepancies are a measurable major roadblock in the development of truly multilingual models.

Prefix Sliding for efficient test-time scaling
Prefix Sliding:面向高效 test-time scaling 的前缀滑动方法
arXiv:2608.26070 LLM 基础设施 方法 OA · 绿色 被引 2 · S2

提出 Prefix Sliding,在推理过程中丢弃不属于前缀或最近几千 token 窗口的 token,从而实现高效的长时程 test-time scaling。Prefix Sliding is proposed, which discards tokens during reasoning that are not part of the prefix or the window of the last few thousand tokens, allowing for efficient long-horizon test-time scaling.