研究库 论文知识库
Papers · organized/paper_cards

论文

1171 张论文卡片 · 方法

开放获取 全部 绿色 · 1640
UniSpace: Unified Visual Representation and Scalable Multimodal Modeling
UniSpace:统一视觉表征与可扩展多模态建模
arXiv:2608.08676 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Patch Reparameterization,在保留原始语义通路的同时,添加一个面向重构的 patch embedding,为同一组冻结的 ViT 块提供细粒度视觉信息,在保持多模态理解能力的同时实现高保真图像重构,并取得有利的重构—生成权衡。Patch Reparameterization is introduced, which preserves the original semantic pathway while adding a reconstruction-aware patch embedding that provides fine-grained visual information to the same frozen ViT blocks, and preserves multimodal understanding while enabling high-fidelity image reconstruction and a favorable reconstruction--generation trade-off.

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection
Task-CoEvolve:通过自适应验证任务选择实现 Harness 高效优化
arXiv:2608.20169 评测基准 方法 OA · 绿色 被引 4 · S2

Task-CoEvolve 基于以下观察:相比被一致解决或一致失败的候选任务,候选 harness 之间存在分歧的任务更能提供区分信息;它利用基于历史结果的方差加权采样,将评估聚焦在能力前沿附近的任务上。Task-CoEvolve builds on the observation that tasks on which candidate harnesses disagree are more informative for distinguishing among them than tasks that are consistently solved or failed, and uses variance-weighted sampling based on past outcomes to focus evaluation on tasks near the capability frontier.

One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders
一页污染足矣:评测 LLM 推荐系统中的网页内容污染
arXiv:2606.13610 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 FORGE(Fake Online Recommendations in Generative Environments),将一组固定检索网页中的真实商品在本地改写为虚假商品,并在 15 个类别、5 种消费场景下的 225 件真实商品上,衡量 LLM 推荐虚假商品的频率。This work introduces FORGE (Fake Online Recommendations in Generative Environments), which locally rewrites real products in a frozen set of retrieved web pages into fake ones and measures how often the LLM recommends the fake product, across 225 real products in 15 categories and 5 consumer scenarios.

Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion
Block3D:通过块级扩散实现高效文本到 3D 生成
arXiv:2608.19567 多模态 方法 OA · 绿色 被引 1 · S2

本文提出 Block3D,一种块级扩散框架,将离散形状 token 序列划分为连续块,自回归地生成各块,并联合去噪当前块内的所有 token,同时引入置信度引导的块内修正机制,在每块定稿前对低置信度 token 进行修订。Block3D is proposed, a block-wise diffusion framework that partitions the discrete shape-token sequence into contiguous blocks, generates the blocks autoregressively, and jointly denoises all tokens within the current block and introduces confidence-guided intra-block correction, which revises low-confidence tokens before each block is finalized.

Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision
释放图像编辑潜力:概念缩放与密集监督
arXiv:2608.16812 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文建立了一个包含超过 1,000 个细粒度编辑概念的综合性层次化分类体系,并提出一种密集监督训练策略,将多个互不干扰的概念合成到单个图像对中,显著提升了训练效率和模型整体性能。A comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts is established and a dense supervision training strategy that synthesizes multiple non-interfering concepts into single image pairs is proposed that significantly enhances both training efficiency and overall model performance.

ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction
ARC:开放式真实交互中的公平相对优势比较
arXiv:2608.13622 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ARC(Advantage Regularization via Conditioning),一种通过策略条件化 rollout 分组来恢复更公平的相对比较、并结合混合奖励与熵正则化的训练方法。The proposed ARC (Advantage Regularization via Conditioning), a training recipe that restores fairer relative comparison through strategy-conditioned rollout grouping, together with hybrid rewards and entropy regularization, is proposed.

The Compaction Cliff in Long-Running AI Agent Memory
长时运行 AI Agent 记忆中的压缩悬崖
arXiv:2608.22752 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本文提出 Knowledge Triage 框架,对 agent 知识库的每一行按类型分类,并为每种类型配置独立的保留策略,同时开源发布 AgentArtifactCorpus 数据集、分类器及参考实现。Knowledge Triage, a framework that classifies each line of an agent's knowledge base by type and routes each type through its own retention policy, is addressed, and AgentArtifactCorpus, the classifier, and the reference implementation are released.

8. TrustMargin:RAG 答案级仲裁框架
arXiv:2606.08397 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 TRUSTMARGIN,一种免训练、即插即用的仲裁层,利用模型自身的似然对两个候选进行打分,在不微调、无需外部评判或额外生成的情况下,在直接回答与 RAG 之间进行选择。TRUSTMARGIN is proposed, a training-free, plug-and-play arbitration layer that scores the two existing candidates with the model's own likelihoods and selects between Direct and RAG without fine-tuning, external judges, or additional generation.

Towards a Densing Law for User Representation Learning at Billion-Scale Capacity
迈向十亿级容量用户表示学习的稠密定律
arXiv:2608.23392 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 User Behavioral Densing Law,为大规模用户表示学习中的 tokenization 配置选择提供实用指导,并开发了 ALGN——一种自适应变长 tokenization 方法,可改善容量分配。The proposed User Behavioral Densing Law is proposed, providing practical guidance for tokenization configuration selection in large-scale user representation learning and ALGN, an adaptive variable-length tokenization method that improves capacity allocation, is developed.

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks
LongWoF-Bench:评估 EvoMap Gene 的可验证长工作流任务基准
arXiv:2608.23200 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

EvoMap 的结果表明,经过验证的执行经验可以被保留并共享为可复用的外部资源,使模型能够提升长工作流完成度,而无需反复承担经验探索的全部成本。EvoMap results show that verified execution experience can be retained and shared as a reusable external resource, enabling models to improve long-workflow completion without repeatedly paying the full cost of experience discovery.

AutoResearch: Insight In, Hallucination Out
AutoResearch:洞察输入,幻觉输出
arXiv:2608.17906 Agent 智能体 方法 OA · 绿色 被引 1 · S2

介绍 AutoResearch,一个连接 Idea Generation 与 Idea Execution 的两阶段系统,分别解决研究思路如何形成与如何通过实验可靠验证的问题,展示「实验前先夯实洞见、接受前先夯实结论」的研究流程。AutoResearch is introduced, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation to demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance.

EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment
EXPL-FR:通过视觉-语言对齐解释人脸识别模型
arXiv:2608.21486 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

覆盖 4 个 FR backbone 与 2 个 VLM 编码器;EXPL-FR 无需访问模型架构,支持身份级、单图及差异式解释,并在三种监督设置(人工标注、VLM 伪标签、完全 prompt 驱动的审计)下针对真实核验行为进行属性级审计基准测试。This work covers four FR backbones and two VLM encoders, EXPL-FR needs no architecture access, and supports identity-level, per-image, and differential explanations, and benchmark attribute-level auditing under three supervision settings, human labels, VLM pseudo-labels, and the authors' fully prompt-driven audit, against real verification behavior.

Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav
直接语料交互中的证据盲区:基于 AtlasNav 的持久化导航
arXiv:2608.24764 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

AtlasNav 减少了 Evidence Blindness,更早实现完整证据,在 PhantomWiki 上对语料结构与规模变化保持鲁棒,并在异构企业数据上取得领先性能。AtlasNav reduces Evidence Blindness, reduces Evidence Blindness, realizes complete evidence earlier, remains robust to corpus-structure and scale shifts on PhantomWiki, and achieves leading performance on heterogeneous enterprise data.

Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training
Game2World Engine:解锁真实游戏视频用于世界模型训练
arXiv:2608.24680 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 GameCleaner,一个无需 mask 的游戏 UI 移除模型,结合多模态语义理解与视频编辑能力,整体 VideoReward 较在带 UI 数据上训练的模型提升 6.83%。GameCleaner is proposed, a mask-free gameplay UI removal model that combines multimodal semantic understanding with video editing capabilities and improves overall VideoReward by 6.83% over those trained on UI-overlaid data.

Length-Adaptive Decoding for Masked Diffusion Machine Translation
掩码扩散机器翻译的长度自适应解码
arXiv:2608.22274 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 Entropy-Valley (EV):一种无需训练的画布长度选择器,通过 all-mask 前向的预测平均熵对候选目标画布打分,并挑选出 backbone 最「准备好」填充的画布。Entropy-Valley (EV), a training-free length selector that scores candidate target canvases by mean predictive entropy from all-mask forward passes and selects the canvas the backbone is most prepared to fill, is introduced.

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs
以标注作为 Rollout:面向视频 MLLMs 的高效可扩展强化学习
arXiv:2608.20492 多模态 方法 OA · 绿色 被引 1 · S2

本文研究视频 MLLM 的 RL 后训练样本效率与可扩展性,并提出 OraRL——一种随模型规模与数据规模共同 scaling 的解耦 advantage estimator,在 0.8B 到 9B backbone 上均超越其基线,并在 100k prompts 规模下超越 GRPO。The sample efficiency and scalability of RL post-training for video MLLMs and introduces OraRL, a decoupled advantage estimator that scales with model size and data, surpassing its backbone from 0.8B to 9B and GRPO up to 100k prompts.

CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild
CyberFactory:基于真实实例扩展网络安全能力
arXiv:2608.23181 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 CyberFactory,一个统一的开源框架,贯通 PoC 生成、漏洞修补与 CyberQA 三大任务中的数据构建、轨迹合成与模型训练。CyberFactory is introduced, a unified open-source framework that connects data construction, trajectory synthesis, and model training across proof-of-concept (PoC) generation, vulnerability patching, and cybersecurity question answering (CyberQA).

DREAM Technical Report
DREAM 技术报告
arXiv:2608.09408 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 DREAM,一种自主优化控制架构:在不替换现有流水线的前提下叠加感知可感知、可编排、可审计的策略层,支持将 agentic meta-control 作为工业推荐的一种可行范式。This work presents DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them, supporting agentic meta-control as a viable paradigm for industrial recommendation.

When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows
当 "Must" 变成 "Maybe":LLM Agent 工作流中的约束弱化
arXiv:2608.24569 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本文指出 LLM agent 中信息抽取与动作之间的 state-transmission 失效,并展示 handoff 变换如何在保留状态内容的同时削弱其对下游动作的约束。This work identifies a state-transmission failure between information extraction and action in large language model agents, and shows how handoff transformations can retain state content while weakening its constraints on downstream action.

MoTE: Mixture of Task Experts for Multi-Task Video Understanding
MoTE:面向多任务视频理解的 Task Expert 混合模型
arXiv:2608.24763 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MoTE(Mixture of Task Experts),一种将大语言模型前馈网络转化为任务特定专家同时保持多模态 backbone 共享的 decoder 架构,并在五个 COIN 基准上使用显式任务路由进行评估。This work proposes MoTE (Mixture of Task Experts), a decoder architecture that converts large language model feed-forward networks into task-specific experts while keeping the multimodal backbone shared and evaluates it on five COIN benchmarks using explicit task routes.

AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace
AgentRoom:基于 CRDT 共享工作空间的并发多 Agent 编程
arXiv:2608.23740 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

AgentRoom 是一种面向并发编码 Agent 的实时协同编辑协议,通过在 CRDT 合并的共享文件系统上将文件级 claim、status 和 broadcast 暴露为 MCP 工具,且运行间的差异小于 CLI-stable 模型。AgentRoom is a realtime collaborative editing protocol for concurrent coding agents that exposes file-level claim, status, and broadcast as MCP tools on a CRDT-merged shared filesystem and has less run-to-run variation than CLI-stable models.

7. Triton Attention Kernel 学术分析 (arXiv 2511.11581)
Triton Attention Kernel 学术分析 (arXiv 2511.11581)
arXiv:2511.11581 LLM 基础设施 方法 OA · 绿色 被引 4 · S2

本工作开发了一个 SOTA 的 paged attention kernel,完全基于领域特定即时编译语言 Triton 构建,在 NVIDIA 与 AMD GPU 上均达到 SOTA 性能。This work develops a state-of-the-art paged attention kernel that builds exclusively on the domain-specific just-in-time compiled language Triton to achieve state-of-the-art performance on both NVIDIA and AMD GPUs.

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
开放世界多 Agent 环境中的自主数学发现
arXiv:2608.23691 Agent 智能体 方法 OA · 绿色 被引 2 · S2

The Station 被评估为一个开放世界多 agent 环境,不同模型家族的 AI agent 在其中无需中心协调器或脚本化流程即可共同追求同一研究目标,并提供发现产生过程的透明记录。The Station is evaluated, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline, providing a transparent record of how discoveries emerged.

SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation
SecOPD:通过 On-Policy Distillation 缓解自适应 Prompt 注入
arXiv:2608.21500 Agent 智能体 方法 OA · 绿色 被引 4 · S2

本文提出 Secure On-Policy Distillation (SecOPD),提供 token 级反馈以指导防御性微调,并能泛化到训练中完全未见过的领域。This paper proposes Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning, and generalizes to domains completely unseen in training.

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
GigaBrain-0.7:以三系统架构将具身基础模型扩展至涌现能力
arXiv:2608.15875 多模态 方法 OA · 绿色 被引 6 · S2

本文提出 GigaBrain-0.7,一种跨多种机器人 embodiment 泛化能力显著增强的 embodied foundation model,并引入一阶段对齐训练,联合优化 vision-language 理解和多 embodiment 动作生成。This work presents GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation.

Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation
以 Rubric 作为视觉修复上下文以实现自演化的 UI-to-Code 生成
arXiv:2608.24138 多模态 方法 OA · 绿色 被引 1 · S2

评估表明,RubSE 在终轮和最佳轮设置下均显著优于朴素 self-evolution,refinement 轨迹更稳定,且轨迹级性能上限更高。Evaluations demonstrate that RubSE substantially outperforms na\"ive self-evolution in final-round and best-round settings, achieving more stable refinement trajectories and a higher trajectory-level performance ceiling.

The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents
交接代价:LLM Agent 中非原生轨迹的延续
arXiv:2608.24358 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文变化 handoff 方向、时机与接口,对比保留仓库状态下的全轨迹传输、压缩与轨迹移除,发现偏好接口随方向反转:减少 LC-model 轨迹信息可提升 escalation 质量,而移除 HC-model 轨迹则会降低 downshift 质量。This work varies handoff direction, timing, and interface, comparing full-trajectory transfer, compaction, and trajectory removal while preserving the repository state, and finds that the preferred interface also reverses with direction: reducing LC-model trajectory information improves escalation quality, whereas removing the HC-model trajectory reduces downshift quality.

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
面向持久化故事与交互式世界的长时音视频生成
arXiv:2608.23383 多模态 方法 OA · 绿色 被引 2 · S2

结果表明,记忆、几何控制以及 rollout-aware 训练为生成连贯故事和持续演化的交互式世界提供了实用基础。Results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds.

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data
Next-chunk reasoning RL 真的优于 SFT 吗?——在 no-CoT 数据下重新审视训练策略
arXiv:2608.23256 工程化 方法 OA · 绿色 被引 1 · S2

Mixed SFT 是一种单阶段监督微调,联合在 no-CoT 和 long-CoT 数据上训练,相比 next-chunk reasoning RL 取得了明显更高的 RLVR 后性能上限,同时训练算力开销减少超过 60 倍。Mixed SFT, a single supervised fine-tuning stage that jointly trains on no-CoT and long-CoT data, achieves a clearly higher post-RLVR performance ceiling than next-chunk reasoning RL while requiring over 60 times less training compute.

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling
FIRM-Video:先核查再评分,实现可靠的 text-to-video 奖励建模
arXiv:2608.21839 多模态 方法 OA · 绿色 被引 1 · S2

本文提出 FHM-Video,一种基于 check-before-score 原则的、由 checklist 驱动的统一数据构建框架,在 FIRM-Video-Bench 上取得最佳综合 MAE,并在三种视频生成器的 Best-of-8 采样中始终获得最高的 VBench Total、Quality 和 Semantic Score。FHM-Video, a unified checklist-driven data construction framework based on a check-before-score principle, is introduced, which achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.

Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction
Real-TurnTurk:用于话轮预测的多模态土耳其语语料库
arXiv:2608.22071 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文发布一个土耳其语多模态对话数据集,包含未脚本化的双人交互,并提供同步的前向视频、可归属到每位说话人的独立音频通道以及时间对齐的转写文本。A multimodal Turkish conversational dataset of unscripted dyadic interactions, comprising synchronized front-facing video, per-speaker audio channels that allow overlapping speech to be attributed to individual speakers, and time-aligned transcriptions is introduced.

7. LLM 压缩:联合剪枝 + 混合精度 PTQ
arXiv:2606.07819 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本工作提出一种新颖的混合精度 PTQ 策略,直接最小化整个模型的全局误差传播,而非孤立地处理逐层误差,并开发了一种新颖的联合优化方法,在统一搜索空间中同时学习结构化剪枝决策与混合精度量化策略。This work proposes a novel mixed-precision PTQ strategy that directly minimizes global error propagation across the entire model, rather than isolating layer-wise errors, and develops a novel joint optimization approach that simultaneously learns structural pruning decisions and mixed-precision quantization policies within a unified search space.

Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models
Stream4D:面向流式自回归扩散视频模型的 4D 一致性
arXiv:2608.19556 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文用前馈 4D 重建奖励替代静态 critic,显式建模场景动态,使连贯运动获得高一致性奖励,并加入对自然 scene-flow 幅值进行奖励同时抑制抖动与非刚性伪影的运动先验。This work replaces the static critic with a feed-forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards, and adds a motion prior that rewards natural scene-flow magnitude while penalizing jitter and non-rigid artifacts.

Multi-Granularity Context-Enhanced RAG over Multimodal Knowledge Graphs
多模态知识图谱上的多粒度上下文增强 RAG
arXiv:2608.25986 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种构建 Context-Enhanced MMKG (CEMMKG) 的新框架,能够有效利用上下文信息提升基于 MMKG 的 RAG 性能,并在不同基于 MMKG 的 RAG 方法上的有效性验证了其广泛适用性。A novel framework for constructing a Context-Enhanced MMKG (CEMMKG) is proposed, effective in leveraging contextual information to improve MMKG-based RAG performance and its effectiveness across different MMKG-based RAG methods demonstrates its broad applicability.

RetrievalRouter: Joint Modality and Architecture Selection for Document Retrieval
RetrievalRouter:面向文档检索的模态与架构联合选择
arXiv:2608.25625 RAG 检索增强 方法 被引 0 · S2

RetrievalRouter 是一种轻量级 query-aware router,仅依据 query 文本即可学习最匹配的检索 pipeline,在面向准确率的设置下 nDCG@5 显著更高,而在面向延迟的设置下,nDCG@5 和延迟均匹配或数值上优于基线。RetrievalRouter is a lightweight query-aware router that learns, from the query text alone, which retrieval pipeline best fits each query, and achieves significantly higher nDCG@5 across accuracy-oriented settings, while matching or numerically outperforming them on both nDCG@5 and latency in latency-oriented settings.

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
LibriBrain100:面向大规模神经语音解码的百小时广深 MEG 数据集
arXiv:2608.25204 多模态 方法 OA · 绿色 被引 4 · S2

本文发布 LibriBrain100,一个面向语音解码的大规模 MEG 数据集,从设计上保证可复现、标准化评估,并展示了广泛多被试数据的价值:对预训练模型进行有监督微调可大幅弥补单被试数据不足。LibriBrain100 is introduced, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation and the value of broad multi-subject data is demonstrated: supervised finetuning of a pre-trained model can substantially compensate for limited per-subject data.