研究库 论文知识库
Papers · organized/paper_cards

论文

1640 张论文卡片 · OA 绿色

开放获取 全部 绿色 · 1640
GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding
GUI-Primitives:诊断视觉语言 GUI grounding 中的空间推理失败
arXiv:2608.21832 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

计算机使用 Agent 将自然语言指令 grounding 到截图中以定位界面元素,但现有基准无法隔离模型是否将关系语言正确绑定到对应元素。我们提出 GUI-Primitives,一个包含 994 个条目的基准,由跨七种空间关系(左右、上下、包含、对齐、邻近、列表序数、遮挡)的对比指令对组成。每对保持截图和锚点不变,仅改变关系表达,使正确目标在两个指定候选之间切换。五位标注者……Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether models bind relational language to the correct element. We introduce GUI-Primitives, a 994-item benchmark of contrastive instruction pairs over seven spatial relations in graphical user interfaces (left/right, above/below, containment, alignment, proximity, list ordinal, occlusion). Each pair holds the screenshot and anchor fixed while changing the relation expression, so the correct target moves between two designated candidates. Five annotators

A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans
一种面向 CT 扫描中空间关系验证的、可靠且可审计的模块化 Agent
arXiv:2608.21140 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个模块化的医学影像 Agent,用于轴位 CT 切片中的二元空间关系验证,采用显式模块化空间验证阶段,并表明显式模块化空间验证可作为未来面向报告的医学影像 Agent 的有前景的构建模块。This work presents a modular medical imaging agent for binary spatial relation verification in axial CT slices using explicit modular spatial verification stages, and suggests that explicit modular spatial verification can serve as a promising building block for future report-oriented medical imaging agents.

6. QBugLM:量子软件调试多智能体框架
arXiv:2606.07314 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本工作提出 QBugLM,一个多 Agent 框架,可自动化量子软件调试流水线,覆盖基于分类法的缺陷注入、基于 LLM 的检测与修复,直至基于仿真的验证,框架无关地支持 OpenQASM 3.0 程序。This work proposes QBugLM, a multi-agent framework that automates the quantum software debugging pipeline, from taxonomy-driven bug injection to LLM-based detection and repair, and finally to simulation-based validation, for framework-agnostic OpenQASM 3.0 programs.

TTPO: Test-Time Policy Optimization
TTPO:测试时策略优化。
arXiv:2608.27448 工程化 方法 OA · 绿色 被引 1 · S2

提出 Test-Time Policy Optimization,一种非对称目标,通过 OPSD 蒸馏一致性 rollout,并使用 Grouped RL 惩罚不一致的 rollout;进一步通过 token 级选择精炼两个分支:蒸馏降低已收敛位置的权重,而 RL 仅惩罚置信的错误。Test-Time Policy Optimization is proposed, an asymmetric objective that distills agreeing rollouts via OPSD and penalizes disagreeing rollouts with Grouped RL and Token-level selection further refines both branches: distillation down-weights already-converged positions, while RL penalizes only confident errors.

PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
PILOT in the Loop:面向长时序 Agent 的实时自我改进。
arXiv:2608.26530 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本文提出 PILOT,一种通过两种耦合机制实现实时自我改进的 supervisor-worker 框架:(1) live steering 允许独立的 supervisor 在执行期间重定向或中止当前 worker;(2) live self-evolution 将执行中发现的过程与失败模式提炼为可复用的 skills 与记忆。PILOT is presented, a supervisor-worker harness for live self-improvement through two coupled mechanisms: (1) live steering lets a separate supervisor redirect or abort the active worker during execution; and (2) live self-evolution distils procedures and failure modes revealed during execution into reusable skills and memory.

Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning
Thinking on Shots:基于 Agentic 推理的一致性多镜头视频编辑。
arXiv:2608.26809 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个 Agentic 视频编辑框架,利用 LLM 与 VLM 的协同实现 shot 级视频解耦与精确指令解析,并构建 MMLVE-Bench,一个聚焦 MMLVE 的数据集,具有复杂的真实世界时空动态、高密度异构指令以及稀疏随机的实体分布。This work introduces an agentic editing framework that leverages the synergy of Large Language Models (LLMs) and Vision-Language Models (VLMs) to achieve shot-level video decoupling and precise instruction parsing, and constructs MMLVE-Bench, which is an MMLVE-focused dataset characterized by complex real-world spatiotemporal dynamics, high-density heterogeneous instructions, and sparse, random entity distributions.

Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning
Aphanta:诊断任务对齐的图像编辑中间态以服务多模态推理。
arXiv:2608.26993 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果将图像编辑定位为一种专门的视觉工作空间而非通用推理机制,并将 Aphanta 确立为可复用的协议,用于度量任务-表征对齐、编辑器实现及下游 pipeline 实用性。The results position image editing as a specialized visual workspace rather than a universal reasoning mechanism, and establish Aphanta as a reusable protocol for measuring task--representation alignment, editor realization, and downstream pipeline utility.

Procedura: Agentic 3D Modeling with Procedural Control
Procedura:基于过程化控制的 Agentic 三维建模
arXiv:2608.26238 Agent 智能体 观点 OA · 绿色 被引 1 · S2

本文探索"以代码表示 3D 形状"的范式,利用并放大 LLM 的编码能力进行 3D 建模,并提出 Procedura 框架——通过编写一个由命名零件构成、并通过类型化、可机器校验的连接关系装配而成的参数化程序来建模对象。The paradigm of 3D shape as code is explored, leveraging and scaling the coding ability of an LLM for 3D modeling, and Procedura, a novel 3D modeling agent framework that writes an object as a procedural assembly, a parametric program whose named parts are joined by typed, machine-checkable mates.

Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
将 Agentic 游戏开发作为可扩展世界模型的可验证轨迹数据引擎
arXiv:2608.25518 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出 Reinforcement Learning with Human-Engine Verification(RLHEV),一种结合稠密引擎信号与开发过程中隐式人类接受反馈的后训练范式,用于支持强化学习后训练。This work proposes Reinforcement Learning with Human-Engine Verification (RLHEV), a post-training paradigm that combines dense engine signals with implicit human acceptance feedback from the development process to support RL post-training.

CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval
CaSKG:用于可扩展 Agent 技能检索的反事实-因果技能图
arXiv:2608.25500 RAG 检索增强 方法 OA · 绿色 被引 3 · S2

提出 CaSKG,一种反事实因果 Skill 图谱框架,在检索前校准程序关系,将边置信度校准定位为大规模紧凑且可执行的 Skill 检索的有效路径。CaSKG, a counterfactual-causal skill graph framework that calibrates procedural relations before retrieval, is proposed, position edge-confidence calibration as an effective route to compact and executable skill retrieval at scale.

GameWAM: A World Action Model for Video Games
GameWAM:面向视频游戏的世界动作模型
arXiv:2608.26200 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 GameWAM,据其所知是首个面向原生闭环游戏与 GUI 控制的 WAM,并发现 Low-Frequency Action Source Imprinting (LASI):在固定条件下,采样动作源的低频分量系统性地引导生成的粗粒度相机运动,揭示了生成式控制中的源敏感性失效模式。This work introduces GameWAM, to its knowledge the first WAM for native closed-loop gameplay and GUI control and uncovers Low-Frequency Action Source Imprinting (LASI), in which low-frequency components of the sampled action source systematically steer coarse generated camera motion under fixed conditioning, revealing a source-sensitivity failure mode in generative control.

Scaling Laws for Neural Language Models
神经语言模型的 Scaling Laws
arXiv:2001.08361 LLM 基础设施 方法 OA · 绿色 被引 9195 · S2

更大的模型显著更具样本效率,因此最优的算力高效训练方式是:在相对适中的数据量上训练非常大的模型,并在远未收敛时显著提前停止训练。Larger models are significantly more sample-efficient, such that optimally compute-efficient training involves training very large models on a relatively modest amount of data and stopping significantly before convergence.

CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes
CritICL:基于小语言模型失败模式的推理时弱到强泛化
arXiv:2608.27455 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

实验结果表明,CritICL 持续优于标准 in-context learning,并以显著更少的生成次数和更低的 token 成本,取得与 test-time scaling 方法相当或更优的性能。Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost.

EditaLive! Unified Character Video Editing for Live Streaming
EditaLive! 面向直播的统一人物视频编辑
arXiv:2608.27123 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

传统视频编辑主要关注场景级内容,而直播更强调人物主体。然而,直接将现有视频编辑方法应用于以人为中心的直播仍具挑战,因为它们可能引入面部表情不一致,且通常依赖多个离线推理步骤,难以满足实时交互需求。我们提出 EditaLive,一个用于实时流式角色视频编辑的新型框架。具体而言,我们从预训练图像动画模型(Wan-Animate)出发,该模型天然解耦...Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically depend on multiple offline inference steps, making them unsuitable for real-time interaction. We propose EditaLive, a novel framework for real-time streaming character video editing. In detail, we start from a pretrained image animation model (Wan-Animate), which naturally decoupl

TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback
TacForcing:融合执行时间触觉反馈的流式动作生成
arXiv:2608.25798 多模态 方法 OA · 绿色 被引 1 · S2

提出 TacForcing,一种融合执行期触觉反馈的流式动作生成框架:按顺序生成动作块并保留未完成块的中间状态,并引入执行感知触觉注意力(EATA),将每次触觉更新的直接访问限制在下一个即将执行的块。TacForcing is introduced, a streaming action-generation framework incorporating execution-time tactile feedback that generates action blocks sequentially while preserving intermediate states of unfinished blocks and introduces Execution-Aware Tactile Attention (EATA), which restricts direct access to each tactile update to the next block scheduled for execution.

Luce: Relightable Gaussians for 3D Asset Generation
Luce:用于 3D 资产生成的可重光照高斯表示
arXiv:2608.23943 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Luce,一种三维表示方法,将几何与 PBR 材质统一到体素化的多模态高斯云中,分别使用专用高斯基元表示反照率、金属度-粗糙度与法线,在单图到三维生成任务上达到 SOTA。Luce is proposed, a 3D representation that unifies geometry and PBR materials within a voxelized multimodal Gaussian cloud, using dedicated Gaussian primitives for albedo, metallic-roughness, and surface normals to achieve state-of-the-art single-image-to-3D generation.

What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals
评估授权了什么?Inspect Evals 中声明相对推理的提交绑定普查
arXiv:2608.19269 LLM 基础设施 观点 OA · 绿色 被引 0 · S2 + OpenAlex

冻结一个大规模评测集合,并尝试对其历史断言进行 replay,以使这一原本隐式的推理步骤变得显式且可执行。A large evaluation collection is frozen and an attempt to replay its historical claims is made to make this otherwise implicit inference step explicit and executable.

LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation
LayerRecall:面向视频生成长程一致性的状态条件记忆路由器。
arXiv:2608.28460 多模态 方法 OA · 绿色 被引 2 · S2

提出 LayerRecall,一种由当前状态条件化、按层选择的 memory router,仅从历史 K/V state 中检索相关状态并注入到 backbone 特定的 memory-sensitive 层中,其余层保留局部 attention;并提出 Cross-Horizon Prediction Matching (CHPM),借助特权的长上下文参考在预测空间中对有界 memory router 进行监督。This work introduces LayerRecall, a current-conditioned, layer-selective memory router that retrieves relevant historical K/V states and injects them only into backbone-specific memory-sensitive layers while preserving local attention elsewhere, and proposes Cross-Horizon Prediction Matching (CHPM), which uses a privileged long-context reference to supervise the bounded-memory router in prediction space.

CrabOS: An Operating System for Human-AI Co-inhabitation
CrabOS:面向人类与 AI 共生的操作系统。
arXiv:2608.28165 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

CrabOS 将人机交替主导复杂任务的支持从依赖桥接的应用层方案提升为原生操作系统能力,为开发与运行 AI agent 提供了新基础。CrabOS elevates support for complex tasks with alternating human and AI leadership from bridge-dependent application-level solutions to native operating-system capabilities, which provide a new foundation for developing and running AI agents.

J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data
J-Zero:从零数据出发的统一 Challenger--Solver--Judge 协同进化
arXiv:2608.26582 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Judge co-adaptation from Zero data (J-Zero),一个统一的 Challenger–Solver–Judge 自进化框架,支持在可验证与不可验证领域的自我提升,并识别出 Judge 协同进化是这一持续改进的关键驱动力。This work proposes Judge co-adaptation from Zero data (J-Zero), a unified Challenger--Solver--Judge self-evolution framework that supports self-improvement across both verifiable and unverifiable domains and identifies Judge co-adaptation as the key driver of this sustained improvement.

Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion
Ring Forcing:面向自回归视频扩散的精确长时记忆
arXiv:2608.26794 多模态 方法 OA · 绿色 被引 5 · S2

提出 Ring Forcing,一个自回归视频扩散框架,旨在稳健构建并精确利用长期记忆,并通过稀疏 RoPE 机制实现灵活、可扩展的记忆适配,同时充分利用预训练先验。Ring Forcing is presented, an autoregressive video diffusion framework designed to robustly construct and precisely utilize long-term memory and a sparse RoPE mechanism to enable flexible, scalable memory adaptation while fully exploiting pre-trained priors.

Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents
画出所见:多模态 Agent 中灵巧视觉工具使用基准评测
arXiv:2608.25417 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 EASEL benchmark,用于评估受控的灵巧视觉工具使用,以参考引导的视觉重建为主代理任务:agent 逐步在画布上绘制以匹配参考图像。EASEL is proposed, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image.

6. LLM Research Papers: The 2026 List (Jan–May) — Sebastian Raschka
LLM 研究论文:2026 年清单(1—5 月)— Sebastian Raschka
arXiv:2601.21204 RAG 检索增强 方法 Open MIND OA · 绿色 被引 14 · S2

本工作将 embedding 缩放作为正交于稀疏度缩放的强有力维度加以探索,并推出 LongCat-Flash-Lite,一个从零训练的 68.5B 参数、约 30 亿激活参数的模型,不仅超越参数等量级的 MoE 基线,还对同规模现有模型展现出卓越竞争力。This work explores embedding scaling as a potent, orthogonal dimension for scaling sparsity and introduces LongCat-Flash-Lite, a 68.5B parameter model with ~3B activated trained from scratch that not only surpasses parameter-equivalent MoE baselines but also exhibits exceptional competitiveness against existing models of comparable scale.

StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
StepGuard:基于可扩展监督与安全-效用平衡的步骤级护栏学习
arXiv:2608.24777 安全与风险 方法 OA · 绿色 被引 5 · S2

提出 StepGuard,一种 step-level guard model,可对已完成的 agent trajectory 进行审计并在工具动作执行前进行检查;并引入 StepGen,一种自动数据引擎,能在风险步生成上下文相同但动作不同的安全与不安全 trajectory。This work proposes StepGuard, a step-level guard model that can audit completed agent trajectories and check tool actions before they are executed, and introduces StepGen, an automatic data engine that generates safe and unsafe trajectories with the same context but different actions at the risky step.

StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments
StarHarness:通过分层搜索为企业环境进化 Harness
arXiv:2608.24804 评测基准 方法 OA · 绿色 被引 1 · S2

StarHarness 提供了一种实用方法,通过根据基线失败行为对任务进行分层、将 proposer 可见的搜索任务与 proposer 隐藏的选择任务分离,并为评估泛化能力保留 held-out 任务,从而缓解工具密集型企业任务中持续的 model-environment mismatch。StarHarness offers a practical way to reduce persistent model-environment mismatch in tool-rich enterprise tasks by stratifying tasks according to baseline failure behavior, separating proposer-visible search tasks from proposer-hidden selection tasks, and reserves held-out tasks for evaluating generalization.

LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering
LoopArena:将模型作为 Loop 工程运行时控制器的基准评测
arXiv:2608.28281 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

提出 LoopArena,一个用于评估一个模型在长时任务中引导另一个独立 coding agent 能力的 benchmark,并在执行范围与成本各不相同的三个互补设置下评估该能力。This work introduces LoopArena, a benchmark for evaluating how well one model can guide a separate coding agent through a long-running task, and evaluates this ability in three complementary settings that differ in execution scope and cost.

Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities
Agentic Artifact Creation:系统、评估、原则与机遇
arXiv:2608.28122 Agent 智能体 综述 OA · 绿色 被引 0 · S2 + OpenAlex

本文综述了 agentic artifact creation,即有状态的构建过程,其中 AI 系统实体地构造或修改交付物,且中间观察结果会重定向后续工作;并提出了若干原则,使承诺与责任显式化、将反馈转化为针对性修复,以及在变更后重新验证受影响的状态。This survey examines agentic artifact creation, which is defined as stateful construction in which an AI system materially constructs or revises a deliverable and intermediate observations redirect later work, and formulate principles for keeping commitments and responsibility explicit, turning feedback into targeted repair, and revalidating affected state after change.

Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding
在视频中定位任意目标:重新思考高效生成式时空视频定位
arXiv:2608.28192 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Decoupled Block Attention,在保留共享 video-query 上下文访问的同时消除跨 box 依赖,并结合用于时间边界与空间几何的 localization-aware policy optimization;并行 tube generation 被证明是视频中 autoregressive 定位的一种高效替代方案。Decoupled Block Attention is introduced, which preserves access to shared video-query context while eliminating cross-box dependencies, together with localization-aware policy optimization for temporal boundaries and spatial geometry, and parallel tube generation is shown to be an efficient and effective alternative to autoregressive localization in videos.

Fast Weight Attention for Continual Learning
用于持续学习的快速权重注意力
arXiv:2608.27763 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该框架将循环序列模型中的 temporal alignment、plasticity、forgetting 与 bounded rehearsal 分离,并结合数值稳定的 positive-decay renormalization,使其在语言建模上保持竞争力,同时提升在可变位数加法任务上的长度外推能力。This framework separates temporal alignment, plasticity, forgetting, and bounded rehearsal in recurrent sequence models, together with numerically stable positive-decay renormalization, to remain competitive in language modeling and improve length extrapolation on variable-digit addition.

GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models
GGSS:用于生成式视觉-语言模型推理时去偏的测地门控球面引导
arXiv:2608.25375 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

GGSS——Geodesic-Gated Spherical Steering——一种保范的干预方法,在单位超球面上发现反事实偏差子空间,沿 geodesic arc 引导视觉 token,并通过自适应门控聚焦于承载更强人口统计信号的 token。GGSS---Geodesic-Gated Spherical Steering---a norm-preserving intervention that discovers a counterfactual bias subspace on the unit hypersphere, steers visual tokens along geodesic arcs, and uses an adaptive gate to focus correction on tokens that carry stronger demographic signal.

Language Chain in Alignment: Cross-lingual Ranking Preference Optimization
对齐中的语言链:跨语言排序偏好优化
arXiv:2608.23149 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Cross-lingual Ranking Preference Optimization (CRPO),一种新框架,利用来自英语的鲁棒偏好知识来促进目标语言的偏好对齐,从而增强语言适应性与输出质量。This paper proposes Cross-lingual Ranking Preference Optimization~ (CRPO), a novel framework that leverages robust preference knowledge from English to facilitate preference alignment in the target language, thereby enhancing language adaptation and output quality.

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models
按意图行动:面向视觉-语言-动作模型的行为意图蒸馏
arXiv:2608.23478 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Intention Distillation (INDI),将行为级意图蒸馏到动作解码器中,并以目标依赖的方式组织下游预测;研究表明,动作解码器显式建模其生成行为的语义目标能够带来收益。Intention Distillation (INDI) is proposed, which distills behavior-level intent into the action decoder and organizes downstream predictions in an objective-dependent manner, and shows that action decoders benefit from explicitly modeling the semantic objective of the behavior they generate.

DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents
DART-SD:面向多轮工具调用 Agent 自蒸馏的菱形拓扑感知检索与微调
arXiv:2608.18524 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本文提出 DART-SD (Diamond-topology Aware Retrieval and Tuning for Self-Distillation),一种将范式从全局 forcing 转向拓扑引导的局部修正的新框架,相对传统 full-trajectory 基线取得显著提升。This work proposes DART-SD (Diamond-topology Aware Retrieval and Tuning for Self-Distillation), a novel framework that shifts the paradigm from global forcing to topology-guided localized correction, and significantly outperforms traditional full-trajectory baselines.

6. LLM Research Papers: The 2026 List (Jan–May) — Sebastian Raschka
LLM 研究论文:2026 年清单(1—5 月)— Sebastian Raschka
arXiv:2602.08071 多模态 方法 Open MIND OA · 绿色 被引 13 · S2

ViT-5 的设计与当代基础模型实践保持一致,可作为对 vanilla ViT 的直接替换升级方案,适用于 2020 年代中期的视觉骨干网络,并为生成建模提供更强大的骨干。With a design aligned with contemporary foundation-model practices, ViT-5 offers a simple drop-in upgrade over vanilla ViT for mid-2020s vision backbones and serves as a stronger backbone for generative modeling.

LINE Conversation History Retrieval for Personal Memory RAG: Evaluating Search Representations and Hybrid Retrieval
基于 LINE 对话历史的个人记忆 RAG 检索:搜索表示与混合检索评估
arXiv:2608.27809 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该研究将 358,896 条消息切分为 22,329 个时间连贯的块,并构建三种搜索表示:raw_text、生成的 summary,以及将 summary 与 raw_text 片段及其他固定文本相结合的 embedding_text。This study segmented 358,896 messages into 22,329 temporally coherent chunks and constructed three search representations: raw_text, a generated summary, and embedding_text, which combines a summary with a raw-text excerpt and other fixed text.

CamoDocs: A Poisoning Attack Against Retrieval-Augmented Language Models Using Camouflaged Documents
CamoDocs:利用伪装文档针对检索增强语言模型的投毒攻击
arXiv:2608.28389 RAG 检索增强 方法 OA · 绿色 被引 1 · S2

提出 CamoDocs,一种通过将对抗文档伪装在良性内容中来避免直接包含查询的投毒攻击,并表明 TrustRAG 等以擦除为主的聚类防御可降低 ASR,但会在 NeoQA 等依赖检索的基准上造成显著的效用下降。CamoDocs is proposed, a poisoning attack that avoids direct query inclusion by camouflaging adversarial documents among benign content, and shows that erasure-heavy clustering defenses such as TrustRAG can reduce ASR, but only with substantial utility drops on retrieval-dependent benchmarks such as NeoQA.