研究库 论文知识库
Papers · organized/paper_cards

论文

1640 张论文卡片 · OA 绿色

开放获取 全部 绿色 · 1640
Compile by Training: Turning Natural-Language Specifications into Local Neural Functions
Compile by Training:将自然语言规约转化为本地神经函数
arXiv:2609.04199 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

许多重复出现的文本函数易于描述却难以用规则实现;而为每个输入调用大型远程模型会带来重复开销、延迟与对服务方的依赖。我们提出 compile by training,将自然语言规约转化为可复用的神经函数。在编译时,教师模型生成任务专属样本,用于为精简解释器训练一个小型适配器。生成的函数可在没有教师模型的情况下运行,并能像普通软件一样被存储、版本化管理与组合。在 FuzzyBench-Hard 这一子集上……Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and dependency on a provider. We present compile by training, which turns a natural-language specification into a reusable neural function. At compile time, teacher models generate task-specific examples that are used to train a small adapter for a compact interpreter. The resulting function runs without the teachers and can be stored, versioned, and composed like ordinary software. On FuzzyBench-Hard, a subset on which

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding
超越检索:面向流式视频理解的渐进式潜在记忆演化
arXiv:2609.04131 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出 LatentStream,一种渐进式 latent working memory 框架,将流式记忆从"存储-检索"转变为"检索-内化",在现有在线和离线视频 benchmark 上取得新的 SOTA 结果。This work introduces LatentStream, a progressive latent working memory framework that shifts streaming memory from store-and-retrieve to retrieve-and-internalize, and achieves new state-of-the-art results on existing online and offline video benchmarks.

PACE: Towards Surfacing Hidden Conflicts in User Requests
PACE:揭示用户请求中的隐性冲突
arXiv:2609.03293 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 PACE 数据集,用于评估模型能否识别以自我中心知识或事件形式表达的潜在约束(这些约束使看似合理的用户请求变得不当),以及 PaceMaker 多 Agent 框架,其中专门 Agent 通过查询重构、多跳图遍历与冲突感知过滤进行协调,以检索上下文决定性证据。PACE is introduced, a dataset for evaluating whether models can identify latent constraints, expressed as egocentric knowledge or events, that render seemingly reasonable user requests inappropriate, and PaceMaker, a multi-agent framework in which specialized agents coordinate across query reformulation, multi-hop graph traversal, and conflict-aware filtering to retrieve contextually decisive evidence.

The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation
缺失的时间纽带:面向剧本驱动音视频生成的时间上下文路由
arXiv:2609.02367 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Temporal Context Routing,将脚本时序映射到视频与音频生成的共享时间轴上,并将每个 prompt 的引导路由到两种模态中的对应位置,同时保持与 baseline 相当的视觉质量与音视频同步性。Temporal Context Routing is introduced, which maps the script timing onto the shared temporal axis of video and audio generation and routes each prompt's guidance to the corresponding positions in both modalities, while maintaining visual quality and audio-visual synchronization comparable to those of the baselines.

Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration
改变置信度,而非改变预测:面向后验校准的预测保持型修复
arXiv:2609.01072 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 CORD,第一个 post-fit adapter,通过从 calibrator 拟合中移除 preservation constraint 来修复完整校准概率向量,从而实现精确的预测保持,并将原始决策的精确恢复交由后续输出修复完成。CORD is proposed, the first post-fit adapter to impose exact prediction preservation by repairing the full calibrated probability vector by removing the preservation constraint from calibrator fitting and assigns exact recovery of the original decision to subsequent output repair.

Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training
知道何时不复用:自主 LLM 后训练中的条件经验迁移
arXiv:2608.26730 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

将条件经验迁移进行形式化,并提出 Boundary-Calibrated Intervention Transfer,一种在权重变化的训练之前即授权经验复用的方法,在相同预算下取得比所评估替代方案更高的最终模型质量。Conditional experience transfer is formulated as conditional experience transfer and Boundary-Calibrated Intervention Transfer is introduced, a method that authorizes experience reuse before weight-changing training and attains higher equal-budget final-model quality than the evaluated alternatives.

Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction
Scal3R:学习高效多相对位姿查询以实现可扩展的在线 3D 重建
arXiv:2609.04201 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该方法将在线重建重构为多参考相对位姿查询,在 Virtual KITTI、Sintel、TUM-Dynamic、ScanNet 和 7-Scenes 上取得 SOTA 性能,并在单 GPU 上 8 小时内收敛。This approach reformulates online reconstruction as multi-reference relative pose querying, which achieves state-of-the-art performance across Virtual KITTI, Sintel, TUM-Dynamic, ScanNet, and 7-Scenes and reaches convergence in 8 hours on a single GPU.

Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance
优化中的渗流动力学:方差级联与离散尺度不变性
arXiv:2609.02373 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

通过将随机梯度流建模为渗流过程来研究随机梯度下降的动态,其中嵌套的架构对称性迫使子网络以离散块的形式合并,而非通过单边附着。The dynamics of Stochastic Gradient Descent is studied by modeling the stochastic gradient flow as a percolation process, in which nested architectural symmetries force subnetworks to merge in discrete blocks rather than by single-edge attachment.

Using Grounded Theory for Agent Behavior Analysis at Scale
大规模 Agent 行为分析中的扎根理论应用
arXiv:2608.30391 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出 AutoTraceGT(Automated Trace analysis through Grounded Theory),首个在 Agent 轨迹上自动化 grounded theory 的多 Agent pipeline,并指出 Grounded Theory 为研究 Agent 实际行为的 ML 研究者和 Agent 开发者提供了可扩展的分析工具。This work proposes AutoTraceGT (Automated Trace analysis through Grounded Theory), the first multi-agent pipeline that automates grounded theory on agent trajectories and suggests Grounded Theory offers a scalable analytic tool for ML researchers and agent developers studying what agents actually do.

QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation
QCell:重组与对齐细胞查询用于重叠实例分割
arXiv:2608.29253 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 QCell,一种新颖的基于查询的模型,用于在显微镜场景中去重叠细胞实例,在多个 benchmark 上优于 SOTA 方法,在 ISBI2014 上取得 +2.2 AP 和 +2.7 AJI。QCell is presented, a novel query-based model that de-overlaps cell instances in microscopy scenes and outperforms state-of-the-art methods across multiple benchmarks, achieving +2.2 AP and +2.7 AJI on ISBI2014.

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training
DRACO:面向长视野 agent 训练的动态 rubric 细粒度信用分配
arXiv:2609.04094 Agent 智能体 方法 OA · 绿色 被引 1 · S2

该工作提出 DRACO:Distributing Rubric-based Advantage for Credit Optimization,在训练期间动态生成 rubric 以追踪 policy 的演化能力,对已完成的轨迹一次性评分,并将该判断重新分配到负责标注 rubric 的步骤上,以在 GRPO 中产生差异化的 per-step advantage。This work proposes DRACO: Distributing Rubric-based Advantage for Credit Optimization, which generates rubrics dynamically during training to track the policy's evolving capability, scores those rubrics once per completed trajectory, and redistributes that judgment over the steps responsible for annotated rubrics to produce differentiated per-step advantages in GRPO.

Last Translation Benchmark
终极翻译基准(Last Translation Benchmark)
arXiv:2609.04173 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 The Last Translation Benchmark,一组由人工编写并经同行评审的样本,可打破领先的机器翻译模型,并提出新评估方法:每个样本附带手工编写的验证规则,描述该样本上的具体失败案例,从而支持可靠且可操作的未来评估。The Last Translation Benchmark is introduced, a collection of human-authored and peer-reviewed examples that break leading machine translation models and a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and actionable future evaluation.

VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement
VeriPhy:用于世界模型评估与精进的 agent 物理推理
arXiv:2609.03153 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 VeriPhy,一种可审计的物理验证系统,其中纯文本 planner 在观察任何帧之前将 prompt 编译为类型化的物理义务与静态验证的执行计划。VeriPhy, an auditable physical-verification system in which a text-only planner compiles the prompt into typed physical obligations and a statically validated execution plan before any frame is observed, is presented.

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space
锁于入口,开于内里:RLVR 收窄解空间的位置
arXiv:2608.29188 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

表面 prompting 未能恢复多样性,而针对入口的干预则成功:使用早期 checkpoint 的 late-layer parameter interpolation 在不损失 pass@1 的情况下将解的覆盖度提高了 37%。While surface prompting fails to recover diversity, entrance-targeted interventions succeed: late-layer parameter interpolation with early checkpoints increases solution coverage by 37% at no loss in pass@1 and late-layer parameter interpolation with early checkpoints increases solution coverage by 37% at no loss in pass@1.

RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning
RoboTok:面向人类演示检索与灵巧操作学习的大规模互联网数据引擎
arXiv:2609.03199 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本工作提出 RoboTok,一个可扩展的数据引擎:利用人类操作视频作为查询,从互联网检索与操作相关的演示以训练灵巧机器人策略,并从以演员为中心的参考系下估计的 3D 手部轨迹中学习一个潜在运动空间This work introduces RoboTok, a scalable data engine that uses a query human manipulation video to retrieve manipulation-relevant internet demonstrations for training dexterous robot policies and learns a latent motion space from 3D hand trajectories expressed in estimated actor-centered reference frames.

Substrate-Aware AI Agents: Execution Context as a First-Class Input
Substrate-Aware AI Agent:将执行上下文作为一等输入。
arXiv:2609.05232 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

最小执行契约可在生成程序中引发主动的结构适应,将计算从无约束分配中移开,在执行前显著改善观察到的资源时间分布,建立了 substrate-aware Agent 规划的受控概念验证。A minimal execution contract induces proactive structural adaptation in generated programs, shifting computation away from unconstrained allocations and substantially improving observed resource-time profiles before execution, establishing a controlled proof of concept for substrate-aware agent planning.

Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
优化之前先提问:交互式优化中的动态预形式化澄清。
arXiv:2609.05258 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出了 OR-Clarify,一个用于预表述澄清的 benchmark,并提出了 Interactive Optimization (InterOPT),这是一个两阶段框架,能识别未解决的、影响表述的关键缺口,并据此引导系统决定是提出下一个问题还是停止提问。This work introduces OR-Clarify, a benchmark for pre-formulation clarification and proposes Interactive Optimization (InterOPT), a two-stage framework that identifies unresolved formulation-critical gaps and uses them to guide whether to ask the next question or to stop.

4️⃣ Tangram · 多轮对话非均匀KV Cache — arXiv:2606.06302(⭐⭐⭐⭐ 新鲜 arXiv)
arXiv:2606.06302 LLM 基础设施 方法 OA · 绿色 被引 2 · S2

Tangram 是一种 serving 框架,将先前系统动态处理的内容静态解析,可作为现有 non-uniform 压缩方法的即插即用底座,在匹配其精度的同时,端到端吞吐量较 full-KV 基线最高提升 2.6×。Tangram is a serving framework that statically resolves what prior systems handle dynamically, and serves as a drop-in substrate for existing non-uniform compression methods, matching their accuracy while improving end-to-end throughput by up to $2.6\times over the full-KV baseline.

RISE: Recursive Improvement via Self-Extrapolating Policy Distillation
RISE:通过自外推策略蒸馏实现递归改进
arXiv:2609.05295 工程化 方法 OA · 绿色 被引 2 · S2

涵盖数学推理、多领域 STEM、代码生成以及多轮 Agent 任务的实验表明,RISE 在所有设置下均优于仅使用 RLVR 的训练以及 on-policy self-distillation。Experiments spanning mathematical reasoning, multi-domain STEM, code generation, and multi-turn agentic tasks show that RISE outperforms RLVR-only training and on-policy self-distillation across all settings.

MaxKernel: Agentic Kernel Generation for TPUs
MaxKernel:面向 TPU 的 Agentic 核函数生成
arXiv:2609.04523 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出了 MaxKernel,一个多 Agent 系统,实现了 TPU kernel 开发的 three distinct paradigms:Human-in-the-Loop (HITL) Agent,用于协作式分步设计;Autonomous (Auto) Agent,执行全自动、由指标和 trace 驱动的优化循环;以及 Graph-Based Autonomous Search,将 Auto Agent 扩展以对设计空间进行全局探索。This work presents MaxKernel, a multi-agent system that implements three distinct paradigms for TPU kernel development: a Human-in-the-Loop (HITL) agent for collaborative, step-by-step design; an Autonomous (Auto) agent that executes a fully automated, metric/trace-driven optimization loop; and a Graph-Based Autonomous Search that scales the Auto agent for global exploration of the design space.

When Models Edit Too Much: On the Fidelity of Minimal Code Edits
模型过度编辑时:最小代码编辑的保真度问题
arXiv:2609.04061 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作通过向参考解中注入受控的 AST 级 corruption,从 400 个 BigCodeBench 问题构建了一个评估框架,为每个修复任务赋予已知的最小 patch,并将 edit fidelity 定位为 code-repair 质量的一个独立维度,表明其可被度量与学习。This work constructs an evaluation framework from 400 BigCodeBench problems by injecting controlled AST-level corruptions into reference solutions, giving each repair task a known minimal patch, and positions edit fidelity as a distinct axis of code-repair quality and shows that it can be measured and learned.

When Quantization Breaks Memory: Recurrent-State Write-Back in Low-Precision Temporal Inference
当量化破坏记忆时:低精度时序推理中的循环状态写回
arXiv:2609.04490 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

循环状态写回被确立为低精度循环动态的关键决定因素,state-storage interface 被识别为量化循环推理的核心设计考量。R recurrent-state write-back is established as a key determinant of low-precision recurrent dynamics and the state-storage interface is identified as a central design consideration for quantized recurrent inference.

Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue
Motion-Omni:面向口语对话的端到端联合语音与全身动作生成
arXiv:2609.04250 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 Motion-Omni,一个端到端框架,其中的 spoken dialogue model 原生输出显式的 facial expression 以及手部、上半身和下半身运动,这些输出直接由生成语音的 hidden states 生成。Motion-Omni is presented, an end-to-end framework in which a spoken dialogue model natively outputs explicit facial expression together with hand, upper-body and lower-body motion, generated directly from the hidden states that produce the speech.

Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs
思维链表象之下:LLM 中推理操作的机制化解读
arXiv:2609.04753 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作发现 reasoning operations 在 held-out 表征中是可分的,且 separability 在中间层达到峰值,并验证该结构无法由词汇或位置混淆因素解释。This work finds that reasoning operations are separable in held-out representations, with separability peaking in middle layers, and verify that this structure is not explained by lexical or positional confounds.

τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction
τ^τ-Bench:面向端到端真实 Agent 构建的环境
arXiv:2609.04611 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 τ^τ-bench(读作 hyper-tau-bench),一个将 Agent 构建作为任务的 benchmark,将协作式 Agent 构建工作转化为面向 coding agent 的可度量目标。The $\tau^\tau$-bench (pronounced hyper-tau-bench), a benchmark that makes agent construction the task, is introduced to turn the work of cooperative agent building into a measurable target for coding agents.

The Attention Triangle in Audio-Video Models
音视频模型中的注意力三角
arXiv:2609.03586 多模态 方法 OA · 绿色 被引 1 · S2

研究揭示沿音视频边的路由是双向的:音频可影响视频生成,视频也可影响音频生成;模型参数中编码的偏差是泄漏的主要来源之一。It is revealed that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation, and biases encoded in the model's parameters and emerges as a major contributor to leakage.

Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems
双层协同反思:多 Agent LLM 系统的博弈论方法
arXiv:2609.02750 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出随机反思记忆提升(SRMA),仅在 grounding 后的评估风险严格下降时才接受候选记忆,为随机评估提供置信度门控,并为分段平稳环境提供重新锚定保证。Stochastic Reflective Memory Ascent (SRMA), which accepts a candidate memory only after a grounded evaluation risk strictly decreases, is introduced and provides confidence gating for stochastic evaluation and re-anchoring guarantees for piecewise-stationary environments.

ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding
ShallowStream:先浅层索引再深层回答的流式视频理解
arXiv:2609.02780 多模态 应用落地 OA · 绿色 被引 2 · S2

ShallowStream 是一个利用 MLLM 浅层同时进行帧编码与检索索引构建的新框架,性能与当前最强流式方法相当,同时将单帧 prefill 延迟与 10 秒端到端延迟分别降低至多 52.1 倍和 11.9 倍。ShallowStream is a novel framework that leverages the shallow layers of an MLLM to simultaneously perform frame encoding and retrieval index building and achieves performance on par with the strongest existing streaming methods, while reducing per-frame prefill latency and 10-second end-to-end latency by up to 52.1x and 11.9x, respectively.

4. The End of Software Engineering(arXiv:2606.05608)
4. 软件工程的终结(arXiv:2606.05608)
arXiv:2606.05608 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文认为,AI agent——即以大语言模型作为主要推理引擎、动态生成与丢弃代码作为工具性资源的系统——的出现构成了对"软件"本身的根本性重构,而非渐进式的工具改进。This paper argues that the emergence of AI agents -- systems where large language models serve as the primary reasoning engine, dynamically generating and discarding code as an instrumental resource -- constitutes a fundamental restructuring of what software is, not an incremental tool improvement.

Enoki: Efficient Multi-Level Hallucination Detection
Enoki:高效多层级幻觉检测
arXiv:2609.00581 安全与风险 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Enoki,一个面向多级幻觉检测的开放信息抽取框架,支持基于 LLM、基于编码器与基于规则的三类抽取模式,通过统一接口平衡准确率与推理成本。This work proposes Enoki, an Open Information Extraction framework for multi-level hallucination detection that supports LLM-based, encoder-based, and rule-based extraction regimes, balancing accuracy and inference cost through a common interface.

To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation
于鲜活情境中观世界:统一的室内-室外城市世界生成
arXiv:2608.05879 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

HoloWorld 是首个在统一连贯的 3D 城市世界中同时支持室内与室外生成的框架,构建于持续更新的跨尺度世界上下文之上。HoloWorld is the first framework to unify indoor and outdoor generation within a coherent 3D urban world, built on a continuously updated cross-scale world context.

UniMate: One Unified Model to Animate Diverse Skeletons
UniMate:一个统一模型驱动多样化骨骼动画
arXiv:2609.05415 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 UniMate,一个统一的 foundation model,可从绑定骨骼的 3D 资产与文本提示合成任意骨骼的关节运动,无需测试时优化或针对每个骨骼的重新训练,在质量、泛化性与效率上均超越 SOTA 基线。UniMate is presented, a unified foundation model that synthesizes articulated motion for arbitrary skeletons from a rigged 3D asset and a text prompt, with no test-time optimization or per-skeleton retraining, and outperforms state-of-the-art baselines in quality, generalization, and efficiency.

Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
知道不该回答什么:视觉语言模型中的选择性不遵从
arXiv:2609.04720 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 KoNA,一个评估 VLM 选择性不遵从的基准,覆盖五类情形:False Premise、Visual Inaccessibility、Universal Unknown、Task Feasibility 与 Safety。实验结果显示,微调后的模型能够区分可回答部分与需要不遵从的部分,并以符合任务要求的方式作答。KoNA is introduced, a benchmark for evaluating selective non-compliance in VLMs across five categories: False Premise, Visual Inaccessibility, Universal Unknown, Task Feasibility, and Safety, and results suggest that the fine-tuned models can distinguish between answerable components and those requiring non-compliance and respond in a task-appropriate manner.

Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models
Refuse without Refusal:面向语言模型安全调优响应的结构性分析以减少误拒
arXiv:2609.04714 安全与风险 方法 OA · 绿色 被引 2 · S2

本文将安全微调数据集中的回复拆分为两个独立部分:模板化的拒答声明与解释拒答的理由,并表明拒答声明会诱导模型依赖表层线索,从而妨碍对有害与良性查询的准确区分。This paper decomposes a response in the safety-tuning dataset into two distinct components: a boilerplate refusal statement and a rationale explaining the refusal, and shows that refusal statements impede accurate discrimination between harmful and benign queries by inducing reliance on superficial cues.

One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing
一个编辑器,多种编辑:一个用于多样化视频编辑的统一免训练框架
arXiv:2609.04190 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

视频编辑涵盖多种编辑范式,但在单一统一框架中同时实现高质量的指令引导与主体引导编辑仍具挑战性。我们提出 EditVid,一个免训练框架,结合用于局部一致性的稀疏因果记忆、用于长程身份保持的基于对应关系的后注意力 token 注入,以及用于编辑局部性的软潜在融合。同一框架支持指令引导和参考引导的编辑,包括风格迁移、属性修改、对象插入、部分级编辑和主体替换。在Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The same framework supports instruction-guided and reference-guided edits, including style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On

HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
HarvestBench:衡量 LLM Agent 是否愿意付出代价以避免杀死动物
arXiv:2609.04444 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

HarvestBench 是首个为避免 side effect 标定价格并将 side effect 命名为生物的 benchmark;在六个模型中有四个的每次作答遭遇 kill rate 对价格变化敏感。HarvestBench is the first benchmark to put a price on avoiding a side effect and name the side effect as a living creature and four out of six models'kill rate per answered encounter were sensitive to price changes.