研究库 论文知识库
Papers · organized/paper_cards

论文

1173 张论文卡片 · 方法

开放获取 全部 绿色 · 1640
Dr. Claw: An AI Scientist Workspace for Vibe Research
Dr. Claw:一个面向氛围研究的 AI 科学家工作台
arXiv:2609.00365 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出 Dr. Claw,一个开源工作空间,将现有编码 Agent 可执行文件封装在可控且可审计的人机协同工作流中,而非引入另一个自主 Agent。Dr. Claw is presented, an open-source workspace that wraps existing coding-agent executors in a controllable and auditable human-in-the-loop workflow rather than introducing another autonomous agent.

3️⃣ arXiv · Memanto: Typed Semantic Memory with Information-Theoretic Retrieval for Long-Horizon Agents(⭐⭐⭐⭐ 高优先级)
3️⃣ arXiv · Memanto:面向长程 Agent 的带类型语义记忆与信息论检索(⭐⭐⭐⭐ 高优先级)
arXiv:2604.22085 Agent 智能体 方法 OA · 绿色 被引 6 · S2

本文提出 Memanto,一种面向 agentic 人工智能的通用记忆层,挑战了"必须依赖知识图谱复杂度才能实现高保真 agent 记忆"的普遍假设,并取得 SOTA 准确率。Memanto is introduced, a universal memory layer for agentic artificial intelligence that challenges the prevailing assumption that knowledge graph complexity is necessary to achieve high fidelity agent memory and achieves state of the art accuracy scores.

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
FlowBalance:基于验证器的 On-Policy 推理经验自改进
arXiv:2609.03241 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

FlowBalance 是一种以验证器为锚定的自改进方法,学习完整响应上的归一化分布,在 Qwen3-4B 和 Qwen3-8B 上相对 FlowRL 提升了平均性能,同时改善了训练速度与稳定性,避免了直接 OPSD 响应长度坍缩,并在受控的 AIME24 诊断中表现出更高的正确策略多样性。FlowBalance, a verifier-grounded self-improvement method that learns a normalized distribution over complete responses, improves average performance over FlowRL on both Qwen3-4B and Qwen3-8B, while also improving training speed and stability, avoiding direct OPSD's response-length collapse, and exhibiting higher correct-strategy diversity in a controlled AIME24 diagnostic.

What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation
还有什么需要修复?探索对话生成制品中修订传播的成本有效 test-time compute
arXiv:2609.03254 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文为该设定引入新基准,并基于该基准评估了九种修订方法,包括序贯反思与并行采样变体,使用 gpt-oss-20b/120b、gpt-5.4-mini 以及 qwen3.5-9b/27b/122b 进行测试。A new benchmark for this setting is introduced, and nine revision methods are evaluated, including sequential reflection and parallel sampling variants, using gpt-oss-20b/120b, gpt-5.4-mini, and qwen3.5-9b/27b/122b on the benchmark.

Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation
蒸馏前先验证:用于 On-Policy 蒸馏的 Prompt 级教师门控
arXiv:2609.02998 多模态 方法 OA · 绿色 被引 4 · S2

本文提出 Teacher-Gated On-Policy Distillation(教师门控的在线策略蒸馏),其核心原则是在引入密集监督前以 prompt 级别验证教师可靠性;在全部六个单领域设定下优于 Vanilla OPD,并在多领域训练下于两种规模上取得更高的七项基准平均成绩。Teacher-Gated On-Policy Distillation is introduced, built on the principle that teacher reliability should be verified at the prompt level before dense supervision is admitted, and outperforms Vanilla OPD in all six single-domain settings and achieves higher seven-benchmark averages at both scales under multi-domain training.

Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing
面向遥感变化检测的基于真实世界知识引导的变化数据合成
arXiv:2608.24263 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 KnowChange,一个知识引导的变化数据合成框架,利用预训练视觉语言模型作为知识源,从变化前场景与期望变化类型推理合理的变化位置与类别转移,在统一框架下灵活合成多样变化类型。KnowChange is introduced, a knowledge-guided change data synthesis framework that leverages pretrained vision-language models as knowledge sources to reason about plausible change locations and class transitions from pre-change scenes and desired change types and enables flexible synthesis of diverse change types within a unified framework.

Unlocking Lossless Speedups in LLMs via Discrete Diffusion
通过离散扩散释放 LLM 的无损加速
arXiv:2609.04010 多模态 方法 OA · 绿色 被引 2 · S2

大语言模型(LLM)的成功很大程度上归功于 next-token prediction(NTP),但其自回归(AR)结构需要缓慢的串行 token 生成。为克服这一瓶颈,我们提出扩散增强 LLM,一类新模型,在使用扩散从该分布中并行采样多个 token 的同时定义 AR 模型分布。我们将这些模型的参数解耦为两组:AR 权重,使用标准 NTP 目标训练;轻量扩散权重,训练用于同时生成多个 token。扩散权重Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overcome this bottleneck, we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion to draw multiple tokens in parallel from that distribution. We decouple the parameters of these models into two sets: AR weights, trained using the standard NTP objective, and lightweight diffusion weights, trained to generate multiple tokens simultaneously. The diffusion weight

ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation
ENEAS:嵌入引导的自适应分割神经集成
arXiv:2609.03756 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

我们提出 ENEAS,一种用于实例追踪与语义发现的统一且文本可提示的方法。包括 SAM 3 在内的文本可提示分割模型仍存在时间幻觉、空间碎片化与语义误分类问题:目标离开视野时无法报告目标缺失;极端特写下只分割局部纹理而非完整目标;将视觉特征置于本体事实之上,从而把雕像、绘画或反射等视觉相似的物体误分割为目标。We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absence when an object leaves the field of view, segment local textures instead of the complete object during extreme close-ups, and prioritize visual features over ontological reality, so that visually similar artifacts such as statues, paintings, or reflections are segmented as targe

Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys
Split-LLM 训练中的隐私失败:返回的梯度使诱饵失效
arXiv:2609.04382 工程化 方法 OA · 绿色 被引 1 · S2

本文对一个双节点 split-LLM 训练系统进行系统安全案例研究:其隐私评估通过,却遗留一条未被测试的可观测信道;系统因此并不安全——包括跨训练步骤累积观测在内的五类攻击从未被测量。A systems-security case study of a two-node split-LLM training system whose privacy evaluation passed while leaving an observable channel untested, but the system is not thereby safe: five classes of attack, including those accumulating observations across training steps, were never measured.

3️⃣ arXiv · MatryoshkaLoRA(⭐⭐⭐⭐ 值得关注)
3️⃣ arXiv · MatryoshkaLoRA(⭐⭐⭐⭐ 值得关注)
arXiv:2605.07850 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MatryoshkaLoRA,一种受 Matryoshka 启发、面向 LoRA 的通用训练框架,通过在已有 LoRA adapter 之间插入一个固定的、经精心设计的对角矩阵来按比例缩放其子秩,从而学习到准确的层次化低秩表示。MatryoshkaLoRA is proposed, a general, Matryoshka-inspired training framework for LoRA that learns accurate hierarchical low-rank representations by inserting a fixed, carefully crafted diagonal matrix between the existing LoRA adapters to scale their sub-ranks accordingly.

Causal Foundation Models
因果基础模型
arXiv:2609.03003 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

因果基础模型是经过预训练的神经网络,可在全新的数据集上通过上下文学习估计因果量(如平均处理效应),无需模型更新。Causal foundation models are pretrained neural networks that estimate causal quantities, such as the average treatment effect, on entirely new datasets using in-context learning without requiring model updates.

Unifying Conformal Language Tasks with In-Context Ensembles
通过上下文集成统一保形语言任务
arXiv:2609.03005 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Conformal Relevance 框架,利用上下文学习的示例筛选与集成构造打分函数,在保持覆盖的同时以极低人工成本提升简洁性。The Conformal Relevance framework is introduced which uses in-context learning example curation and ensembling to create a score function which maintains coverage while improving conciseness with minimal manual input.

Omni Interaction Agent Technical Report
Omni Interaction Agent 技术报告
arXiv:2609.08977 多模态 方法 OA · 绿色 被引 2 · S2

提出了 Gander,一个原生多模态双工交互模型,基于 MiniCPM-o 4.5 构建,并通过异步 Agent 循环进一步适配实时交互,同时开源其模型、代码和数据,以推动社区进一步研究与开发。Gander is presented, a native multimodal duplex interaction model that builds on MiniCPM-o 4.5 and is further adapted for realtime interaction with an asynchronous agent loop and is released together with its models, code, and data to facilitate further research and development in the community.

Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout
Mask Forcing:通过双噪声掩码 rollout 改进自回归视频扩散蒸馏
arXiv:2609.09123 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Mask Forcing 是一种双噪声掩码展开策略,通过扰动自回归学生的自展开来缓解由 reverse-KL 模式寻求引发的模式坍缩,能以更高视觉质量高效改进多种自回归视频扩散蒸馏方法,且无需引入真实视频数据或额外后训练阶段。Mask Forcing, a Dual-Noise Masking Rollout strategy that perturbs the AR student self-rollout to mitigate mode collapse induced by reverse-KL mode seeking, improves multiple AR video diffusion distillation methods with higher visual quality efficiently, without incorporating real video data or additional post-training stages.

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
AuK 技术报告:用于语音生成与编辑的开源基础模型
arXiv:2609.08936 多模态 方法 OA · 绿色 被引 2 · S2

本文提出 AuK,一个开源基础模型,通过自然语言指令与音频上下文的统一接口整合语音生成与编辑,在零样本与指令控制的语音生成以及通用指令引导编辑上取得领先性能。AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context, is introduced and leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing is demonstrated.

CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
CoVeR:基于覆盖度的 token 剪枝,用于 VLM 中的多视图 3D 推理
arXiv:2609.08345 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文证明空间覆盖与三维推理性能相关,并提出 CoVeR,一种仅使用 token 坐标、不依赖学习信号的确定性、无训练选择器,在三项三维推理基准上均超越既有 SOTA,并可作为即插即用模块泛化到四种 VLM。It is shown that spatial coverage is associated with 3D reasoning performance and CoVeR, a deterministic, training-free selector that uses only token coordinates, with no learned signals is introduced, which outperforms prior SOTAs on all three 3D reasoning benchmarks and generalizes as a plug-and-play module tested across four VLMs.

TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
TANGO:基于全身 Vision-Language-Action 模型的杂乱环境中人形机器人导航
arXiv:2609.09158 多模态 方法 OA · 绿色 被引 2 · S2

本文提出 TANGO,首个面向语言条件人形机器人在杂乱环境中通行的全身视觉语言导航框架,在视觉语言导航任务中达到 SOTA,并在需要避障的困难场景中超越强模块化基线。This work introduces TANGO, the first whole-body vision-language navigation framework for language-conditioned humanoid traversal in cluttered environments, and demonstrates state-of-the-art performance in vision-language navigation, while outperforming strong modular baselines in navigating challenging scenes requiring obstacle negotiation.

TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation
TransNormal-2:基于几何约束的 Rectified Flow 与边缘感知解码用于精确法向量估计
arXiv:2609.06665 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

TransNormal-2 是基于 FLUX.2 的整流流框架,采用单步确定性推理,针对 VAE 解码器两侧的重建退化问题,施加 RGB 引导的残差修正以降低局部于边界的解码误差,且不自由改写粗预测。TransNormal-2, a FLUX.2-based rectified-flow framework with single-step deterministic inference that addresses VAE reconstruction degradation on both sides of the VAE decoder, and applies an RGB-guided residual correction to reduce boundary-localized decoding errors without freely rewriting the coarse prediction.

BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference
BeaconKV:由 Beacon 查询引导的 KV Cache 压缩,用于高效大型推理模型推理
arXiv:2609.04971 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 BeaconKV,一种免训练的 KV cache 压缩方法,通过为每个全局查询簇维护紧凑代表性 beacon query 来预测哪些 KV 对将被重访,无需存储完整查询历史。BeaconKV is proposed, a training-free KV cache compression method that maintains beacon queries, compact representatives for each global query cluster, to anticipate which KV pairs will be revisited without storing the entire query history.

A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM
A*-Thought-V2:通过 LLM 几何动力学实现高效潜在推理
arXiv:2609.07821 安全与风险 方法 OA · 绿色 被引 1 · S2

提出 A*-Thought-V2,一个由 LLM 引导的几何动力学框架,将 CoT 建模为隐状态轨迹,并以显式-隐式交错潜在架构替代硬删除,引入更广义的软目标以促进更丰富的步骤级特征学习。A*-Thought-V2 is presented, a geometric dynamics of LLM guided framework that models CoT as a hidden-state trajectory and replaces hard deletion with an explicit-implicit interleaved latent architecture that reflects broader soft targets that encourage richer step-level feature learning.

SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation
SynthGait-19K:用于步态参数估计的物理约束合成视频数据集
arXiv:2609.08108 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

发现空间步态参数对视觉域偏移更敏感,且更好的 HMR 重建并不一定带来下游步态估计的改进;提出 GaitXFormer,作为直接基于 RGB 的参考模型用于步态参数估计。It is found that spatial gait parameters are more sensitive to visual domain shift and that improved HMR reconstruction alone does not necessarily translate to improved downstream gait estimation, and GaitXFormer is introduced as a direct RGB reference model for estimating gait parameters.

OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining
OpenWAM:面向系统化世界-动作模型预训练的开源模块化探索
arXiv:2609.07398 工程化 方法 OA · 绿色 被引 14 · S2

提出 OpenWAM,一个开源研究栈,将世界-动作预训练转化为可控的实验项目,并发布完整栈,包括基础设施、评估协议、预训练模型和数据配方,以促进未来研究。This work introduces OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program, and releases the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.

Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise
Encoded Early, Used Late:Transformer 在何处开始作用于所推断的对话伙伴专业度
arXiv:2609.07139 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

发现合作者专业能力在模型早期层最易被解码,到网络中点前降至接近随机水平,并基于合成语料以单个模型作为初步验证。It is found that partner expertise is most decodable in the early layers and falls to near chance before the midpoint of the network, and one model on a synthetic corpus is used as an initial demonstration.

Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection
借助 CLIP 与 DINO:面向可泛化深度伪造图像检测的不确定性感知级联融合网络
arXiv:2609.07670 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 UCF-Net,一个不确定性感知级联融合网络,利用 CLIP 语言对齐的语义先验与 DINO 自监督视觉结构先验,在域内与跨域评估中均取得最优平均 AUC。UCF-Net is proposed, an uncertainty-aware cascaded fusion network that harnesses CLIP's language-aligned semantic priors and DINO's self-supervised visual-structure priors and achieves the best mean AUC among the evaluated methods in both in-domain and cross-domain evaluations.

RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting
RelightFormer:用于多视角物体重光照的前馈生成式 Transformer
arXiv:2609.07414 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出一个前馈式生成 Transformer,用于直接的单视图与多视图图像重光照,完全绕过显式本征属性估计,并采用置换不变的位置编码对称处理无序多视图输入,避免序列偏差。This work introduces a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation and employs permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias.

Cadence: Error-Bounded Lossy Compression of Demand Time Series with a Time-Series Foundation Model
Cadence:基于时序基础模型的需求时序误差有界有损压缩
arXiv:2609.06008 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

我们提出 Cadence,一种面向数值时序的误差有界有损压缩器,将 3.3 亿参数的时序基础模型(Google TimesFM-3)与自适应算术编码器相结合,保证每个样本满足 |x_t - x̂_t| ≤ τ。一项负面结论限定了设计空间:在无损编码场景下,基础模型毫无价值,因为节省的比特数仅与预测器精度呈对数关系 Δb = log_2(MAE_old / MAE_new)。因此 TimesFM-3 相对 32 阶线性预测器 1.51 倍的精度优势,在 20.28 比特中仅换取 0.60 比特,中位数增益仅 +0.03%。误差有界编码仅在一点上突破了这一限制。We present Cadence, an error-bounded lossy compressor for numeric time series pairing a 330M-parameter time-series foundation model (Google TimesFM-3) with an adaptive arithmetic coder, guaranteeing |x_t-x_t|leτ on every sample. One negative result constrains the design space: for lossless coding a foundation model is worth nothing, because bits saved are logarithmic in predictor accuracy, Δb=log_2(MAE_{old}/MAE_{new}). So the 1.51times advantage TimesFM-3 holds over a 32-tap linear predictor buys 0.60 bits of 20.28, a median gain of +0.03%. Error-bounded coding escapes this at one point: once

NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting
NOAH:学习完整患者旅程的纵向多模态时序感知表示与预测模型
arXiv:2609.09140 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Noah 是一个时间感知、任务无关的生成式 Transformer 模型,对完整多模态患者旅程进行表征与预测,是该领域首个真正整体化的生成模型,支持自回归预测,并具备可选的时间控制、零样本分类与反事实干预模拟能力。Noah is a time-aware, task-agnostic, generative transformer model representing and forecasting the full multimodal patient journey, and is the first truly holistic generative model in its field, enabling autoregressive forecasting with optional time control, zero-shot classification, and counterfactual intervention simulation.

3️⃣ Speculative Decoding 延迟可解释模型 — arXiv:2605.15051(⭐⭐⭐⭐ 调优必读)
arXiv:2605.15051 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文为 LLM serving 中的 SD(投机解码)提出了一种简单且可解释的 latency 模型,能够准确刻画实际观测到的 latency,解释为何加速比常随服务器负载上升而下降,并系统刻画了 draft length、acceptance rate 以及 verifier 与 drafter 规模在不同 serving 条件下对 latency 的影响。A simple and interpretable latency model for SD in LLM serving is developed that accurately describes observed latency, explains why speedups often diminish as server load increases, and characterizes how draft length, acceptance rate, and verifier-drafter size shape latency across serving conditions are characterized.

Counter-Swarm Doctrine: Containing Coordinated Agent Intrusions
反集群 doctrine:遏制协同化 Agent 入侵
arXiv:2609.06140 Agent 智能体 方法 OA · 绿色 被引 1 · S2

主张防御的运行单元应是可修订的协同 episodes,将观察到的迁移、任务权限与响应历史关联起来,并提出跨执行监控建议具有可测试性,但并不声称提出新的检测器或测得具体的遏制收益。It is argued that the operational unit of defence should be a revisable coordination episode linking observed transfers, task authority, and response history, and it makes the recommendation to monitor across executions testable without claiming a new detector or a measured containment benefit.

RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives
RenderFormer-V2:基于异构场景基元的神经渲染
arXiv:2609.05738 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

我们提出 RenderFormer-V2,一个统一的基于 transformer 的学习型神经渲染模型,可与现代基于物理的渲染系统互补,无需逐场景训练或专用代码,即可处理焦散、体积散射、环境光照、带纹理与置换的表面以及分布外材质等多种光传输效果。RenderFormer-V2 将全局光传输建模为序列到序列变换。沿袭前作,它仍采用两阶段流程:先是与视图无关的阶段,解析场景内基元到……We present 'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to modern physics-based rendering systems, that can handle diverse light-transport effects such as caustics, volumetric scattering, environment lighting, textured and displaced surfaces and out-of-distribution materials without per-scene training or specialized code. RenderFormer-V2 models global light transport as a sequence-to-sequence transformation. Following its predecessor, RenderFormer-V2 also employs a two stage process: a view-independent stage that resolves intra-scene primitive to

Graph Machine: Towards Better Pretraining via Edges
Graph Machine:通过边实现更优的预训练
arXiv:2609.02881 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Graph Machine,一种保持 O(n) 规模状态并通过稀疏动态路由访问的架构,使用边——由类似指针追逐的引用机制以可微分方式更新的指针类对象。The Graph Machine is introduced, an architecture that maintains an O(n)-sized state and accesses it through sparse, dynamic routing and uses edges - pointer-like objects updated differentiably by a referral mechanism resembling pointer chasing.

MasterControl Seventeen Every Time
MasterControl:每次都 Seventeen
arXiv:2609.03209 数据与向量库 方法 OA · 绿色 被引 0 · S2 + OpenAlex

我们研究一种受治理的企业分析方法:语言模型负责解读问题,确定性 policy 负责选取并运行预先批准的分析程序,返回结果与证据。我们证明,在限定的分析类内(包含关系运算,以及聚合、比较、窗口、排序和相似度),这种限制仍可保持表达力。固定的语义、policy、数据和执行规则也使结果可复现。在 440 次运行中,三个 8B 模型生成 SQL 并在运行时选取工具,而 Qwen3-8B 仅解读意图,policyWe study a governed approach to enterprise analytics: a language model interprets the question, while deterministic policy selects and runs a pre-approved analytical program that returns both results and evidence. We show that this restriction can remain expressive within a defined analytical class, using relational operations plus aggregation, comparison, windows, ranking, and similarity. Fixed meaning, policy, data, and execution rules also make results replayable. Across 440 runs, three 8B models generated SQL and selected tools at runtime, while Qwen3-8B interpreted intent only and policy

Glyph: A Multi-Strategy Agentic System for Column Description and Sensitivity-Ontology Tagging of Enterprise Data Catalogs
Glyph:面向企业数据目录的列描述与敏感度本体标注的多策略 Agent 系统
arXiv:2609.10430 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Glyph 是一个生产系统,将列描述生成与列类型标注这两个耦合问题建模为协同工作的 LLM Agent,并以有状态图形式编排,使多 Agent LLM 目录编制可审计且可作为生产服务运行。Glyph, a production system that frames two coupled problems, column description generation and column type annotation for data classification, as cooperating LLM agents orchestrated as stateful graphs, makes multi-agent LLM cataloging auditable and operable as a production service.

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
SAEScientist-Bench:AI Agent 能否开展自主 SAE 可解释性研究?
arXiv:2609.09113 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

评估 AI Agent 能否作为科学家利用 SAE 工具开展自主机理发现,旨在将实验性模型理解确立为可测量的能力,推动闭环自主 AI R&D。Whether AI agents can act as scientists utilizing SAE tools for autonomous mechanistic discovery is evaluated to establish experimental model understanding as a measurable capability for closed-loop autonomous AI R&D.

Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
协同进化的 Harness 与模型:On-policy 修正助力弱模型在模仿失效处迎头赶上
arXiv:2609.09134 评测基准 方法 OA · 绿色 被引 3 · S2

开发一条 on-policy 专家修正流水线,由元层级 MLE Agent 自动化,在弱模型自身的 rollout 中定位失败回合,并请专家仅重写该回合,从而保留模型的规划风格,融合 Harness 进化与模型适配带来的收益。An on-policy expert-correction pipeline is developed, automated by a meta-level MLE agent, that localizes the failing turn in the weaker model's own rollout and asks the expert to rewrite only that turn, which preserves the model's planning style and combines the gains of harness evolution and model adaptation.

3. Tutti:让 SSD 后备 KV Cache 成为长上下文生产方案
arXiv:2605.03375 LLM 基础设施 方法 OA · 绿色 被引 6 · S2

Tutti 是一种高效的 SSD-backed KV caching 方案,将 CPU 从 HBM 与 SSD 之间的关键数据与 I/O 控制路径中彻底移除,在提供近乎无限容量的同时,实现了与 DRAM-backed LMCache 几乎相当的 inference 性能。Tutti is an efficient SSD-backed KV caching solution that eliminates CPU intervention from the critical data and I/O control paths between HBM and SSDs, and achieves nearly the same inference performance as DRAM-backed LMCache, while providing almost infinite capacity.