研究库 论文知识库
Papers · organized/paper_cards

论文

1640 张论文卡片 · OA 绿色

开放获取 全部 绿色 · 1640
OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
OSWorld-Science:面向科学与科研软件学习与使用的计算机操作 Agent 基准
arXiv:2609.39903 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 OSWorld-Science,这是一个结合科学意义任务、基于 artifact 的评估以及高效 agent harness 的基准与评估环境,用于研究科学领域的计算机使用,从而在科学工作流中系统评估 agent 能力与 harness 设计。OSWorld-Science is introduced, a benchmark and evaluation environment that combines scientifically meaningful tasks, artifact-based evaluation, and an efficient agent harness for studying computer use in the scientific domain, enabling systematic evaluation of both agent capabilities and harness design in scientific workflows.

Scaling Laws for Looped Mixture of Experts
循环化 Mixture of Experts 的扩展定律
arXiv:2609.40316 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Loop Scaling Laws,是首个将 recurrence 和 sparsity 与模型规模、数据联合建模的 scaling law,为在算力和显存约束下设计 looped MoE 模型提供了原则性基础。Loop Scaling Laws are introduced, the first scaling law to jointly model recurrence and sparsity alongside model size and data, and provide a principled foundation for designing looped MoE models under compute and memory constraints.

Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation
打破巴别塔:面向长字幕翻译的自进化多 Agent 系统
arXiv:2609.38660 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

SMART,面向长篇字幕翻译的自演化多 Agent 系统,构建持久化的剧集级 memory,并通过动态 router 和 Mixture-of-Agents 层翻译部分句子,配合术语验证、字幕约束校验和上下文检索工具。SMART, a Self-evolving Multi-Agent system for long-foRm subtitle Translation, builds persistent series-level memory and translates a subset of sentences through a dynamic router and Mixture-of-Agents layer with tools for terminology verification, subtitle constraint validation, and contextual retrieval.

EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making
EVOKE:在 Agent 中引出世界知识以实现可迁移的决策
arXiv:2609.38334 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 EVOKE,一种后训练方法,通过在固定状态下以目标多样性对直接决策施加监督,来施加压力以激发模型内化的、可迁移动作的世界知识,并提供了一种通过直接决策监督激发内化世界知识以获得可迁移动作的新视角。EVOKE is introduced, a post-training method that supplies pressure on eliciting internalized world knowledge for transferable action through direct decision supervision through goal diversity at fixed states, and offers a new perspective on eliciting internalized world knowledge for transferable action through direct decision supervision.

Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models
揭示状态之谜:Masked Diffusion Language Models 中状态适应何时重要
arXiv:2609.33355 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

转移层面的结果表明,state adaptation 选择性而非统一地应用时最为有效,且 adaptation 的价值取决于策略反转的频率与幅度。The transition-level results suggest that state adaptation is most useful when applied selectively rather than uniformly, and that adaptation value depends on both the frequency and magnitude of strategy reversals.

RAGScope: A Leakage-Controlled, Cost-Aware Evidence-Gating Protocol for RAG Hallucination Triage
RAGScope:面向 RAG 幻觉分诊的泄漏可控、成本感知的证据门控协议
arXiv:2609.39075 RAG 检索增强 应用落地 OA · 绿色 被引 1 · S2

RAGScope 是一个泄漏受控的评估协议,用于评估仅使用任务输入、检索上下文和答案文本的本地证据门,结合了上下文分组划分、折范围预处理、组自举区间、部署工作点、端到端运行时以及显式的源偏移压力测试。RAGScope, a leakage-controlled protocol for evaluating local evidence gates that use only the task input, retrieved context, and answer text is presented, which combines context-grouped splits, fold-scoped preprocessing, group bootstrap intervals, deployment operating points, end-to-end runtime, and explicit source-shift stress tests.

RoPE at the End of Its Rope? Theory, Diagnosis, and Mitigation of Long-Context Failures
RoPE 走到尽头了吗?长上下文失效的理论、诊断与缓解
arXiv:2609.39929 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

基于 RoPE 的语言模型出现长上下文失效,根源在于 RoPE 在维持稳定 token 偏好与区分相近位置之间存在固有权衡。要判断应处理哪种弱点以及如何处理,需要更精确地刻画 RoPE 在不同上下文长度下训练后模型中的行为。我们通过允许 RoPE 各频率下 query-key 尺度不一致,弥补了先前理论的一个关键局限,使之与实际经验观测高度吻合。我们的理论使得上述脆弱性对单个注意力头与输入可测量,并量化了高频Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and distinguishing nearby positions. Determining which weakness to address, and how, requires a more precise characterization of RoPE's behavior in trained models across context lengths. We address a key limitation of prior theory by allowing unequal query-key scales across RoPE frequencies, which aligns well with practical empirical observations. Our theory makes both vulnerabilities measurable for individual heads and inputs, and quantifies how high-freque

4.1 LogicalRAG:把 Agentic RAG 的重点从“更重 backend”转向“更强 retrieval control”
arXiv:2605.27123 RAG 检索增强 方法 OA · 绿色 被引 3 · S2

本文提出一个 Agentic RAG 框架,使 LLM 能够使用逻辑表达式构建检索意图,同时将检索后端简化为基于倒排索引的系统,并表明将检索过程锚定在逻辑查询上可显著降低生成响应中的幻觉。This paper proposes an agentic RAG framework that enables LLMs to formulate retrieval intents using logical expressions while simplifying the retrieval backend to an inverted-index-based system, and shows that anchoring the retrieval process in logical queries substantially reduces hallucinations in generated responses.

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
DyRAD:面向动态驾驶场景的雷达新视角合成
arXiv:2609.39841 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

DyRAD 使用静态背景反射器和运动追踪的动态点反射器建模动态驾驶场景,渲染完整的距离-方位-多普勒 (RAD) 张量,并通过从雷达信号处理链推导出的固定解析点扩散函数渲染反射器,避免传感器引起的扩散被烘焙到场景表示中。DyRAD is presented, which models dynamic driving scenes using static background reflectors and motion-tracked dynamic point reflectors to render complete range-azimuth-Doppler (RAD) tensors, and renders reflectors through a fixed analytic point-spread function derived from the radar's signal-processing chain, preventing sensor-induced spread from being baked into the scene representation.

Safety of Latent Communication in Multi-Agent Systems
多 Agent 系统中潜在通信的安全
arXiv:2609.39788 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

即使是良性的链接训练也会在底层安全对齐 agent 不变的情况下,相对于基于文本的通信增加有害合规性;安全对齐需要将多智能体系统作为整体来考虑。This work shows that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged, and shows that safety alignment requires considering the multi-agent system as a whole.

ATLAS: Aligned Transport of Latent Structure for Reliable World Model Planning
ATLAS:面向可靠世界模型规划的对齐潜在结构传输
arXiv:2609.36333 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ATLAS(Aligned Transport of Latent Structure),一种在显式保持关系几何结构的同时校准全局潜空间分布的训练目标,通过一维 Wasserstein-2 传输进行 Wasserstein 嵌入匹配来校准其边缘分布。This work introduces Aligned Transport of Latent Structure (ATLAS), a training objective that explicitly preserves relational geometry while calibrating the global latent distribution and uses Wasserstein embedding matching to calibrate its marginal through one-dimensional Wasserstein-2 transport.

It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them
关键不在图像本身:无关上下文会扰动 VLM 评判且不提供有效信息
arXiv:2609.37863 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MIST(Misleading-Image Stress Test):200 个英文句子,每句围绕一个可作比喻或字面理解的短语,配以对齐图像(描绘其读法)、误导图像(描绘相反读法)或无图像三种条件。MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all is introduced.

The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence
无尽之试:当代模型迈向超智能的数学构造
arXiv:2609.24555 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

我们提出 Endless Exam,一个涵盖十四个参数化数学构造问题族的 benchmark,具有可验证的分数,能够区分在已发表数学前沿之前与之后的进展。每个提交的对象会被自动检验有效性,并依据已发表前沿或构造基线获得相对质量分数,不将改进上限设为 1。该 benchmark 从开放性数学问题中汲取长期挑战,并通过改变参数生成更大规模的实例。紧凑证书使得大型构造能被快速验证We introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems, with verifiable scores that distinguish progress before and beyond published mathematical frontiers. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at 1. The benchmark draws long-term challenges from open mathematical problems and generates larger instances by varying their parameters. Compact certificates allow large constructions to be ve

CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering
CUA-SWE:当 Computer-Use Agents 遇见可视化软件工程
arXiv:2609.32600 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本评估刻画了前沿 agent 如何结合源代码级执行、应用截图与图形交互来生成经过验证的软件变更,考察了跨领域与不同任务信息需求下的表现,以及与成功修复相关的开发行为。This evaluation characterizes how frontier agents combine source-level execution with application screenshots and graphical interaction to produce verified software changes, and examines performance across domains and task information requirements, alongside the development behaviors associated with successful repairs.

Generative Moment Matching Networks
生成矩匹配网络
arXiv:1502.02761 多模态 方法 OA · 绿色 被引 954 · S2

本文提出一种方法,通过多层感知机的一次前馈传播生成独立样本(与近期提出的 GAN 类似),并使用 MMD 学习生成可被解码为样本的 codes。This work forms a method that generates an independent sample via a single feedforward pass through a multilayer perceptron, as in the recently proposed generative adversarial networks, using MMD to learn to generate codes that can then be decoded to produce samples.

Training LLM Judges from Language Feedback via Position-Selective Self-Distillation
通过位置选择性自蒸馏从语言反馈训练 LLM 评判器
arXiv:2609.38792 工程化 观点 OA · 绿色 被引 0 · S2 + OpenAlex

实验表明,屏蔽高熵偏移位置相比朴素的 SD 提升了分布外泛化能力,由此得到的自蒸馏 judges 在所评估的主观子类别上比基于结果监督 RL 训练的 judges 高出 2-9 个百分点,同时在客观子类别上保持竞争力。Experiments show that masking higher-entropy-shift positions improves out-of-distribution generalization over naive SD, and the resulting self-distilled judges outperform judges trained with outcome-supervised RL by 2-9 percentage points on the evaluated subjective subcategories, while remaining competitive on objective ones.

DAGent: Evaluate-then-Grow Planning for Deep Research Agents
DAGent:面向深度研究 Agent 的 Evaluate-then-Grow 规划
arXiv:2609.39154 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

DAGent 是一个基于 DAG 的多智能体框架,采用先评估再生长的增量规划:Orchestrator 逐批扩展任务图,每一步扩展都以已完成节点的置信度与不确定性信号为条件;证据条件化规划在更低的每任务 token、工具调用和步骤开销下达到了比 Plan-then-Patch 更高的准确率。DAGent, a DAG-based multi-agent framework with Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph one batch at a time, conditioning each expansion on confidence and uncertainty signals from completed nodes, which shows that evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart.

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
MILO:通过编排式多 Agent 进化实现自动化 Harness 发现
arXiv:2609.38349 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MILO(Meta-evolutionary Island Orchestration),一个共同演化 agent harness 及其发现策略的框架,使用前沿模型(Opus 4.8)与开源权重模型(gpt-oss-120b)超越了八个 SOTA harness 与六种搜索方法。This work introduces MILO (Meta-evolutionary Island Orchestration), a framework that co-evolves agent harnesses and the strategy used to discover them and outperforms eight state-of-the-art harnesses and six search methods using frontier (Opus 4.8) and open-weight (gpt-oss-120b) models.

2.3 本轮补充公开检索
arXiv:2606.14589 Agent 智能体 方法 OA · 绿色 被引 3 · S2

本文对一个自 2026 年 3 月起持续运行的个人助理 Agent 运行时中的静默失败进行纵向研究,该系统包含约 40 个定时任务、8 个 LLM 提供商、一个工具治理代理以及一个知识库记忆层,由 4,286 个单元测试和 827 项治理检查守护。A longitudinal study of silent failures in a personal-assistant agent runtime in continuous production since March 2026, with roughly 40 scheduled jobs, 8 LLM providers, a tool-governance proxy, and a knowledge-base memory plane, defended by 4,286 unit tests and 827 governance checks is presented.

Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Tacit-TTS:从自回归解码到掩码预测的高效免转录语音克隆
arXiv:2609.38658 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Tacit-TTS 是一个从 IndexTTS2 蒸馏而来的高效无需转录的零样本语音克隆系统,用掩码非自回归生成替换自回归的文本到语义解码,引入无需训练的声学长度估计,并通过 ReFlow 蒸馏加速流匹配渲染器。Tacit-TTS is presented, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2 that replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation.

PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation
PixelUMM:无编码器的统一图像与视频理解与生成
arXiv:2609.38597 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

统一多模态模型(UMM)通常依赖独立的视觉表示分别完成理解与生成,这增加了视觉上下文长度,并使其难以与既有视觉-语言预训练流程集成。近期 pixel-space modeling 的进展提供了一种无编码器的替代方案,但将该范式从图像扩展到视频并非易事:视频理解与生成采用不同的时间表示,统一视觉接口的设计仍是开放问题。本文提出 PixelUMM,一种用于统一图像与视频理解的无编码器模型……Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending this paradigm from images to videos is non-trivial: video understanding and generation adopt different temporal representations, leaving the design of a unified visual interface an open question. We present PixelUMM, an encoder-free model for unified image and video understanding

BIABench: Evaluating AI agents on real-world bioimage analysis tasks
BIABench: 在真实生物图像分析任务上评估 AI agent
arXiv:2609.34274 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 BIABench,一个由 16 个从已发表生物研究重建的任务组成的基准,保留了其科学问题、成像数据与真值标注,为评估并最终训练可靠的、面向长程的生物图像分析 agent 提供可验证的框架。BIABench is introduced, a benchmark of 16 tasks reconstructed from published biological studies that retain their scientific questions, imaging data and ground truth, and provides a verifiable framework for evaluating, and eventually training, agents for reliable long-horizon bioimage analysis.

Aligning One-Step Generative Models with Reward-Weighted Transport Distillation
用奖励加权传输蒸馏对齐单步生成模型
arXiv:2609.30840 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

理论分析表明 RWTD 的不动点分布在参考策略的 off-policy 奖励倾斜与当前模型的 on-policy 倾斜之间插值,提供了一种在奖励适配与保留先验知识之间取得平衡的原则性方法。Theoretical analysis shows that the fixed-point distributions of RWTD interpolate between off-policy reward tilting of the reference and on-policy tilting of the current model, providing a principled approach to balancing reward adaptation with retention of prior knowledge.

Walking the Embedding Space: Datastore Extraction from Multimodal RAG
漫步嵌入空间:来自多模态 RAG 的数据存储提取
arXiv:2610.01871 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出一种自适应、自动化的数据提取攻击流程,在黑盒设置下针对 MRAG(其中检索到的视觉产物本身就是答案)发起攻击,表明亟需专门面向多模态数据设计的安全防护。An adaptive and automatic data extraction attack procedure operating in a black box setting against MRAG, a configuration in which the retrieved visual artifact is itself the response, and shows the urgent need for safeguards specifically designed for multimodal data.

InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
InterEvolve:用于人形机器人 loco-manipulation 的奖励程序测试时演化
arXiv:2610.02196 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

InterEvolve 提出了一个物体感知的前向-后向行为基础模型,其在冻结身体先验上的物体残差可在测试时将关于身体或物体的奖励转化为 loco-manipulation 行为,并将任务以奖励程序的形式指定:带完成条件与可调常数的分阶段奖励。InterEvolve develops an object-aware forward-backward behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time, and specifies tasks as reward programs: staged rewards with completion conditions and tunable constants.

Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation
在部分可观测机器人操作任务上对技能级记忆的基准评测与增强
arXiv:2609.38886 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 HIDE,一个用于在部分可观测条件下评估操作记忆的基准,并提出 SEEK 框架,结合三种互补的记忆机制来保留历史证据并追踪执行状态。This work introduces HIDE, a benchmark for evaluating manipulation memory under partial observability, and proposes $SEEK, a framework combining three complementary memory mechanisms to retain historical evidence and track execution state, and proposes a framework combining three complementary memory mechanisms to retain historical evidence and track execution state.

Smaller Models, Better Rejects: Preference Distillation Scaling
更小的模型,更好的拒绝:偏好蒸馏的规模扩展
arXiv:2609.38987 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,有效的拒绝能够保留任务结构,同时限制与参考策略的耦合,并且较小的冻结模型可以低成本地提供此类拒绝能力。The results suggest that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide them at low cost.

2.3 本轮补充公开检索
arXiv:2606.14061 Agent 智能体 方法 OA · 绿色 被引 5 · S2

结果表明,纯视觉设置会降低准确率并增加 token 成本,因为 Agent 缺乏足够的符号化细节,需通过重复的视觉查询进行补偿;研究指向一种面向下一代编码 Agent 的实用文本与视觉混合设计。The results show that a strictly vision-only setup degrades accuracy and increases token cost, because agents lack sufficient symbolic detail and compensate with repeated visual queries, and point to a practical hybrid text-and-vision design for next-generation coding agents.

Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
面向异构多 Agent LLM 的免预填充跨系列 KV Cache 迁移
arXiv:2609.32259 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,HeteroFold 能够在接收端无需预填充的前提下实现高效的跨系列 KV 复用,并在全部四个长上下文基准和大多数短上下文设置上取得最佳的 cache 迁移性能。Results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill, and achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings.

DataMagic: Authoring Data Videos through Declarative Multi-Agent Orchestration
DataMagic:通过声明式多 Agent 编排制作数据可视化视频
arXiv:2609.33403 Agent 智能体 方法 OA · 绿色 被引 2 · S2

DataMagic 通过声明式多 Agent 编排,从原始表格数据自动生成数据可视化视频,在完全自动化与细粒度人工控制之间架起桥梁,并提升了创作效率、降低了感知认知负荷。DataMagic, which authors data videos from raw tabular data through declarative multi-agent orchestration, bridging full automation with fine-grained human control, and improves creation efficiency and reduces perceived cognitive load.

Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs
泛化即稳定性,而非准确率:LLM 的多轴评估
arXiv:2610.01428 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出稳定性感知的泛化目标(SAGO),一个用于衡量模型在同一输入上面对不同扰动和基准时行为变化程度的评估框架,涵盖生成一致性、内部激活、置信度以及响应镜像等多个维度的变异性。The Stability-Aware Generalization Objective (SAGO) is introduced, a framework that measures how much model behavior changes for the same input under different variations and benchmarks, capturing variability across several dimensions including generation consistency, internal activations, confidence, and response mirroring.

LOCI: Spatial Linear Memory for Streaming World Models
LOCI:面向流式世界模型的空间线性记忆
arXiv:2609.40222 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 LOCI,一种混合的空间记忆架构,同时维护键值缓存与循环记忆两种表示,在重访内容的复现上比代表性世界模型以及同配置的 full-softmax 模型都更为忠实。LOCI is introduced, a hybrid spatial-memory architecture that keeps both representations of key-value caches and recurrent memory that reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model.

PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop
PhysVista:通过感知-推理-评估闭环评测 VLM 的物理智能
arXiv:2610.00559 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

PhysVista 是一个旨在通过借鉴人类“感知-推理-评估”过程的认知闭环框架来评测 VLM 物理智能的基准,揭示了视觉识别与真实物理理解之间持续存在的差距,并为面向物理基础的多模态智能设计提供了更具原则性的方向。PhysVista is a benchmark designed to evaluate physical intelligence in VLMs through a closed cognitive loop framework inspired by the human seeing-reasoning-assessment process, highlighting a persistent gap between visual recognition and genuine physical understanding, and pointing toward more principled designs for physically grounded multimodal intelligence.

Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans
Architect-Ant:可编辑的建筑平面图自动家具布置
arXiv:2606.10953 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 AntPlan,一个包含 505 张真实专业建筑平面图、覆盖 92 类物体和十类住宅房间且具有密集家具标注的精选数据集,以及 Architect-Ant,一个用于生成家具布局的框架,可在不依赖高成本迭代式 Agent 推理的情况下直接进行约束感知的布局生成。AntPlan, a curated dataset of 505 real professional architectural floor plans with dense furniture annotations spanning 92 object classes and ten residential room categories, and Architect-Ant, a framework for generating furniture layouts are introduced, enabling direct constraint-aware layout generation without relying on costly iterative agentic inference.

JevSpawn: Adaptive Agentic Inference through Compositional Action Spaces
JevSpawn:通过组合动作空间实现自适应 Agent 推理
arXiv:2610.00437 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 JevSpawn,一种将自然语言任务规范连接到有限概率探索的组合策略,并将 JevSpawn 确立为结构化 Agent 推理的一种有前景的方法,在任务性能和导航速度上均有所提升。This work introduces JevSpawn, a compositional policy that connects natural language task specifications to finite probabilistic exploration, and establishes JevSpawn as a promising approach to structured agentic inference, with improved task performance and faster navigation.

When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs
纠错何时变为修复?面向工具使用 LLM 内部干预的机制审计
arXiv:2609.36138 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

SAKIKO 是一个审计框架,通过定向错误发现、路由器条件干预、目标解析验证以及前瞻性冻结统计许可来形式化表示修复,并确立了在声明内部修复之前必须进行结果解析裁定的必要性。SAKIKO is presented, an auditing framework that formalizes representation repair via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing and establishes the necessity of outcome-resolved adjudication before claiming internal repair.