本文探索"以代码表示 3D 形状"的范式,利用并放大 LLM 的编码能力进行 3D 建模,并提出 Procedura 框架——通过编写一个由命名零件构成、并通过类型化、可机器校验的连接关系装配而成的参数化程序来建模对象。The paradigm of 3D shape as code is explored, leveraging and scaling the coding ability of an LLM for 3D modeling, and Procedura, a novel 3D modeling agent framework that writes an object as a procedural assembly, a parametric program whose named parts are joined by typed, machine-checkable mates.
论文
430 张论文卡片 · Agent 智能体
本文提出 Reinforcement Learning with Human-Engine Verification(RLHEV),一种结合稠密引擎信号与开发过程中隐式人类接受反馈的后训练范式,用于支持强化学习后训练。This work proposes Reinforcement Learning with Human-Engine Verification (RLHEV), a post-training paradigm that combines dense engine signals with implicit human acceptance feedback from the development process to support RL post-training.
PLENA 是一个软硬件协同设计的系统,采用三条核心优化路径,具备新颖的扁平化 systolic-array 架构以及支持非对称量化方案的高效计算与存储单元(路径 2)。PLENA is a hardware-software codesigned system that applies three core optimization pathways that features a novel flattened systolic-array architecture and efficient compute and memory units that support an asymmetric quantization scheme (Pathway 2).
CrabOS 将人机交替主导复杂任务的支持从依赖桥接的应用层方案提升为原生操作系统能力,为开发与运行 AI agent 提供了新基础。CrabOS elevates support for complex tasks with alternating human and AI leadership from bridge-dependent application-level solutions to native operating-system capabilities, which provide a new foundation for developing and running AI agents.
提出 EASEL benchmark,用于评估受控的灵巧视觉工具使用,以参考引导的视觉重建为主代理任务:agent 逐步在画布上绘制以匹配参考图像。EASEL is proposed, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image.
提出 LoopArena,一个用于评估一个模型在长时任务中引导另一个独立 coding agent 能力的 benchmark,并在执行范围与成本各不相同的三个互补设置下评估该能力。This work introduces LoopArena, a benchmark for evaluating how well one model can guide a separate coding agent through a long-running task, and evaluates this ability in three complementary settings that differ in execution scope and cost.
本文综述了 agentic artifact creation,即有状态的构建过程,其中 AI 系统实体地构造或修改交付物,且中间观察结果会重定向后续工作;并提出了若干原则,使承诺与责任显式化、将反馈转化为针对性修复,以及在变更后重新验证受影响的状态。This survey examines agentic artifact creation, which is defined as stateful construction in which an AI system materially constructs or revises a deliverable and intermediate observations redirect later work, and formulate principles for keeping commitments and responsibility explicit, turning feedback into targeted repair, and revalidating affected state after change.
本文提出 DART-SD (Diamond-topology Aware Retrieval and Tuning for Self-Distillation),一种将范式从全局 forcing 转向拓扑引导的局部修正的新框架,相对传统 full-trajectory 基线取得显著提升。This work proposes DART-SD (Diamond-topology Aware Retrieval and Tuning for Self-Distillation), a novel framework that shifts the paradigm from global forcing to topology-guided localized correction, and significantly outperforms traditional full-trajectory baselines.
结果表明,可靠的 Agent 自我进化需要协同设计验证、状态接地、见证语义和恢复语言表达能力,而非仅依赖迭代提示。The results indicate that reliable agent self-evolution requires co-designing verification, state grounding, witness semantics, and recovery-language expressivity rather than relying on iterative prompting alone.
提出 MnIST-PRO 基准,将 Agent 感知能力隔离测试:把 MNIST 数字识别转换为带回看约束的序列式 glimpse 搜索任务,并评估了十个多模态模型,结果表明仅获取视觉证据是不够的,Agent 还必须能够构建并更新可靠的感知状态。MnIST-PRO is addressed, a benchmark that isolates agentic perception by converting MNIST digit recognition into a sequential, glimpse-based search task with lookback constraints, and ten multimodal models are evaluated, showing that simply acquiring visual evidence is not enough and agents must also be able to build and update a reliable perceptual state.
Pera 描述了一种持久化 Agent,围绕感知和控制组件组织,持续从情景任务执行、上下文及周围环境变化中感知服务相关信号,并利用这些信号构建生命周期任务。Pera describes a persistent agent organized around perception and control components that continually perceive service-relevant signals from episodic task executions, internal context, and changes in the surrounding environment, and use these signals to construct lifecycle tasks.
GLM-5 在真实编码任务中展现出前所未有的能力,在端到端软件工程挑战的处理上超越既有基线,并提出了新颖的异步 Agent RL 算法,进一步提升了 RL 质量。GLM-5 demonstrates unprecedented capability in real-world coding tasks, surpassing previous baselines in handling end-to-end software engineering challenges and proposing novel asynchronous agent RL algorithms that further improve RL quality.
大量实验表明,SOTA embedding 模型和静态 decompose-and-rerank 范式存在关系盲区,而 BundleWeaver 取得了显著性能提升,凸显了从原子打分转向动态关系组合的必要性。Extensive experiments demonstrate that while state-of-the-art embedding models and static decompose-and-rerank paradigms suffer from relational blindness, BundleWeaver achieves substantial performance gains, highlighting the necessity of shifting from atomic scoring to dynamic relational composition.
该工作提出 Super Library Agent 问题:Agent 顺序生成 N 个相关应用组成的组合,同时维护一个共享的 Super Library 用于跨应用可复用组件,并解决了基于候选引导的代码块摘要抽取、抽取前的代码库整合,以及利用抽取 trace 和调用图信息的上下文感知迁移。This work introduces the Super Library Agent problem, where an agent sequentially generates a portfolio of N related applications while maintaining a shared Super Library of reusable cross-application components, and addresses candidate-guided extraction over code chunk summaries, pre-extraction codebase consolidation, and context-aware migration using extraction traces and call-graph information.
本文提出 UI-Venus-2,一个面向移动、Web 和桌面环境的通用 GUI 基础 Agent,采用统一的闭环推理-行动框架,并通过集成安全感知机制来确保关键操作的可控执行。UI-Venus-2 is presented, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework that integrates safety-aware mechanisms to ensure controlled execution of consequential actions.
本工作在采购、网络安全与金融场景下评估了 5 个 LLM 作为记忆写入者、2 个 LLM 作为执行者,并提出 EAL-Bench,用于衡量持久记忆在多大程度上准确保留不断演化的授权状态,以及错误是否会向下游传播为未授权操作。This work evaluates five LLMs as memory writers and two as executors across procurement, cybersecurity, and finance and introduces EAL-Bench, which measures how accurately persistent memory preserves evolving authorization state and whether errors propagate to downstream unauthorized actions.
本文提出 agentic data cracking,将非结构化数据自适应、推测性地结构化为推理自身的副产品,是面向非结构化数据上 Agentic reasoning 的下一代数据基础设施的第一步。This work proposes agentic data cracking, a method that structures unstructured data adaptively and speculatively as a byproduct of reasoning itself, a first step toward next-generation data infrastructure for agentic reasoning over unstructured data.
为解决 LLM 与非语言 Agent 的协作问题,本文提出 latent state internalization,将子 Agent 的连续表示直接投射到 LLM 的 token 流中作为习得的状态 token,并随着动作推进环境状态而进行动态重编码。To solve LLM collaboration with non-language agents, latent state internalization is introduced, which projects the subagent's continuous representations directly into the LLM's token stream as learned state tokens, with dynamic re-encoding as actions advance the environment state.
本文提出 FoldingAgent,一个从折纸演示视频中推断显式参数化折纸程序的 Agent 框架,利用预训练 Vision-Language Model 的推理能力,并配备可模拟几何变换、验证物理合理性、检索与比较视觉内容以及评估自身预测的专用工具集。FoldingAgent is presented, an agentic framework for inferring explicit parametric folding programs directly from origami demonstration videos that leverages the reasoning power of a pre-trained Vision-Language Model equipped with a suite of specialized tools that enable the agent to simulate geometric transitions, verify physical plausibility, retrieve and compare visual content, and evaluate its own predictions.
本文提出一种知识门控的任务构建协议,将任务指令与一个包含私有约定、参考表与效用算子的紧凑工件分离,并证明被保留的任务能够改善 post-training。A knowledge-gated task-construction protocol is introduced that separates a task instruction from a compact artefact containing private conventions, reference tables, and utility operators, and it is shown that the retained tasks improve post-training.
本教程综合了推动这些汇聚方向的算法、系统与设计原则,为数据科学与数据挖掘研究者提供统一视角,涵盖将 LLM、图数据管理、图挖掘、图 ML 与 agentic 计算融合到下一代 graph-native AI 系统中。This tutorial synthesizes the algorithms, systems, and design principles driving these converging directions, offering data science and data mining researchers a unified perspective on integrating LLMs, graph data management, graph mining, graph ML, and agentic computation into next-generation graph-native AI systems.
证明了一个应用层定理:若用户在 UART 控制台输入 echo hello world,系统唯一能产生的输出即为 hello world;这证明了基于 LLM 的 Agent 能够对如此底层的细节进行推理。An application-level theorem is proved: if the user types echo hello world as input on the UART console, the only output the system can produce is hello world, which proves LLM-based agents are capable of reasoning about such low-level details.
本文提出一种分层隔离架构,结合策略感知的 ingestion、retrieval-time gating 与共享推理,并通过服务端 agentic 编排加以执行,在为多租户隔离提供天然强制点的同时,允许客户端框架保留对 agent 组合与延迟敏感操作的控制权。A layered isolation architecture combining policy-aware ingestion, retrieval-time gating, and shared inference, enforced through server-side agentic orchestration is introduced, creating natural enforcement points for multitenant isolation while allowing client-side frameworks to retain control over agent composition and latency-sensitive operations.
该工作提出 AutoTraceGT(Automated Trace analysis through Grounded Theory),首个在 Agent 轨迹上自动化 grounded theory 的多 Agent pipeline,并指出 Grounded Theory 为研究 Agent 实际行为的 ML 研究者和 Agent 开发者提供了可扩展的分析工具。This work proposes AutoTraceGT (Automated Trace analysis through Grounded Theory), the first multi-agent pipeline that automates grounded theory on agent trajectories and suggests Grounded Theory offers a scalable analytic tool for ML researchers and agent developers studying what agents actually do.
该工作提出 DRACO:Distributing Rubric-based Advantage for Credit Optimization,在训练期间动态生成 rubric 以追踪 policy 的演化能力,对已完成的轨迹一次性评分,并将该判断重新分配到负责标注 rubric 的步骤上,以在 GRPO 中产生差异化的 per-step advantage。This work proposes DRACO: Distributing Rubric-based Advantage for Credit Optimization, which generates rubrics dynamically during training to track the policy's evolving capability, scores those rubrics once per completed trajectory, and redistributes that judgment over the steps responsible for annotated rubrics to produce differentiated per-step advantages in GRPO.
提出 VeriPhy,一种可审计的物理验证系统,其中纯文本 planner 在观察任何帧之前将 prompt 编译为类型化的物理义务与静态验证的执行计划。VeriPhy, an auditable physical-verification system in which a text-only planner compiles the prompt into typed physical obligations and a statically validated execution plan before any frame is observed, is presented.
最小执行契约可在生成程序中引发主动的结构适应,将计算从无约束分配中移开,在执行前显著改善观察到的资源时间分布,建立了 substrate-aware Agent 规划的受控概念验证。A minimal execution contract induces proactive structural adaptation in generated programs, shifting computation away from unconstrained allocations and substantially improving observed resource-time profiles before execution, establishing a controlled proof of concept for substrate-aware agent planning.
该工作提出了 MaxKernel,一个多 Agent 系统,实现了 TPU kernel 开发的 three distinct paradigms:Human-in-the-Loop (HITL) Agent,用于协作式分步设计;Autonomous (Auto) Agent,执行全自动、由指标和 trace 驱动的优化循环;以及 Graph-Based Autonomous Search,将 Auto Agent 扩展以对设计空间进行全局探索。This work presents MaxKernel, a multi-agent system that implements three distinct paradigms for TPU kernel development: a Human-in-the-Loop (HITL) agent for collaborative, step-by-step design; an Autonomous (Auto) agent that executes a fully automated, metric/trace-driven optimization loop; and a Graph-Based Autonomous Search that scales the Auto agent for global exploration of the design space.
提出了 τ^τ-bench(读作 hyper-tau-bench),一个将 Agent 构建作为任务的 benchmark,将协作式 Agent 构建工作转化为面向 coding agent 的可度量目标。The $\tau^\tau$-bench (pronounced hyper-tau-bench), a benchmark that makes agent construction the task, is introduced to turn the work of cooperative agent building into a measurable target for coding agents.
提出随机反思记忆提升(SRMA),仅在 grounding 后的评估风险严格下降时才接受候选记忆,为随机评估提供置信度门控,并为分段平稳环境提供重新锚定保证。Stochastic Reflective Memory Ascent (SRMA), which accepts a candidate memory only after a grounded evaluation risk strictly decreases, is introduced and provides confidence gating for stochastic evaluation and re-anchoring guarantees for piecewise-stationary environments.
本文认为,AI agent——即以大语言模型作为主要推理引擎、动态生成与丢弃代码作为工具性资源的系统——的出现构成了对"软件"本身的根本性重构,而非渐进式的工具改进。This paper argues that the emergence of AI agents -- systems where large language models serve as the primary reasoning engine, dynamically generating and discarding code as an instrumental resource -- constitutes a fundamental restructuring of what software is, not an incremental tool improvement.
HarvestBench 是首个为避免 side effect 标定价格并将 side effect 命名为生物的 benchmark;在六个模型中有四个的每次作答遭遇 kill rate 对价格变化敏感。HarvestBench is the first benchmark to put a price on avoiding a side effect and name the side effect as a living creature and four out of six models'kill rate per answered encounter were sensitive to price changes.
本文提出 Dr. Claw,一个开源工作空间,将现有编码 Agent 可执行文件封装在可控且可审计的人机协同工作流中,而非引入另一个自主 Agent。Dr. Claw is presented, an open-source workspace that wraps existing coding-agent executors in a controllable and auditable human-in-the-loop workflow rather than introducing another autonomous agent.
本文提出 Memanto,一种面向 agentic 人工智能的通用记忆层,挑战了"必须依赖知识图谱复杂度才能实现高保真 agent 记忆"的普遍假设,并取得 SOTA 准确率。Memanto is introduced, a universal memory layer for agentic artificial intelligence that challenges the prevailing assumption that knowledge graph complexity is necessary to achieve high fidelity agent memory and achieves state of the art accuracy scores.
提出一个自演化 Agent 框架,通过识别 Agent 轨迹中不稳定、低一致性的步骤并将其转化为情节记忆,以供后续运行调用,从而缩小一致性 gap。This work presents a self-evolving agent framework that reduces the consistency gap by identifying unstable, low-consistency steps in agent trajectories and converting them into episodic memory the agent can draw on in future runs.
主张防御的运行单元应是可修订的协同 episodes,将观察到的迁移、任务权限与响应历史关联起来,并提出跨执行监控建议具有可测试性,但并不声称提出新的检测器或测得具体的遏制收益。It is argued that the operational unit of defence should be a revisable coordination episode linking observed transfers, task authority, and response history, and it makes the recommendation to monitor across executions testable without claiming a new detector or a measured containment benefit.