研究库 论文知识库
Papers · organized/paper_cards

论文

432 张论文卡片 · Agent 智能体

开放获取 全部 绿色 · 1640
Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architecture
面向 AI Agent 签名工作流的硬件密钥存储:一种零信任 MCP 强制执行架构
arXiv:2608.06130 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

两个维度之间的取舍在于:操作者事先能承诺的内容越少,所得到的保证就越不确定——极端情况下就只能求助于人工。The trade-off across both planes is that the less an operator can commit to in advance, the less deterministic the resulting guarantee, down to asking a human.

From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models
从经济主体到智能体经济:面向经济世界模型的系统蓝图
arXiv:2608.06020 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出构建经济世界模型(作为生成式引擎)的实施路线图:异质 Agent 在其中行动、交互、适应并与市场和制度共同演化,由此从内部生成经济动态。This paper develops an implementation roadmap for building economic world models as generative engines in which heterogeneous agents act, interact, adapt, and co-evolve with markets and institutions, thereby producing economic dynamics from the inside.

ASGE-RR: Agentic Service Graph Embedding with Revisable Reservations for Dynamic AI-Agent Calls
ASGE-RR:面向动态 AI-Agent 调用的可修订预留 Agentic 服务图嵌入
arXiv:2608.06033 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ASGE-RR,一种支持可修订预留的在线 ASGE 控制器:该在线网络控制问题将运行时揭示的工作流调用映射到服务副本与网络路径上,受容量、成本和截止时间约束;研究表明,运行时揭示的工作流结构创造了新的网络控制机会。This work presents ASGE-RR, an online ASGE controller with revisable reservations, an online network-control problem that maps runtime-revealed workflow calls to service replicas and network paths under capacity, cost and deadline constraints and suggests that runtime-revealed workflow structure creates a new network control opportunity.

Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay
[标题中文] Activity Frames:面向 Agent 记忆与回放的确定性屏幕活动编译
arXiv:2608.05784 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

将一个确定性、无模型的流水线编译进 Agent 记忆:该流水线将本地采集流切分为类型化的活动帧与有界事件片段,携带应用、站点、时间、输入量以及回指原始行的证据指针,全程无模型参与。A deterministic, zero-model pipeline is compiled into agent memory with a deterministic, zero-model pipeline that segments a local capture stream into typed activity frames, bounded episodes carrying application, site, timing, input volume, and evidence pointers back to the raw rows, with no model in the loop.

FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
FactorJEPA:将单体未来分解为布局-智能体-交互通道,面向拥挤混沌的全球南方城市场景
arXiv:2608.01049 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 FactorJEPA,将世界结构作为一等预测原语,并通过 visibility gate 与分离的子空间来组合布局、实体与交互,以保留部分可观测的 Agent 并抑制跨因子捷径。FactorJEPA is introduced, which makes world structure a first-class predictive primitive, and composes layout, entities, and interactions, using a visibility gate and separated subspaces to preserve partially observed agents and discourage cross-factor shortcuts.

PHOENIX: Fine-Tuned SLM-Powered Autonomous Satellite Lifetime Extension via Predictive Self-Healing and Multi-Agent AI Recovery
PHOENIX:基于微调 SLM 的卫星自主延寿,通过预测性自愈与多 Agent AI 恢复实现
arXiv:2608.07126 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 PHOENIX(Predictive Health On-orbit Edge Neural Intelligence eXtension),为卫星赋予自主故障推理能力,并在 ESA Anomaly Detection Benchmark 上报告了初步结果。PHOENIX (Predictive Health On-orbit Edge Neural Intelligence eXtension) is proposed to give the satellite its own fault reasoning capability, and preliminary results on the ESA Anomaly Detection Benchmark are reported.

Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events
AI 人格会成长吗?LLM Agent 在生活事件后的人格演化分析与基准测试
arXiv:2608.06485 Agent 智能体 评测集 OA · 绿色 被引 2 · S2

本文研究 11 项重大生活事件引发的人格变化,以大五人格作为心理测量锚点,并将所得轨迹与人类人格心理学的纵向证据进行对照,指出当前 PC-Agents 模拟了人类人格动态的均值,但未能模拟其形态。This work studies event-induced personality change after 11 major life events, using the Big Five traits as a psychometric anchor and interpreting the resulting trajectories against longitudinal evidence from human personality psychology, and suggests that current PC-Agents simulate the mean of human personality dynamics, but not its shape.

PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
PrivacyPeek:审计 LLM Agent 获取了什么,而不仅仅是说了什么
arXiv:2606.00152 Agent 智能体 评测集 OA · 绿色 被引 5 · S2

实验表明,对敏感信息的不必要获取广泛存在,并观察到任务完成能力与获取阶段泄露之间的相关性,因此对获取阶段隐私进行审计既紧迫也必要。The experiments show that the unnecessary acquisition of sensitive information is widespread, and a correlation between the task-completion capability and acquisition-stage leakage, and auditing acquisition-stage privacy both urgent and necessary is observed.

When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents
特权引导失配时:面向多轮 Agent 的状态匹配路由与情境化自蒸馏
arXiv:2608.05219 Agent 智能体 方法 OA · 绿色 被引 5 · S2

SMRC-SD(State-Matched Routing and Contextualized Self-Distillation)显式地决定特权轨迹应在何时、以何种方式指导 on-policy student,其表现始终优于无条件的成功全路径蒸馏。State-Matched Routing and Contextualized Self-Distillation (SMRC-SD), which explicitly determines when and how a privileged trajectory should guide an on-policy student, consistently outperforms unconditional successful full-path distillation.

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
Ouroboros:基于已评审核心演化的自进化前沿编程 Agent
arXiv:2608.08311 Agent 智能体 方法 OA · 绿色 被引 4 · S2

我们提出 Ouroboros,一个自进化的 Agent harness,其工具、提示、上下文组装与核心实现通过已评审的 commit 持续改进,并成为后续工作的运行时。核心演化以两种模式推进:在递归自由演化中,改进本身就是任务,完成一个演化周期即可调度下一周期;在经验驱动核心演化中,常规工作和社交交互暴露的 bug、粗糙之处及低效上下文构造会引发已评审的结构变更。在 Terminal-Bench 2.1 上,Opus 5 运行取得 86.74% 的得分,为该基准报告的最佳结果。We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evolution proceeds in two modes. In recursive free evolution, improvement is itself a task, and completing one evolution cycle can schedule the next. In experience-driven core evolution, ordinary work and social interaction expose bugs, rough edges, and inefficient context construction that lead to reviewed structural changes. On Terminal-Bench 2.1, an Opus 5 run scores 86.74%, the best result reporte

Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
Agent 记忆蒸馏:通过层级教师记忆赋能小型 LLM Agent
arXiv:2608.07169 Agent 智能体 方法 OA · 绿色 被引 3 · S2

本文提出 Agent Memory Distillation (AMD),一个无需训练、通过分层记忆将结构化知识从大型教师智能体迁移到小型学生智能体的框架,一致优于现有基于记忆的基线方法。Agent Memory Distillation (AMD), a training-free framework that transfers structured knowledge from a large teacher agent to a small student agent through hierarchical memory, is proposed, which consistently outperforms existing memory-based baselines.

CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems
CEAA:面向交互式计算系统的认知具身 Agent 架构
arXiv:2608.09848 Agent 智能体 应用落地 OA · 绿色 被引 1 · S2

所提架构通过提供模块化、面向实现的具身认知能力 IVA 部署框架,弥合高层智能体推理模型与实时具身执行之间的鸿沟,助力在复杂交互虚拟环境中构建可扩展、自适应且可解释的智能体。The proposed architecture contributes by providing a modular, implementation-oriented framework for the deployment of embodied, cognitive-capable IVAs and bridges the gap between high-level agent reasoning models with real-time embodied execution, for scalable, adaptive, and explainable agents in complex interactive virtual environments.

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
WeClawArena:人本 Agent 网络中跨用户 Agent 协作与安全的可审计沙箱与基准
arXiv:2608.03499 Agent 智能体 评测集 OA · 绿色 被引 3 · S2

本文提出 WeClawArena,一个面向个人工作空间多参与方 owned-agent 协作的可审计基准与运行时沙盒,基于有界运行时证据审计攻击成功情况,支持任务分解失败、隐私泄露、证据投毒以及权限路径失效等问题的诊断。WeClawArena is introduced, an auditable benchmark and runtime sandbox for multi-party owned-agent collaboration over personal workspaces and audits attack success from bounded runtime evidence, supporting diagnosis of task breakdown, privacy leakage, poisoned evidence, and invalid authority paths.

VeriForge: Mitigating Latent Knowledge Gaps in Narrative Drafting via Mixed-Initiative Scaffolding
VeriForge:通过混合主动式支架缓解叙事起草中的潜在知识缺口
arXiv:2608.09698 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

VeriForge 是一种混合主动写作系统,通过划分认知劳动,使系统在领域发现上承担主动权,而作者保留对叙事合成的完全主动权;在受控的冷启动写作任务中,专家评审者认为其产出在领域扎根方面更强。VeriForge is a mixed-initiative writing system that divides cognitive labor so that the system assumes initiative over domain discovery while the author retains full initiative over narrative synthesis, and is perceived by expert raters to produce passages with stronger domain grounding in a controlled cold-start writing task.

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
Business Arena:在真实市场环境中基准测试 LLM Agents
arXiv:2608.08621 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

本文提出 Business Arena——一个受控环境,AI agent 在其中经营跨境店铺,在长周期内向供应商采购并向买家销售,迈出了构建面向端到端商业 agent 的真实可信测试床的第一步。Business Arena, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon, is introduced, taking a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
ComBodied Agents:以人为本的 Agentic AI 新范式
arXiv:2608.10915 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本文提出 Combodied Agents——一种以人为中心的范式,借助软件工具、传感器、可穿戴设备、机器人与人工服务作为行动通道而非终极目标,在时间维度上感知、建模、预测并支持个体的人体状态轨迹。Combodied Agents is introduced, a human-centered paradigm that perceives, models, predicts, and supports individual human-state trajectories over time, using software tools, sensors, wearables, robots, and human services as action channels rather than end goals.

Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
不值得再花一个 token:面向高效深度研究 Agent 的边际价值估计
arXiv:2608.08389 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,剪枝效果更取决于剪枝应用的位置,而非具体的评分规则:早期剪枝带来最大的端到端节省,后期剪枝主要用于细化最终的合成上下文。The results show that pruning effectiveness depends more on where pruning is applied than on the specific scoring rule: early pruning yields the largest end-to-end savings, while later pruning mainly refines the final synthesis context.

InSight-doc: Agentic Visual Perception for Long-Document Understanding
InSight-doc:面向长文档理解的 Agent 视觉感知
arXiv:2608.10628 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 InSight-doc,一种智能体视觉感知框架,将视觉分辨率视为一种自适应的推理时资源,从低分辨率起步,选择性地放大高分辨率区域以获取更细粒度的证据,且不依赖任何外部检索器。This work proposes InSight-doc, an agentic visual perception framework that treats visual resolution as an adaptive reasoning-time resource that starts from low resolution and selectively zooms into high-resolution regions for finer evidence, without relying on any external retriever.

Decision Transformer: Reinforcement Learning via Sequence Modeling
Decision Transformer:通过序列建模实现强化学习
arXiv:2106.01345 Agent 智能体 方法 OA · 绿色 被引 2520 · S2

尽管方法简单,Decision Transformer 在 Atari、OpenAI Gym 和 Key-to-Door 任务上达到或超过 SOTA 无模型离线 RL 基线的性能Despite its simplicity, Decision Transformer matches or exceeds the performance of state-of-the-art model-free offline RL baselines on Atari, OpenAI Gym, and Key-to-Door tasks.

Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents
超越记忆:面向长寿 AI Agent 的事务性连续性内核
arXiv:2608.11632 Agent 智能体 方法 OA · 绿色 被引 3 · S2

本文提出连续性核 (Continuity Kernel, CK),一种激活契约,将提交前候选评估与原子状态激活解耦,将连续性定义为已接受分支头的连续且经过授权的谱系。The Continuity Kernel (CK), an activation contract that decouples off-commit candidate evaluation from atomic state activation from atomic state activation is presented, defining continuity as an unbroken, authorized lineage of accepted branch heads.

Persistent Recursive Worlds Enable Autonomous Software Evolution
持久化递归世界使自主软件演化成为可能
arXiv:2608.10450 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,长周期软件开发可以围绕持久化项目而非持久化智能体来组织;本文提出 EvoX Genesis,使软件项目保持持久,同时允许局部智能体保持有限生命周期。Results show that long-horizon software development can be organized around a persistent project rather than a persistent agent, and EvoX Genesis is introduced, which instead makes the software project persistent while allowing local agents to remain finite-lived.

The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images
视觉工具使用的幻觉:对"用图像思考"的因果审计
arXiv:2608.06270 Agent 智能体 方法 OA · 绿色 被引 2 · S2

尽管聚合准确率有所提升,但视觉工具使用在广泛的 rollout 中并不具备因果有效性。Despite aggregate accuracy gains, visual tool-use is not causally effective across a broad range of rollouts: despite aggregate accuracy gains, visual tool-use is not causally effective across a broad range of rollouts.

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
OpenART:通过开放式环境演化扩展 Agent 红队测试
arXiv:2608.00677 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出进化式马尔可夫超图攻击(EMHA),这是一种黑盒策略,通过协调授权状态转移执行反馈驱动的环境演化,无需参数更新,并将 OpenART 确立为在复杂演化环境中研究 Agent 安全性的可扩展基础。This work proposes the Evolutionary Markov Hypergraph Attack (EMHA), a black-box policy that performs feedback-driven environment evolution by coordinating authorized state transitions without requiring parameter updates, and establishes OpenART as a scalable foundation for studying agent safety in complex, evolving environments.

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
Mechanist: 作为科学仪器的 AI,用于发现智能的机制
arXiv:2608.12036 Agent 智能体 方法 OA · 绿色 被引 1 · S2

Mechanist 是一个 agentic 系统,将 AI 作为科学仪器,用于自主发现 AI 内在机制,揭示模型如何表征世界知识、形成信念、推断他人信念,以及这些机制如何在预训练中涌现。Mechanist is an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretraining.

Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control
Ready Cohorts: 在 LLM-Agent 控制中界定 GPU 机会并避免 Host 来回
arXiv:2608.12123 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

两项研究为 GPU Agent 控制建立了两个可度量的门槛:deadline 可达的 cohort 供给与观测放置,并使用固定分区份额 F、精确离线份额 P*、局部上界 U 和在线达成份额 A 对 ready-cohort 边界进行了形式化。Two studies establish two measurable gates for GPU agent control: deadline-feasible cohort supply and observation placement and formalize the ready-cohort boundary using fixed-partition share F, exact offline share P*, local upper bound U, and online achieved share A.

Self-Evolving Embodied Agents via Skill-Harness Evolution
基于 Skill-Harness 演化的自演化具身智能体
arXiv:2608.11350 Agent 智能体 方法 OA · 绿色 被引 7 · S2

本文提出 SHAPER,一种免训练具身自适应的自演化框架,保持模型参数冻结,通过目标环境 rollout 演化可复用的 skills 和 context-code harness 来改进非参数化 Agent 系统。This work proposes SHAPER, a self-evolving framework for train-free embodied adaptation that keeps model parameters frozen and improves the non-parametric agent system by evolving reusable skills and a context-code harness through target-environment rollouts.

SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries
SkillZip:面向可扩展 Agent 技能库的契约保持型图压缩
arXiv:2608.05604 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本文提出 SkillZip,一种执行感知的程序化抽象框架,对 section 级图执行保持契约的压缩,加载一个紧凑、依赖闭合的 context,并仅在需要时展开宏。SkillZip is proposed, an execution-aware procedural abstraction framework that performs contract-preserving compression over section-level graphs that hydrates a compact, dependency-closed context and expands macros only when required.

The Lumiere Project: Bayesian User Modeling for Inferring the Goals and Needs of Software Users
The Lumiere Project: Bayesian User Modeling for Inferring the Goals and Needs of Software Users
arXiv:1301.7385 Agent 智能体 方法 OA · 绿色 被引 894 · S2

本工作综述了可用于推断用户需求的贝叶斯用户模型研究,这些模型综合考虑用户的背景、操作和查询,并提出了一种智能用户界面的整体架构。This work reviews work on Bayesian user models that can be employed to infer a user's needs by considering a users' background, actions, and queries and proposes an overall architecture for an intelligent user interface.

Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
arXiv:2210.03629 Agent 智能体 方法 OA · 绿色 被引 11980 · S2

探索以交错方式使用 LLM 同时生成推理轨迹和任务特定动作,使两者产生更大协同:推理轨迹帮助模型归纳、跟踪和更新动作计划以及处理异常,而动作使其与外部源交互以获取额外信息。The use of LLMs are explored to generate both reasoning traces and task-specific actions in an interleaved manner, allowing for greater synergy between the two: reasoning traces help the model induce, track, and update action plans as well as handle exceptions, while actions allow it to interface with external sources to gather additional information.

SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models
SKILLER:面向小型语言模型可复用技能提取的语言级强化学习
arXiv:2608.10538 Agent 智能体 应用落地 OA · 绿色 被引 2 · S2

SKILLER 是一个由自然语言驱动的强化学习框架,旨在为小模型自动生成执行器特定的 skills,使用强模型作为 actor 和 critic,将小模型 Agent 系统视为环境,并通过自然语言完全传递所有强化学习信号。SKILLER is a natural-language-driven reinforcement learning framework designed to automatically generate executor-specific skills for small models, which employs a strong model as the actor and critic, treats the small-model agent system as the environment, and propagates all reinforcement learning signals entirely via natural language.

AVA-Encoder: Towards Agent-Native Video Representation Learning
AVA-Encoder:迈向面向 Agent 原生的视频表征学习
arXiv:2608.12313 Agent 智能体 方法 OA · 绿色 被引 1 · S2

提出 Agentic Video Auto-Encoder(AVA-Encoder),一种由 agentic 自我进化驱动的新型自编码框架,用于学习 agent-native 视频表示,在 shot-level 和 keyframe-level system-prompt token 使用量减少 74.3% 的同时,性能优于精心人工调优的策略。The Agentic Video Auto-Encoder (AVA-Encoder), a novel auto-encoding framework driven by agentic self-evolution to learn agent-native video representations that outperforms a carefully human-tuned policy while using 74.3% fewer shot-level and keyframe-level system-prompt tokens.

Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review
规范优先收敛与 AI 编码 Agent:在一 717k 行代码库中跨 189 个文件拆除核心架构不变量的案例研究(无测试预言机、无人工代码审查)
arXiv:2608.12440 Agent 智能体 综述 OA · 绿色 被引 0 · S2 + OpenAlex

本文报告了一项完整的、有完整记录的案例研究:在规范优先协议下,由 AI 编码 Agent 对大规模架构进行重构,期间无人工代码审查、无预先存在的预言机来验证目标行为。该任务是在一个大型相互依赖的代码库中拆除核心不变量,作者评估认为通过增量重构基本上不可行,这类变更通常需要重写。本文所述协议下,Agent 成功完成了任务。该系统包含 717,725 行This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated code and no pre-existing oracle to validate the target behaviour. The task, dismantling a central invariant across a large interdependent codebase, was assessed by the author as effectively infeasible through incremental refactoring, the kind of change that conventionally calls for a rewrite instead. Under the protocol described here, the agent completed it successfully. The system is a 717,725-li

Second Thought: Reasoning in Parallel as LLM Agents Act and Observe
Second Thought:让 LLM Agent 在执行与观察时并行推理
arXiv:2608.13667 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Second Thought,一种免训练的推理框架,在每个 Thought 阶段结束时立即 fork 四个辅助分支,与主循环并发解码,并在环境 observation 到达时将生成的 thought 合并回去。This work proposes Second Thought, a training-free inference framework that forks four auxiliary branches the instant each Thought phase concludes, decodes them concurrently with the main loop, and merges the generated thoughts back when the environment observation arrives.

Latent On-Policy Self-Distillation
Latent On-Policy Self-Distillation
arXiv:2608.13040 Agent 智能体 方法 OA · 绿色 被引 5 · S2

提出 Latent On-Policy Self-Distillation(LOPD),不再提出另一种手工设计、附带新形式 privileged context 的 OPSD 变体,而是让 teacher 的 privileged context 本身可从经验端到端学习。This work introduces Latent On-Policy Self-Distillation (LOPD), which, rather than proposing another hand-crafted OPSD variant with a newly prescribed form of privileged context, makes the teacher's privileged context itself learnable end-to-end from experience.

Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead
Nanbeige4.2-3B on Apple Silicon:修复部署 Bug 并降低 Looped Transformer 显存开销
arXiv:2608.13987 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出一种 chunked-prefill 策略,可缓解由此带来的内存容量惩罚,在 32 GiB 共享内存上将允许的上下文宽度扩展 $2.7 \times$,但即便降低了内存开销,仍需打补丁才能使 Nanbeige4.2-3B 可用。A chunked-prefill strategy is introduced which alleviates the incurred memory-capacity penalty, extending allowable context width by $2.7 \times$ on 32~GiB shared memory, however, even with the reduced memory overhead, it is shown that patches are required to render Nanbeige4.2-3B usable.

Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
Agent 抓 Agent:临床多 Agent 系统中的捷径级联与 benchmark 作弊
arXiv:2608.03744 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

探讨共享工作空间上 LLM Agent 委员会的审议过程是否能被捷径和线索(benchmark 所奖励但临床医生会忽略的)所博弈,以及委员会的社会可信度所构成的游戏。It is asked whether committees of language-model agents deliberating on a shared workspace can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore, and what games a committee is social plausibility.