研究库 论文知识库
Papers · organized/paper_cards

论文

1640 张论文卡片 · OA 绿色

开放获取 全部 绿色 · 1640
Generative Adversarial Text to Image Synthesis
生成对抗式文本到图像合成
arXiv:1605.05396 多模态 方法 OA · 绿色 被引 3461 · S2

提出一种新颖的深度架构与 GAN 形式化方法,有效衔接文本与图像建模领域的进展,将视觉概念从字符转换为像素。A novel deep architecture and GAN formulation is developed to effectively bridge advances in text and image modeling, translating visual concepts from characters to pixels.

Activation Functions: Comparison of trends in Practice and Research for Deep Learning
激活函数:深度学习实践与研究趋势对比
arXiv:1811.03378 工程化 综述 OA · 绿色 被引 1486 · S2

本文首次系统梳理了深度学习研究迄今为止的激活函数应用趋势,将实践中的使用情况与文献中的研究成果进行对比。This paper will be the first, to compile the trends in AF applications in practice against the research results from literature, found in deep learning research to date.

MMAgent-R$^2$: Learning to Rerank and Reject for Agentic mRAG
MMAgent-R$^2$:面向 Agentic mRAG 的重排序与拒答学习
arXiv:2607.07383 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 MMAgent-R$^2$,一种将视觉重排序与主动拒答作为内部验证机制的 Agentic mRAG 框架,并通过 GRPO 训练实现外部检索、内部验证与答案生成的联合优化。MMAgent-R$^2$, an agentic mRAG framework that integrates visual reranking and active rejection as its internal verification mechanism, is proposed and achieves joint optimization of external retrieval, internal verification, and answer generation via GRPO training.

A Matryoshka Hierarchical RAG for Efficient Multi-Hop Question Answering
用于高效多跳问答的 Matryoshka 分层 RAG
arXiv:2610.01767 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

MatRAG 在检索质量上优于其最强的竞争者;同时,通过避免 KG 构建和基于 LLM 的摘要降低了索引成本,并通过维度感知的相似度降低了查询成本。MatRAG outperforms its strongest competitors in terms of retrieval quality; furthermore, it reduces indexing costs by avoiding KG construction and LLM-based summarization, and lowers query-time costs through dimension-aware similarity.

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation
Ego2Act:评估自我中心视频生成中的目标导向操作
arXiv:2610.01092 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Ego2Act:一个面向目标的基准,包含来自 110 个真实日常任务的 2.640 段视频,覆盖不同的物体杂乱度与多步复杂度;同时提出 Ego2ActJudge,一种无参考评估流水线,在任务完成度与物理合理性评估上与人类共识的对齐效果优于相关基线。Ego2Act is introduced, a goal-directed benchmark featuring 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity, and Ego2ActJudge, a reference-free evaluation pipeline that achieves better task completion and physics plausibility evaluation alignment with human consensus compared to relevant baselines.

Honeycomb: Constant-Size Scene Memory Representation for Video World Models
Honeycomb:面向视频世界模型的恒定大小场景记忆表征
arXiv:2609.37690 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 Honeycomb,一种基于 HexMemory 构建的视频世界模型;HexMemory 是作者提出的低秩表示,用于在固定大小的内存中存储场景特征,共包含六个空间与时空平面。Honeycomb is introduced, a video world model built on HexMemory, the authors' proposed low-rank representation for storing scene features in a fixed-size memory with a total of six spatial and spatiotemporal planes.

X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization
X-Tree:面向高效 Agent 泛化的可复用经验 token 化方法
arXiv:2609.32993 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

在数据和预算匹配的条件下,X-Tree 相较标准方案在 WebArena 上 SR 提升最高 4.5%,在 ScienceWorld 上提升 5.8%,在 WebShop 上成功率提升 4.1%;匹配分析表明增益源自 X-Tree 结构及其三项集成设计。X-Tree improves over standard recipes at matched data and budget by up to 4.5% SR on WebArena, 5.8% SR on ScienceWorld and 4.1% success on WebShop, and matched analyses attribute the gains to the X-Tree structure and the three integrations.

OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
OpenTumorBoard:多学科肿瘤委员会讨论轨迹的真实世界基准
arXiv:2609.32810 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 OpenTumorBoard,一个包含 611 个患者案例、19,157 轮讨论、涵盖十个专科角色的基准,源自 YouTube 上 12,534 分钟公开肿瘤委员会会议录音的转写,并发布了自动化整理流水线,以支持多学科、个性化癌症决策中 LLM 的开发与评估。OpenTumorBoard is introduced, a benchmark with 611 patient cases and 19,157 discussion turns across ten specialist roles, transcribed from 12,534 minutes of publicly available tumor board recordings on YouTube, and automated curation pipeline is released to support the development and evaluation of LLMs for multidisciplinary, personalized cancer decision-making.

MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization
MemFold:通过在策略优化学习用于长上下文个性化的紧凑软记忆
arXiv:2609.36435 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MemFold,根据所支持的行为来优化固定预算的软记忆;在 PersonaMem-32K 和 PersonaMem-128K 上取得了作者所测得的最高准确率,且在更长历史长度下优势进一步扩大。This work presents MemFold, which optimizes a fixed-budget soft memory by the behavior it supports, and attains the highest accuracy the authors measure on PersonaMem-32K and PersonaMem-128K, with margins that widen at the longer history length.

Replacing Large Language Models with Jev Decision Models for Low-Latency Edge Service Orchestration
用 Jev 决策模型替代大语言模型以实现低延迟边缘服务编排
arXiv:2609.22753 Agent 智能体 应用落地 OA · 绿色 被引 10 · S2

将 Jev 面向决策的 API 集成到边缘服务编排中,在保持服务完成度的同时降低开销,并支持在有界契约下针对延迟受限的准入进行决策模型替换。Jev's decision-oriented application programming interface (API) is integrated into edge service orchestration to reduce overhead while retaining service completion, and decision-model substitution for latency-bound admission on bounded contracts is supported.

Lower bounds for multivariate independence polynomials and their generalisations
多元独立多项式及其推广的下界
arXiv:2602.02450 工程化 方法 OA · 绿色 被引 12 · S2

在统计物理中,多元硬核模型描述一个粒子系统,每个粒子拥有各自的逸度。用图论语言表述,该模型的配分函数对应多元独立多项式,即独立多项式的多重仿射推广,定义为 $Z_G(λ_1,\dots,λ_n) := \sum_{I\in\mathcal{I}(G)} \prod_{v\in I}λ_v$,其中 $\mathcal{I}(G)$ 表示 $[n]:=\{1,2,\dots,n\}$ 上图 $G$ 的所有独立集。我们证明对于 $[n]$ 上的每个简单图 $G$ 以及 $λ_1,\dots,λ_n\geq 0$,\[ Z_G(λ_1,\dots,In statistical physics, the multivariate hard-core model describes a system of particles, each of which receives its own fugacity. In graph-theoretic language, the partition function of the model translates to the multivariate independence polynomial, i.e., the multiaffine generalisation of the independence polynomial, defined by $Z_G(λ_1,\dots,λ_n) := \sum_{I\in\mathcal{I}(G)} \prod_{v\in I}λ_v$, where $\mathcal{I}(G)$ denotes the set of all independent sets in a graph $G$ on $[n]:=\{1,2,\dots,n\}$. We prove that for every simple graph $G$ on $[n]$ and $λ_1,\dots,λ_n\geq 0$, \[ Z_G(λ_1,\dots,

"I think this is the most disruptive technology": Exploring Sentiments of ChatGPT Early Adopters using Twitter Data
"我认为这是最具颠覆性的技术":基于 Twitter 数据探索 ChatGPT 早期采用者的情感
arXiv:2212.05856 评测基准 应用落地 OA · 绿色 被引 284 · S2

基于 10,732 条早期 ChatGPT 用户推文的混合方法研究,对每个主题进行深入定性情感分析,结果显示大多数早期采用者在软件开发颠覆性、娱乐与创意发挥等主题上表达了压倒性的积极情感。A mixed-method study using 10,732 tweets from early ChatGPT users to conduct an in-depth qualitative sentiment analysis of each topic, showing that the majority of the early adopters have expressed overwhelmingly positive sentiments related to topics such as Disruptions to software development, Entertainment and exercising creativity.

Interpretable Uncertainty for Adaptive Retrieval and Reasoning in Question Answering
问答中面向自适应检索与推理的可解释不确定性
arXiv:2607.07380 RAG 检索增强 方法 OA · 绿色 被引 1 · S2

提出一种基于不确定性感知框架的自适应问答方法,通过 LLM 内部表征中区分知识不足与知识歧义/冲突的显式信号,在单次前向传播中即可由隐状态高效估计。This work proposes an uncertainty-aware framework for adaptive QA based on explicit signals derived from LLM internal representations that distinguish between knowledge insufficiency and knowledge ambiguity or conflict, and efficiently estimate these from hidden states in a single forward pass.

Tracking State Footprints: How Agents Can Transact
追踪状态足迹:Agent 如何进行事务处理
arXiv:2610.03140 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文将多智能体系统(MAS)协调建模为数据管理问题,提出用智能体的状态足迹来刻画它们:即在其自身局部上下文与状态、以及编排器和外部系统状态上的读写行为。This work frames MAS coordination as a data management problem and proposes to describe agents by their state footprint: the state they read and write across their own local context and state, as well as the state of the orchestrator and external systems.

DyadMem: A Long-Term Memory Benchmark of How Agents Work with Users
DyadMem:关于 Agent 如何与用户协作的长期记忆基准
arXiv:2610.03020 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 DyadMem,一个双领域、全流程的记忆基准,伴随大量标注工作,并提出了新定义——用户条件关系型智能体记忆(URAM),以推动该领域发展。DyadMem is a dual-domain, full-pipeline memory benchmark with extensive annotation efforts for advancing the domain's development and is proposed with the proposed new definition User-conditioned Relational Agent Memory (URAM).

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Transformer 过早停止思考,一个微型 LoRA 即可修复
arXiv:2609.36585 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

预训练 Transformer 仅利用其深度的一小部分来跟踪上下文中的引用。十三个基础模型仅能可靠地跟随 1.4–3.6 行,额外的预训练循环收益甚微。在一个早期层上训练的 rank-8 LoRA 在所有模型权重冻结的情况下扩展了这一计算能力。Qwen3-8B 在 24 行链上的精确准确率从 15.5% 提升至 99%;更长训练的 LoRA 可达 50 行。Ouro-1.4B 经过四轮循环达到 60 行,八轮后至少达到 160 行。该 LoRA 启动了一场接力:程序行通过中间层的一段短距离传递其链身份。冻结的 head 逐层读取渐进式进展信号……Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressi

Persona Dosing: Calibrated Activation Steering for Graded Trait Control
人设剂量:用于分级特性控制的校准激活引导
arXiv:2609.36388 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果将控制器所学到的行为范围与该范围内请求的准确性分离开来,尽管连贯性下限并非对每种特性都成立。The results separate the behavioral range learned by a controller from the accuracy of requests within that range, although the coherence floor does not hold for every trait there.

World Embedding Benchmark
World Embedding Benchmark(世界嵌入基准)
arXiv:2610.03632 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 World Embedding Benchmark,包含来自 80 个族的 8,000 个受控仿真案例,涵盖流体力学、固体力学、动力学以及光学与电磁学,旨在强调需要联合评估物理一致性与物理属性可恢复性,并展示物理表示对改进视频生成的效用。The World Embedding Benchmark is introduced, comprising 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics, dynamics, and optics&electromagnetism, to highlight the need to evaluate physical alignment and property recoverability jointly and demonstrate the utility of physical representations for improving video generation.

CLIMB: Confidence-Guided Complementary Evidence for Multimodal Retrieval-Augmented Generation
CLIMB:面向多模态检索增强生成的置信度引导互补证据
arXiv:2610.03421 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

在 Encyclopedic-VQA 和 InfoSeek 上的实验表明,CLIMB 持续优于基于检索增强的多模态基线;消融实验显示互补池化、基于评论家的打分以及迭代置信度控制的精炼各自对最终性能均有贡献。Experiments on Encyclopedic-VQA and InfoSeek show that CLIMB consistently improves over retrieval-augmented multimodal baselines, andlations indicate that complementary pooling, critic-based scoring, and iterative confidence-controlled refinement each contribute to the final performance.

ProAR: Learning Prospective Reasoning with Autoregressive Video Models
ProAR:基于自回归视频模型的前瞻性推理学习
arXiv:2610.03664 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

ProAR 引入两个关键组件:一是将生成锚定到长远结果,二是通过非对称注意力掩码将目标帧预测集成到自回归循环中,使预测的目标帧可指导中间状态的生成而不被其干扰。ProAR introduces two key components: to anchor generation to the long-range outcome, and to integrate goal-frame prediction into the autoregressive loop via an asymmetric attention mask, enabling the predicted goal frame to guide the generation of intermediate states without being disrupted by them.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
面向具身智能的 Mixture-of-Experts 视频预训练规模化
arXiv:2607.07675 多模态 方法 OA · 绿色 被引 15 · S2

提出 LingBot-Video,一种专为具身智能设计的基于 DiT 的视频预训练范式,并将其作为首个大规模开源 MoE 视频基础模型贡献给社区,致力于在数字创意与物理执行之间搭建桥梁。LingBot-Video is presented, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence, and is contributed as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.

LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models
LexReward:面向法律语言模型的分类驱动奖励框架
arXiv:2609.39071 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

实验表明,基于评分细则的奖励能可靠地区分不同质量的法律回答,且在该偏好数据上进行 DPO 训练在所有三个维度上均提升了性能;逐维度分析进一步支持了所提分类法与奖励构建的有效性。Experiments show that the rubric-based rewards reliably distinguish legal responses of different quality and that DPO training on the preference data improves performance across all three dimensions, and Dimension-wise analyses further support the effectiveness of the proposed taxonomy and reward construction.

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation
MotorMind:为通用视觉语言模型搭建零样本机器人操控脚手架
arXiv:2609.38078 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 MotorMind,一个机器人操作框架,将 VLM 提出的中级动作与确定性机器人控制及反馈相连接,支持异步监控与后台记忆更新;研究表明,通用 VLM 在配备合适的中级动作表示和异步执行框架后,可执行有效的零样本机器人操作。MotorMind is introduced, a robot manipulation harness that connects VLM-proposed mid-level actions to deterministic robot control and feedback, with asynchronous monitoring and background memory updates, and shows that a general-purpose VLM, when equipped with an appropriate mid-level action representation and asynchronous execution harness, can perform effective zero-shot robotic manipulation.

Multilingual GSM-Symbolic: What determines capability transfer across languages?
多语言 GSM-Symbolic:跨语言能力迁移由什么决定?
arXiv:2610.03367 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Multilingual GSM-Symbolic,一个可扩展的多语言数学数据集,包含 30,000 对题项匹配的问答对,覆盖 15 种语言;并将能力差异的最大决定因素量化为模型规模、语言资源水平、推理能力与类型学距离。This work introduces Multilingual GSM-Symbolic, an extensible multilingual mathematical dataset covering 30,000 item-matched question-answer pairs and spanning 15 languages and quantifies the largest determinants of capability as model size, language resource level, reasoning, and typological distance.

Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability
高效推理训练并不总是损害 CoT 忠实性与可监控性
arXiv:2610.03509 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文用三种以不同方式施加长度压力的方法(固定生成预算、逐样本长度目标、组相对长度奖励)对多种模型进行微调,发现它们对忠实性和可监控性具有不同的影响。This work fine-tune a variety of models with three methods that apply length pressure differently, namely a fixed generation budget, a per-example length target, and a group-relative length reward, and finds that it affects faithfulness and monitorability differently.

Collective Bias Mitigation via Model Routing and Collaboration
Collective Bias Mitigation:通过模型路由与协作的集体偏见缓解
arXiv:2610.03240 工程化 应用落地 OA · 绿色 被引 1 · S2

本文首次系统地探索了不同 LLM 的有效选择与组织,以培育更公平的 LLM 回答,并展示了 CBM 显著优于独立基线。This work is the first to systematically explore the effective selection and organization of distinct LLMs to cultivate fairer LLM responses and show CBM substantially outperforms standalone baselines.

World Action Modeling with Progressive Visual Planning
World Action Modeling with Progressive Visual Planning:基于渐进视觉规划的世界动作建模
arXiv:2610.02508 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ProWAM,一种渐进式世界动作模型,联合预测动作和稀疏视觉子目标的有序序列,为任务执行过程中的动作生成提供显式视觉引导,展示了进度索引视觉前瞻对闭环控制的价值。ProWAM is presented, a progressive world action model that jointly predicts actions and an ordered sequence of sparse visual sub-goals, providing explicit visual guidance to anchor action generation throughout task execution, demonstrating the value of progress-indexed visual foresight for closed-loop control.

LVMT: Video Mask Transformer for Long-term Video Segmentation
LVMT:用于长期视频分割的视频掩码 Transformer
arXiv:2609.34895 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 TQP,一种训练策略,将视频按帧分块处理,块间传递被追踪对象信息,但仅在单个块内进行反向传播,从而在无显存溢出、推理开销或梯度消失问题的情况下实现更长时间范围的监督。TQP is introduced, a training strategy in which the model processes a video in chunks of frames, where information about tracked objects is propagated between chunks but backpropagation is only conducted in individual chunks, enabling longer temporal supervision without out-of-memory issues, inference overhead, or vanishing gradients.

Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation
Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation:面向长程自回归视频生成的 Rollout-Marginal 蒸馏
arXiv:2609.37925 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Rollout-Marginal Distillation,在 AR 预测中保留生成历史,但对每个 chunk 独立对照 chunk teacher 打分,确保其质量修正不受不完美的时序上下文影响。Rollout-Marginal Distillation is introduced, which retains the generated history for AR prediction but scores each chunk independently against a chunk teacher, ensuring its quality correction is not compromised by an imperfect temporal context.

PointWAM: 3D World Action Modeling for Dexterous Robotic Manipulation
PointWAM:面向灵巧机器人操作的 3D World Action Modeling
arXiv:2610.02840 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

一种 3D 世界动作模型,将世界分解为场景与手,并在共享时空坐标系内将二者联合预测为 3D 点轨迹,无需任务特定物体或关键点选择即可在大量人类演示视频上进行有效预训练。A 3D world action model that decomposes the world into a scene and hands and hands, and jointly forecasts both as 3D point trajectories within a shared space-time coordinate frame, which enables effective pre-training on large-scale human demonstration videos without requiring any task-specific object or keypoint selection.

EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras
EyeRobot 2.0:无需腕部摄像头的精确操作主动注视
arXiv:2610.03710 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

EyeRobot 2.0 通过旋转两个眼动视角将其注视中心对准场景中的 3D 注视点实现物理上的视觉聚焦,并将夹爪信息规范化到以注视点为参考的 SE(3) 坐标系,从而压缩待学习的动作分布规模。EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it, and takes advantage of fixation by canonicalizing gripper information into a fixation-relative SE(3) frame, which compacts the size of the action distribution to learn.

Infinite Worlds with Versatile Interactions
支持多样化交互的无限世界
arXiv:2607.07534 多模态 方法 OA · 绿色 被引 39 · S2

首次在 World Modeling 领域引入 Agentic Harness 框架:由 pilot agent 负责规划并执行角色行为,director agent 负责随着场景推进合成新的环境要素。The integration of an agentic harness within the domain of world modeling is pioneered, wherein a pilot agent is tasked with planning and executing character behaviors, while a director agent is responsible for synthesizing novel environmental elements as the scene progresses.

Improving Atomic-Fact Recall via Focused Views in Unstructured Knowledge Editing
在非结构化知识编辑中通过聚焦视图提升原子事实回忆
arXiv:2610.02772 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 FOVEATED,一种即插即用框架,通过随机偏移分配给前文上下文 key 的 Rotary Position Embedding 位置来构建每个句子的聚焦视图,并从理论上分析了 FOVEATED 如何抵消上下文引发的难度低估。FOVEATED, a plug-and-play framework that constructs focused views of each sentence by randomly shifting the Rotary Position Embedding positions assigned to the keys of its preceding context, is proposed and theoretically analyzed how FOVEATED counteracts context-induced difficulty underestimation is analyzed.

Self-Supervised Scaling of Terminal Environments for Scientific Domains
科学领域终端环境的自监督规模化
arXiv:2610.02710 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明现有科学软件能够为终端 Agent 提供可扩展且经过行为验证的监督,并提出 software-in-the-loop reconstruction,一种自监督框架,从现有软件工作流(即把结构化输入映射为输出的可执行程序)中获取参考输出与验证目标。Results indicate that existing scientific software can provide scalable, behaviorally verified supervision for terminal agents, and introduces software-in-the-loop reconstruction, a self-supervised framework that obtains reference outputs and verification targets from existing software workflows, executable programs mapping structured inputs to outputs.

CUAWright: A Minimal Unified Interface for Digital Agents
CUAWright:面向数字 Agent 的极简统一接口
arXiv:2610.04116 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

数字环境的可编程性远超其 GUI 界面所呈现的程度,且一个极简的、以终端为中心的 harness 是获得更优性能与效率的关键。Digital environments are far more programmable than their GUI interfaces imply, and a minimal terminal-focused harness is the key for better performance and efficiency, according to these results.

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
4DCodeBench:动态场景逆向图形中的 Agent 基准测试
arXiv:2610.03715 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 4DCodeBench,一个面向 4D 逆向图形(通过代码生成)的基准,Agent 以可执行图形程序的形式从视频重建动态场景,结果表明强大的静态重建能力尚未转化为对复杂动态的可靠重建。4DCodeBench is introduced, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics.