4DAnyone 在新视角视频质量和下游 4DGS 重建上均优于先前方法,并具有稳健的野外泛化能力。4DAnyone outperforms prior methods in both novel-view video quality and downstream 4DGS reconstruction, with robust in-the-wild generalization.
论文
1749 张论文卡片
提出 Repo0,一个面向零到全代码生成的持续结构演化框架,维护显式的架构状态,实例化为双有向无环图(Dual-DAG),由需求级 DAG、组件级 DAG 及其对齐关系组成。Repo0 is presented, a continuous structural evolution framework for zero-to-all code generation that maintains an explicit architectural state instantiated as a Dual-Directed-Acyclic-Graph (Dual-DAG), consisting of a requirement-level DAG, a component-level DAG, and their alignment relation.
本文提出 IAR(Inject, Align, and Recover),一个三阶段后训练框架,将结构化文档知识注入、问答行为对齐与通用能力恢复解耦,提升面向无检索文档内化的"领域主—领域通"前沿。This work proposes IAR (Inject, Align, and Recover), a three-stage post-training framework that separates structured document knowledge injection, QA behavior alignment, and general ability recovery and improves the domain-primary domain-general frontier for retrieval-free document internalization.
本文提出一种量化基准优化的方法论,聚焦于音频对参考转写不充分确定的情形,指出高性能模型会表现出基准条件化行为,从而虚高基准得分,却未必反映通用转写能力的真正提升。This work presents a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript, and indicates that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.
一项配对消融实验在保留仓库与可执行工程上下文的同时移除显式科学指导,表明科学知识并非一律有益:可靠信息能约束修复、提升平均表现与 token 效率,而错位指导则会诱发锚定,不必然提升精确修复成功率。A paired ablation that removes explicit scientific guidance while preserving the repository and executable engineering context shows that scientific knowledge is not uniformly beneficial: well-grounded information can constrain repair and improve average performance and token efficiency, whereas poorly aligned guidance can induce anchoring and does not necessarily improve exact repair success.
本文研究 LLM 在测试时从迭代经验中学习的机制,称之为 Chain-of-Experience (CoE):模型通过与自身或环境反馈的迭代交互积累经验痕迹,形成超越零样本推理的持续改进循环。This study studies how LLMs learn from iterative experience at test time, a setting the authors refer to as Chain-of-Experience (CoE), where models accumulate experiential traces through iterative interactions with self or environmental feedback to form a continual improvement loop beyond zero-shot inference.
本文指出,持续技能演化的关键瓶颈既非编辑能力亦非迭代轮数,而在于评估反馈能否持续提供可信的演化梯度;据此提出 SkillEvo,由可信反馈生成梯度,由可控治理约束方向。This work argues that the binding constraint on sustained skill evolution is neither editing capability nor the number of iterations, but whether the evaluation feedback keeps supplying trustworthy evolution gradients, and introduces SkillEvo, in which trustworthy feedback generates the gradient and controllable governance constrains its direction.
泄露带来两类现实攻击:一个训练好的分类器可从常规自然语言输出中推断用户记忆的语义谓词;一个由 RL 训练的对抗者可从生产级风格的 Agent 中完整提取社会安全号码。Leakage enables two practical attacks: a trained classifier that infers semantic predicates about user memories from routine natural-language outputs, and an RL-trained adversary that extracts full Social Security Numbers from a production-style agent.
本文提出 NAPE(Next-Audio-Patch-Embedding prediction),一个自监督框架:因果 Transformer 仅依据因果掩码与 stop-gradient,从先前 patch 嵌入预测对数梅尔频谱图的下一 patch 嵌入。NAPE (Next-Audio-Patch-Embedding prediction), a self-supervised framework in which a causal Transformer predicts each next patch embedding of a log-mel spectrogram from the previous ones, using causal masking and stop-gradient as its sole training signal is introduced.
选取三个前沿混合专家模型在低资源语言上进行推理微调,提出六个可度量的行为维度,且每维度均设门拒绝任何与输出长度相关的指标,并报告其自家评测工具为何失效。Take three frontier mixture-of-experts models and fine-tune them to reason in a low-resource language and propose six behavioural dimensions that make changes measurable, each gated to reject any metric that correlates with output length, and report how their own instruments lied.
Bespoke-Card 在传统通用估计器与学习型估计器架构之外开辟了一条新的基数估计路径,它是一个 Agent 驱动的系统,将面向特定工作负载的基数估计器合成为可执行代码。Bespoke-Card is opening a new avenue for cardinality estimation next to classical generic estimators and learned estimator architectures, an agent-driven system that synthesizes workload-specific cardinality estimators as executable code.
介绍 NARU,一个用于评估日语长视频中叙事演进和文化理解推理能力的基准;该工作提出一种基于分层记忆的标注流水线,可将原始视频转换为结构化的事件、叙事和文化标注,并通过任务导向合成与迭代式捷径去除来生成问题。NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video, is introduced, a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal.
本文为阿拉伯法学(fiqh)构建了一个检索测试集,并基于此评估稠密、词法、混合、微调及教法学派感知(madhhab-aware)等检索策略;错误分析表明,主要挑战在于区分包含答案的段落与主题相似但不含目标教法的段落。This work builds a retrieval test collection for Arabic fiqh and uses it to evaluate dense, lexical, hybrid, fine-tuned, and madhhab-aware retrieval strategies, and presents an error analysis showing that the main challenge is distinguishing answer-bearing passages from topically similar passages that do not contain the requested ruling.
FlowEvo 是一个免训练框架,在推理时让工作流与技能协同进化:它将成功的工作流编译为可调用技能,存入持久化技能库,并通过直接执行或作为上下文来检索使用这些技能,以构建新的工作流。FlowEvo, a training-free framework in which workflows and skills co-evolve at inference time, compiles successful workflows into callable skills, stores them in a persistent bank, and uses retrieved skills either through direct execution or as context for constructing new workflows.
在领域内和分布偏移设置下,增加测试时计算可显著提升下一子任务预测准确率,这些增益进一步转化为长时序机器人操作任务中更高的闭环成功率。Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.
我们提出 TinyCast,一个注意力无关的零样本预测器,仅用 146,505 个参数输出预测分布,其前提是在该规模下,上下文中的周期结构值得通过计算而非学习方式得到。一个零参数谱检测器给出主导周期,上下文按其相位进行折叠,再由一个膨胀卷积编码器和一个分块自回归分位数解码器建模其余部分。它在 GIFT-Eval 榜单上所有可确认参数量的零样本条目中体积最小;在概率准确性方面,它划定了 size-accuracy 前沿。We introduce TinyCast, an attention-free zero-shot forecaster that emits a predictive distribution from 146,505 parameters, on the premise that at this size the periodic structure of a context is worth computing rather than learning. A zero-parameter spectral detector supplies the dominant periods, the context is folded on their phase, and a dilated convolutional encoder and a block-autoregressive quantile decoder model the rest. It is smaller than every zero-shot entry on the GIFT-Eval board whose parameter count can be established. On probabilistic accuracy it defines the size-accuracy front
这些结果支持一种分工:使用 embedding 模型处理相似度、分类与聚类任务,将 LLM 留给推理密集型的检索任务。These results support a division of labour: use embedding models for similarity, classification, and clustering, and reserve LLMs for reasoning-intensive retrieval, and reserve LLMs for reasoning-intensive retrieval.
对指令下发 Agent 的评估应报告模型配置、生成契约、执行路径、工作点及终态校验器,而不应将匹配分数视为模型的内在属性。Evaluations of command-issuing agents should report the model configuration, generation contract, execution path, operating point, and final-state validator rather than treat a matched score as an intrinsic model property.
结果表明,任务特定的 harness 进化是改进冻结 LLM Agent 的一条可行路径,但其效果存在明确的经验性边界。Results demonstrate task-specific harness evolution as a viable axis for improving frozen LLM agents under clear empirical limits.
本文提出 EnSI-RAG(Entity-Structure-Indexed Retrieval-Augmented Generation),通过构建查询无关、以实体为中心的索引,将证据定位与答案合成解耦,同时保留可追溯的源证据。This work proposes EnSI-RAG (Entity-Structure-Indexed Retrieval-Augmented Generation), a framework that constructs a query-independent, entity-centered index that separates evidence localization from answer synthesis while preserving traceable source evidence.
本文提出 AID-Guard,一种有状态的授权到效果(authorization-to-effect)闭包协议:在提交(commit)时重新校验已批准的请求与提供方状态;在歧义下仅保留一个预订;仅在收到终态结果或经认证的“无效果”并设置投递栅栏(delivery fence)后,才允许释放或生成一个后继。This work presents AID-Guard, a stateful authorization-to-effect closure protocol that revalidates the approved request and provider state at commit, retains one reservation under ambiguity, and permits release or one successor only after a terminal result or certified no effect with a delivery fence.
本研究对新兴的内存键值存储进行了全面的性能与可行性评估,突出了性能、兼容性与长期可行性(包括项目成熟度、社区支持与持续开发)之间的权衡。This study presents a comprehensive performance and viability assessment of the emerging in-memory key-value stores and highlights trade-offs between performance, compatibility, and long-term viability, including project maturity, community support, and sustained development.
介绍 Dis2Pat,一个从披露到专利(disclosure-to-patent)的数据集,其设计贴合真实专利工作流,要求直接从发明人风格的、去法律化的披露生成完整专利申请;并提出强基线 Patent-MAF,一个可在本地部署、面向专利撰写的多 Agent 框架。Dis2Pat is introduced, a disclosure-to-patent dataset that reflects realistic patenting workflows by requiring the generation of complete patent applications directly from inventor-style, de-legalized disclosures and a strong baseline named Patent-MAF is proposed, a multi-agent framework for locally deployable patent drafting.
结果显示,规约规模本身并不能预测实现质量,跨 Agent 迁移可能导致显著的、依赖具体 Agent 的性能下降;因此在异构 SDD 工作流中,不应将规约默认视为与 Agent 无关的工件(artifact)。The results show that specification size alone does not predict implementation quality and that cross-agent transfer can produce substantial agent-dependent degradation, and suggest that specifications in heterogeneous SDD workflows should not automatically be treated as agent-neutral artifacts.
PhysCaP 在 code-as-policy 框架上增加了物理信息驱动的探索层,使其能够通过交互进行显式的信息获取,并引入免训练的物理属性提取模块,仅凭机器人本体感知即可估计物体质量与刚度,无需额外传感器。PhysCaP augments code-as-policy frameworks with a physics-informed exploration layer that enables explicit information-seeking through interaction, and introduces training-free physical property extraction modules that estimate object mass and stiffness from robot proprioception without additional sensors.
在所测试的同伴排序信息流下,鲁棒的结论是词法层面的趋同,而非对一般意见的捕获,也不是合成 LLM-Agent 群体中的一般性协调优势;该结果未估计其对人类或真实生产平台的影响。The robust result is lexical convergence under the tested peer-ranked feed, not general opinion capture or a general coordination advantage in the synthetic LLM-agent populations; it does not estimate effects on people or production platforms.
本文提出 FlavorBench:一个基于版本化烹饪嵌入模型编译稠密确定性答案映射的基准,并报告了一项 3 seed 的后训练研究——在 Epicure 的 270 条最优答案上对 Qwen3-0.6B checkpoint 进行 LoRA SFT 后,在该任务集上获得 13.3 分的提升。This work introduces FlavorBench: a benchmark for Compiling Dense Deterministic Answer Maps from a Versioned Culinary Embeddings Model and presents a 3-seed post-training study where LoRA SFT of a Qwen3-0.6B checkpoint on 270 optimal answers for Epicure to score on this task-set resulted in a 13.3 point gain.
本文提出 SparsePR,一种无需训练的方法,将响应耦合划分(Response-Coupled Partitioning)与探针拟合残差重建(Probe-Fitted Residual Reconstruction)相结合,并发现划分方式既影响同一分组内查询偏好支持集之间的重叠程度,也影响稀疏输出的仿射函数对稠密与稀疏输出差异的拟合能力。This work introduces SparsePR, a training-free method combining Response-Coupled Partitioning with Probe-Fitted Residual Reconstruction, and finds that partition choice affects both the overlap among grouped queries' preferred supports and how well an affine function of the sparse output can represent the difference between dense and sparse outputs.
我们提出 Hydra-0,一种以动作流为条件的通用世界模型,将机器人动作表示为像素运动。这种共享的视觉接口使得跨具身、任务、环境和视频生成 backbone 的通用世界建模与控制成为可能,学习动作在不同场景下的后果。我们的最佳配置相比动作条件 baseline,机器人运动误差降低 90.4%,物体运动误差降低 60.2%,同时支持零样本组合与数据高效适配。在 RoboLab 基准上,Hydra-0 在 replayeWe introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configuration achieves 90.4% lower robot-motion error and 60.2% lower object-motion error than our action-conditioned baseline, while supporting zero-shot composition and data-efficient adaptation. On the RoboLab benchmark, Hydra-0 achieves a Pearson correlation of r=0.96 between replaye
提出一个全谱系的人类上下文分类法,将六个相互关联的层级整合在一起:将人类视为通过视觉外观与空间几何可观测的主体、视为通过运动学动态与交互建模的动态行动者,以及视为通过世界仿真与具身智能定位的处境化 Agent。A full-spectrum human context taxonomy is introduced that integrates six interconnected levels by viewing humans as observable subjects through visual appearance and spatial geometry, as dynamic actors through kinematic dynamics and interaction modeling, and as situated agents through world simulation and embodied agency.
提出 Patch Reparameterization,在保留原始语义通路的同时,添加一个面向重构的 patch embedding,为同一组冻结的 ViT 块提供细粒度视觉信息,在保持多模态理解能力的同时实现高保真图像重构,并取得有利的重构—生成权衡。Patch Reparameterization is introduced, which preserves the original semantic pathway while adding a reconstruction-aware patch embedding that provides fine-grained visual information to the same frozen ViT blocks, and preserves multimodal understanding while enabling high-fidelity image reconstruction and a favorable reconstruction--generation trade-off.
通过将交互式函数调用沙箱与基于证据的验证相结合,MobilePA-Bench 既可作为实用的诊断基准,也是 Agent 强化学习的交互式基础,加速可靠移动 Agent 的开发。By pairing an interactive function-calling sandbox with evidence-based verification, MobilePA-Bench serves as both a practical diagnostic benchmark and an interactive foundation for agentic reinforcement learning---accelerating the development of dependable mobile agents.
本文对 LLM 推理服务方法进行了全面综述,涵盖基础的实例级方法、深入的集群级策略、新兴的场景方向以及其他重要但零散的领域。This paper provides a comprehensive survey of LLM inference serving methods, covering fundamental instance-level approaches, in-depth cluster-level strategies, emerging scenario directions, and other miscellaneous but important areas.
Task-CoEvolve 基于以下观察:相比被一致解决或一致失败的候选任务,候选 harness 之间存在分歧的任务更能提供区分信息;它利用基于历史结果的方差加权采样,将评估聚焦在能力前沿附近的任务上。Task-CoEvolve builds on the observation that tasks on which candidate harnesses disagree are more informative for distinguishing among them than tasks that are consistently solved or failed, and uses variance-weighted sampling based on past outcomes to focus evaluation on tasks near the capability frontier.
基于语音的应用在接入检索模块之前需先通过自动语音识别(ASR)处理口头查询,因此 ASR 错误会以固定的上游约束进入 pipeline。我们通过实验验证标准检索增强生成(RAG)的两项扩展——实体图链接与迭代式 query 改写——是吸收还是放大了这些错误。基于神经 TTS 合成的四种英语口音,我们在三个多跳 QA 基准(HotpotQA、2WikiMultiHopQA 和 MuSiQue)上评测四种 RAG 配置,以干净文本 oracle 为对照。尽管结构上更丰富的 configuratioSpeech-based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module, so ASR errors enter the pipeline as a fixed upstream constraint. We empirically test whether two extensions to standard retrieval-augmented generation (RAG), entity-graph linking and iterative reformulation, absorb or amplify these errors. Using four English accents synthesized through neural TTS, we evaluate four RAG configurations on three multi-hop QA benchmarks (HotpotQA, 2WikiMultiHopQA and MuSiQue) against a clean-text oracle. Although the structurally richer configuratio
本文提出 FORGE(Fake Online Recommendations in Generative Environments),将一组固定检索网页中的真实商品在本地改写为虚假商品,并在 15 个类别、5 种消费场景下的 225 件真实商品上,衡量 LLM 推荐虚假商品的频率。This work introduces FORGE (Fake Online Recommendations in Generative Environments), which locally rewrites real products in a frozen set of retrieved web pages into fake ones and measures how often the LLM recommends the fake product, across 225 real products in 15 categories and 5 consumer scenarios.