Bespoke-Card 在传统通用估计器与学习型估计器架构之外开辟了一条新的基数估计路径,它是一个 Agent 驱动的系统,将面向特定工作负载的基数估计器合成为可执行代码。Bespoke-Card is opening a new avenue for cardinality estimation next to classical generic estimators and learned estimator architectures, an agent-driven system that synthesizes workload-specific cardinality estimators as executable code.
论文
1640 张论文卡片 · OA 绿色
介绍 NARU,一个用于评估日语长视频中叙事演进和文化理解推理能力的基准;该工作提出一种基于分层记忆的标注流水线,可将原始视频转换为结构化的事件、叙事和文化标注,并通过任务导向合成与迭代式捷径去除来生成问题。NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video, is introduced, a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal.
本文为阿拉伯法学(fiqh)构建了一个检索测试集,并基于此评估稠密、词法、混合、微调及教法学派感知(madhhab-aware)等检索策略;错误分析表明,主要挑战在于区分包含答案的段落与主题相似但不含目标教法的段落。This work builds a retrieval test collection for Arabic fiqh and uses it to evaluate dense, lexical, hybrid, fine-tuned, and madhhab-aware retrieval strategies, and presents an error analysis showing that the main challenge is distinguishing answer-bearing passages from topically similar passages that do not contain the requested ruling.
FlowEvo 是一个免训练框架,在推理时让工作流与技能协同进化:它将成功的工作流编译为可调用技能,存入持久化技能库,并通过直接执行或作为上下文来检索使用这些技能,以构建新的工作流。FlowEvo, a training-free framework in which workflows and skills co-evolve at inference time, compiles successful workflows into callable skills, stores them in a persistent bank, and uses retrieved skills either through direct execution or as context for constructing new workflows.
在领域内和分布偏移设置下,增加测试时计算可显著提升下一子任务预测准确率,这些增益进一步转化为长时序机器人操作任务中更高的闭环成功率。Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.
我们提出 TinyCast,一个注意力无关的零样本预测器,仅用 146,505 个参数输出预测分布,其前提是在该规模下,上下文中的周期结构值得通过计算而非学习方式得到。一个零参数谱检测器给出主导周期,上下文按其相位进行折叠,再由一个膨胀卷积编码器和一个分块自回归分位数解码器建模其余部分。它在 GIFT-Eval 榜单上所有可确认参数量的零样本条目中体积最小;在概率准确性方面,它划定了 size-accuracy 前沿。We introduce TinyCast, an attention-free zero-shot forecaster that emits a predictive distribution from 146,505 parameters, on the premise that at this size the periodic structure of a context is worth computing rather than learning. A zero-parameter spectral detector supplies the dominant periods, the context is folded on their phase, and a dilated convolutional encoder and a block-autoregressive quantile decoder model the rest. It is smaller than every zero-shot entry on the GIFT-Eval board whose parameter count can be established. On probabilistic accuracy it defines the size-accuracy front
这些结果支持一种分工:使用 embedding 模型处理相似度、分类与聚类任务,将 LLM 留给推理密集型的检索任务。These results support a division of labour: use embedding models for similarity, classification, and clustering, and reserve LLMs for reasoning-intensive retrieval, and reserve LLMs for reasoning-intensive retrieval.
对指令下发 Agent 的评估应报告模型配置、生成契约、执行路径、工作点及终态校验器,而不应将匹配分数视为模型的内在属性。Evaluations of command-issuing agents should report the model configuration, generation contract, execution path, operating point, and final-state validator rather than treat a matched score as an intrinsic model property.
结果表明,任务特定的 harness 进化是改进冻结 LLM Agent 的一条可行路径,但其效果存在明确的经验性边界。Results demonstrate task-specific harness evolution as a viable axis for improving frozen LLM agents under clear empirical limits.
本文提出 EnSI-RAG(Entity-Structure-Indexed Retrieval-Augmented Generation),通过构建查询无关、以实体为中心的索引,将证据定位与答案合成解耦,同时保留可追溯的源证据。This work proposes EnSI-RAG (Entity-Structure-Indexed Retrieval-Augmented Generation), a framework that constructs a query-independent, entity-centered index that separates evidence localization from answer synthesis while preserving traceable source evidence.
本文提出 AID-Guard,一种有状态的授权到效果(authorization-to-effect)闭包协议:在提交(commit)时重新校验已批准的请求与提供方状态;在歧义下仅保留一个预订;仅在收到终态结果或经认证的“无效果”并设置投递栅栏(delivery fence)后,才允许释放或生成一个后继。This work presents AID-Guard, a stateful authorization-to-effect closure protocol that revalidates the approved request and provider state at commit, retains one reservation under ambiguity, and permits release or one successor only after a terminal result or certified no effect with a delivery fence.
本研究对新兴的内存键值存储进行了全面的性能与可行性评估,突出了性能、兼容性与长期可行性(包括项目成熟度、社区支持与持续开发)之间的权衡。This study presents a comprehensive performance and viability assessment of the emerging in-memory key-value stores and highlights trade-offs between performance, compatibility, and long-term viability, including project maturity, community support, and sustained development.
介绍 Dis2Pat,一个从披露到专利(disclosure-to-patent)的数据集,其设计贴合真实专利工作流,要求直接从发明人风格的、去法律化的披露生成完整专利申请;并提出强基线 Patent-MAF,一个可在本地部署、面向专利撰写的多 Agent 框架。Dis2Pat is introduced, a disclosure-to-patent dataset that reflects realistic patenting workflows by requiring the generation of complete patent applications directly from inventor-style, de-legalized disclosures and a strong baseline named Patent-MAF is proposed, a multi-agent framework for locally deployable patent drafting.
结果显示,规约规模本身并不能预测实现质量,跨 Agent 迁移可能导致显著的、依赖具体 Agent 的性能下降;因此在异构 SDD 工作流中,不应将规约默认视为与 Agent 无关的工件(artifact)。The results show that specification size alone does not predict implementation quality and that cross-agent transfer can produce substantial agent-dependent degradation, and suggest that specifications in heterogeneous SDD workflows should not automatically be treated as agent-neutral artifacts.
PhysCaP 在 code-as-policy 框架上增加了物理信息驱动的探索层,使其能够通过交互进行显式的信息获取,并引入免训练的物理属性提取模块,仅凭机器人本体感知即可估计物体质量与刚度,无需额外传感器。PhysCaP augments code-as-policy frameworks with a physics-informed exploration layer that enables explicit information-seeking through interaction, and introduces training-free physical property extraction modules that estimate object mass and stiffness from robot proprioception without additional sensors.
在所测试的同伴排序信息流下,鲁棒的结论是词法层面的趋同,而非对一般意见的捕获,也不是合成 LLM-Agent 群体中的一般性协调优势;该结果未估计其对人类或真实生产平台的影响。The robust result is lexical convergence under the tested peer-ranked feed, not general opinion capture or a general coordination advantage in the synthetic LLM-agent populations; it does not estimate effects on people or production platforms.
本文提出 FlavorBench:一个基于版本化烹饪嵌入模型编译稠密确定性答案映射的基准,并报告了一项 3 seed 的后训练研究——在 Epicure 的 270 条最优答案上对 Qwen3-0.6B checkpoint 进行 LoRA SFT 后,在该任务集上获得 13.3 分的提升。This work introduces FlavorBench: a benchmark for Compiling Dense Deterministic Answer Maps from a Versioned Culinary Embeddings Model and presents a 3-seed post-training study where LoRA SFT of a Qwen3-0.6B checkpoint on 270 optimal answers for Epicure to score on this task-set resulted in a 13.3 point gain.
本文提出 SparsePR,一种无需训练的方法,将响应耦合划分(Response-Coupled Partitioning)与探针拟合残差重建(Probe-Fitted Residual Reconstruction)相结合,并发现划分方式既影响同一分组内查询偏好支持集之间的重叠程度,也影响稀疏输出的仿射函数对稠密与稀疏输出差异的拟合能力。This work introduces SparsePR, a training-free method combining Response-Coupled Partitioning with Probe-Fitted Residual Reconstruction, and finds that partition choice affects both the overlap among grouped queries' preferred supports and how well an affine function of the sparse output can represent the difference between dense and sparse outputs.
我们提出 Hydra-0,一种以动作流为条件的通用世界模型,将机器人动作表示为像素运动。这种共享的视觉接口使得跨具身、任务、环境和视频生成 backbone 的通用世界建模与控制成为可能,学习动作在不同场景下的后果。我们的最佳配置相比动作条件 baseline,机器人运动误差降低 90.4%,物体运动误差降低 60.2%,同时支持零样本组合与数据高效适配。在 RoboLab 基准上,Hydra-0 在 replayeWe introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configuration achieves 90.4% lower robot-motion error and 60.2% lower object-motion error than our action-conditioned baseline, while supporting zero-shot composition and data-efficient adaptation. On the RoboLab benchmark, Hydra-0 achieves a Pearson correlation of r=0.96 between replaye
提出一个全谱系的人类上下文分类法,将六个相互关联的层级整合在一起:将人类视为通过视觉外观与空间几何可观测的主体、视为通过运动学动态与交互建模的动态行动者,以及视为通过世界仿真与具身智能定位的处境化 Agent。A full-spectrum human context taxonomy is introduced that integrates six interconnected levels by viewing humans as observable subjects through visual appearance and spatial geometry, as dynamic actors through kinematic dynamics and interaction modeling, and as situated agents through world simulation and embodied agency.
提出 Patch Reparameterization,在保留原始语义通路的同时,添加一个面向重构的 patch embedding,为同一组冻结的 ViT 块提供细粒度视觉信息,在保持多模态理解能力的同时实现高保真图像重构,并取得有利的重构—生成权衡。Patch Reparameterization is introduced, which preserves the original semantic pathway while adding a reconstruction-aware patch embedding that provides fine-grained visual information to the same frozen ViT blocks, and preserves multimodal understanding while enabling high-fidelity image reconstruction and a favorable reconstruction--generation trade-off.
通过将交互式函数调用沙箱与基于证据的验证相结合,MobilePA-Bench 既可作为实用的诊断基准,也是 Agent 强化学习的交互式基础,加速可靠移动 Agent 的开发。By pairing an interactive function-calling sandbox with evidence-based verification, MobilePA-Bench serves as both a practical diagnostic benchmark and an interactive foundation for agentic reinforcement learning---accelerating the development of dependable mobile agents.
本文对 LLM 推理服务方法进行了全面综述,涵盖基础的实例级方法、深入的集群级策略、新兴的场景方向以及其他重要但零散的领域。This paper provides a comprehensive survey of LLM inference serving methods, covering fundamental instance-level approaches, in-depth cluster-level strategies, emerging scenario directions, and other miscellaneous but important areas.
Task-CoEvolve 基于以下观察:相比被一致解决或一致失败的候选任务,候选 harness 之间存在分歧的任务更能提供区分信息;它利用基于历史结果的方差加权采样,将评估聚焦在能力前沿附近的任务上。Task-CoEvolve builds on the observation that tasks on which candidate harnesses disagree are more informative for distinguishing among them than tasks that are consistently solved or failed, and uses variance-weighted sampling based on past outcomes to focus evaluation on tasks near the capability frontier.
基于语音的应用在接入检索模块之前需先通过自动语音识别(ASR)处理口头查询,因此 ASR 错误会以固定的上游约束进入 pipeline。我们通过实验验证标准检索增强生成(RAG)的两项扩展——实体图链接与迭代式 query 改写——是吸收还是放大了这些错误。基于神经 TTS 合成的四种英语口音,我们在三个多跳 QA 基准(HotpotQA、2WikiMultiHopQA 和 MuSiQue)上评测四种 RAG 配置,以干净文本 oracle 为对照。尽管结构上更丰富的 configuratioSpeech-based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module, so ASR errors enter the pipeline as a fixed upstream constraint. We empirically test whether two extensions to standard retrieval-augmented generation (RAG), entity-graph linking and iterative reformulation, absorb or amplify these errors. Using four English accents synthesized through neural TTS, we evaluate four RAG configurations on three multi-hop QA benchmarks (HotpotQA, 2WikiMultiHopQA and MuSiQue) against a clean-text oracle. Although the structurally richer configuratio
本文提出 FORGE(Fake Online Recommendations in Generative Environments),将一组固定检索网页中的真实商品在本地改写为虚假商品,并在 15 个类别、5 种消费场景下的 225 件真实商品上,衡量 LLM 推荐虚假商品的频率。This work introduces FORGE (Fake Online Recommendations in Generative Environments), which locally rewrites real products in a frozen set of retrieved web pages into fake ones and measures how often the LLM recommends the fake product, across 225 real products in 15 categories and 5 consumer scenarios.
本文提出 GameXpert-Bench,将游戏开发的三个生命周期阶段以 coding agent 操作化为三条互补的基准轨道,并发现现有 agent 在生成可玩基础框架和实现显式需求方面更可靠,而在发现缺陷、验证运行时行为以及跨变更保持功能一致性方面能力较弱。GameXpert-Bench is introduced, which operationalizes the three lifecycle stages of game development with a coding agent as three complementary benchmark tracks, and finds current agents are more reliable at producing playable foundations and implementing explicit requirements than at discovering defects, verifying runtime behavior, and preserving functionality across changes.
本文提出 Block3D,一种块级扩散框架,将离散形状 token 序列划分为连续块,自回归地生成各块,并联合去噪当前块内的所有 token,同时引入置信度引导的块内修正机制,在每块定稿前对低置信度 token 进行修订。Block3D is proposed, a block-wise diffusion framework that partitions the discrete shape-token sequence into contiguous blocks, generates the blocks autoregressively, and jointly denoises all tokens within the current block and introduces confidence-guided intra-block correction, which revises low-confidence tokens before each block is finalized.
本文建立了一个包含超过 1,000 个细粒度编辑概念的综合性层次化分类体系,并提出一种密集监督训练策略,将多个互不干扰的概念合成到单个图像对中,显著提升了训练效率和模型整体性能。A comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts is established and a dense supervision training strategy that synthesizes multiple non-interfering concepts into single image pairs is proposed that significantly enhances both training efficiency and overall model performance.
本文提出 ARC(Advantage Regularization via Conditioning),一种通过策略条件化 rollout 分组来恢复更公平的相对比较、并结合混合奖励与熵正则化的训练方法。The proposed ARC (Advantage Regularization via Conditioning), a training recipe that restores fairer relative comparison through strategy-conditioned rollout grouping, together with hybrid rewards and entropy regularization, is proposed.
本文提出 Knowledge Triage 框架,对 agent 知识库的每一行按类型分类,并为每种类型配置独立的保留策略,同时开源发布 AgentArtifactCorpus 数据集、分类器及参考实现。Knowledge Triage, a framework that classifies each line of an agent's knowledge base by type and routes each type through its own retention policy, is addressed, and AgentArtifactCorpus, the classifier, and the reference implementation are released.
本文介绍 TianoForge,这是面向 TianoCore 开源 UEFI 固件开发生态中 bug 分诊的集成方案,部署 AI(具体为机器学习)领域的 SOTA 方法以实现自动化 bug 分诊。This integrated approach to bug triage in the TianoCore open-source UEFI firmware development ecosystem, called TianoForge, deploys the state of the art in artificial intelligence, specifically machine learning, to enable automated bug triage.
本文通过全面超越传统开环基线,证明了当前主流的单体上下文扩展策略是一种因相关性衰减而受到惩罚的架构陷阱,并确立了以顺序、反馈驱动的编排作为生成式搜索的确定性范式。By dominating classical open-loop baselines, this work proves that the prevailing strategy of monolithic context widening is an architectural trap penalized by relevance decay and establishes sequential, feedback-driven orchestration as the definitive paradigm for generative search.
提出 TRUSTMARGIN,一种免训练、即插即用的仲裁层,利用模型自身的似然对两个候选进行打分,在不微调、无需外部评判或额外生成的情况下,在直接回答与 RAG 之间进行选择。TRUSTMARGIN is proposed, a training-free, plug-and-play arbitration layer that scores the two existing candidates with the model's own likelihoods and selects between Direct and RAG without fine-tuning, external judges, or additional generation.
本文提出 User Behavioral Densing Law,为大规模用户表示学习中的 tokenization 配置选择提供实用指导,并开发了 ALGN——一种自适应变长 tokenization 方法,可改善容量分配。The proposed User Behavioral Densing Law is proposed, providing practical guidance for tokenization configuration selection in large-scale user representation learning and ALGN, an adaptive variable-length tokenization method that improves capacity allocation, is developed.
EvoMap 的结果表明,经过验证的执行经验可以被保留并共享为可复用的外部资源,使模型能够提升长工作流完成度,而无需反复承担经验探索的全部成本。EvoMap results show that verified execution experience can be retained and shared as a reusable external resource, enabling models to improve long-workflow completion without repeatedly paying the full cost of experience discovery.