空间资产定价模型将企业间交互结构视为已知,并利用语言模型表征从企业的信息环境中推断该结构;语言模型表征充当资本市场中潜在企业间信息结构的测量工具。Spatial asset-pricing models take the structure of inter-firm interaction as given and infer that structure from firms' information environments using language-model representations, which serve as a measurement instrument for latent inter-firm information structure in capital markets.
论文
1173 张论文卡片 · 方法
结果表明:领域驱动的合成数据一致优于零样本提示与朴素合成基线;主动选择相比随机采样提升了数据效率;元数据与协同过滤信号各自提升选择质量;在低资源场景下,合成数据可优于稀缺的真实对话,并进一步对真实对话形成补充。The results show that domain-grounded synthetic data consistently outperforms zero-shot prompting and naive synthetic baselines; active selection improves data efficiency over random sampling; metadata and collaborative filtering signals each improve selection quality; and, in low-resource settings, synthetic data can outperform scarce real dialogues while further complementing them.
证明了一个应用层定理:若用户在 UART 控制台输入 echo hello world,系统唯一能产生的输出即为 hello world;这证明了基于 LLM 的 Agent 能够对如此底层的细节进行推理。An application-level theorem is proved: if the user types echo hello world as input on the UART console, the only output the system can produce is hello world, which proves LLM-based agents are capable of reasoning about such low-level details.
提出 Puffin-World,一种统一的多模态架构,集成物理理解、空间仿真与 3D 世界生成重建,无需依赖外部离线模块,可支持需要多任务协同的交错式闭环应用。Puffin-World is proposed, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules and enables interleaved closed-loop applications requiring synergy across multiple tasks.
许多重复出现的文本函数易于描述却难以用规则实现;而为每个输入调用大型远程模型会带来重复开销、延迟与对服务方的依赖。我们提出 compile by training,将自然语言规约转化为可复用的神经函数。在编译时,教师模型生成任务专属样本,用于为精简解释器训练一个小型适配器。生成的函数可在没有教师模型的情况下运行,并能像普通软件一样被存储、版本化管理与组合。在 FuzzyBench-Hard 这一子集上……Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and dependency on a provider. We present compile by training, which turns a natural-language specification into a reusable neural function. At compile time, teacher models generate task-specific examples that are used to train a small adapter for a compact interpreter. The resulting function runs without the teachers and can be stored, versioned, and composed like ordinary software. On FuzzyBench-Hard, a subset on which
该工作提出 LatentStream,一种渐进式 latent working memory 框架,将流式记忆从"存储-检索"转变为"检索-内化",在现有在线和离线视频 benchmark 上取得新的 SOTA 结果。This work introduces LatentStream, a progressive latent working memory framework that shifts streaming memory from store-and-retrieve to retrieve-and-internalize, and achieves new state-of-the-art results on existing online and offline video benchmarks.
提出 PACE 数据集,用于评估模型能否识别以自我中心知识或事件形式表达的潜在约束(这些约束使看似合理的用户请求变得不当),以及 PaceMaker 多 Agent 框架,其中专门 Agent 通过查询重构、多跳图遍历与冲突感知过滤进行协调,以检索上下文决定性证据。PACE is introduced, a dataset for evaluating whether models can identify latent constraints, expressed as egocentric knowledge or events, that render seemingly reasonable user requests inappropriate, and PaceMaker, a multi-agent framework in which specialized agents coordinate across query reformulation, multi-hop graph traversal, and conflict-aware filtering to retrieve contextually decisive evidence.
提出 CORD,第一个 post-fit adapter,通过从 calibrator 拟合中移除 preservation constraint 来修复完整校准概率向量,从而实现精确的预测保持,并将原始决策的精确恢复交由后续输出修复完成。CORD is proposed, the first post-fit adapter to impose exact prediction preservation by repairing the full calibrated probability vector by removing the preservation constraint from calibrator fitting and assigns exact recovery of the original decision to subsequent output repair.
将条件经验迁移进行形式化,并提出 Boundary-Calibrated Intervention Transfer,一种在权重变化的训练之前即授权经验复用的方法,在相同预算下取得比所评估替代方案更高的最终模型质量。Conditional experience transfer is formulated as conditional experience transfer and Boundary-Calibrated Intervention Transfer is introduced, a method that authorizes experience reuse before weight-changing training and attains higher equal-budget final-model quality than the evaluated alternatives.
该方法将在线重建重构为多参考相对位姿查询,在 Virtual KITTI、Sintel、TUM-Dynamic、ScanNet 和 7-Scenes 上取得 SOTA 性能,并在单 GPU 上 8 小时内收敛。This approach reformulates online reconstruction as multi-reference relative pose querying, which achieves state-of-the-art performance across Virtual KITTI, Sintel, TUM-Dynamic, ScanNet, and 7-Scenes and reaches convergence in 8 hours on a single GPU.
过程中发现 AKS baseline 自身存在实现 bug,且两个 harness 在相同预算下运行相同已发布规则存在 0.74 分的差距,这表明此类对比应在同一受控 harness 内进行,而非跨论文比较。Along the way, an implementation bug in the own AKS baseline and a 0.74 point gap between two harnesses running the same published rules at the same budget show why these comparisons need to happen inside one controlled harness rather than across papers.
通过将随机梯度流建模为渗流过程来研究随机梯度下降的动态,其中嵌套的架构对称性迫使子网络以离散块的形式合并,而非通过单边附着。The dynamics of Stochastic Gradient Descent is studied by modeling the stochastic gradient flow as a percolation process, in which nested architectural symmetries force subnetworks to merge in discrete blocks rather than by single-edge attachment.
该工作提出 AutoTraceGT(Automated Trace analysis through Grounded Theory),首个在 Agent 轨迹上自动化 grounded theory 的多 Agent pipeline,并指出 Grounded Theory 为研究 Agent 实际行为的 ML 研究者和 Agent 开发者提供了可扩展的分析工具。This work proposes AutoTraceGT (Automated Trace analysis through Grounded Theory), the first multi-agent pipeline that automates grounded theory on agent trajectories and suggests Grounded Theory offers a scalable analytic tool for ML researchers and agent developers studying what agents actually do.
本文提出 LoRAFusion,一种面向 LLM 的高效 LoRA 微调系统,可消除不必要的内存访问,在不付出重算或同步代价的前提下保持 compute-bound GEMM 的性能,并引入面向多任务微调的自适应批处理算法。LoRAFusion is introduced, an efficient LoRA fine-tuning system for LLMs that eliminates unnecessary memory accesses and preserves the performance of compute-bound GEMMs without incurring the cost of recomputation or synchronization and introduces an adaptive batching algorithm for multi-job fine-tuning.
提出 QCell,一种新颖的基于查询的模型,用于在显微镜场景中去重叠细胞实例,在多个 benchmark 上优于 SOTA 方法,在 ISBI2014 上取得 +2.2 AP 和 +2.7 AJI。QCell is presented, a novel query-based model that de-overlaps cell instances in microscopy scenes and outperforms state-of-the-art methods across multiple benchmarks, achieving +2.2 AP and +2.7 AJI on ISBI2014.
该工作提出 DRACO:Distributing Rubric-based Advantage for Credit Optimization,在训练期间动态生成 rubric 以追踪 policy 的演化能力,对已完成的轨迹一次性评分,并将该判断重新分配到负责标注 rubric 的步骤上,以在 GRPO 中产生差异化的 per-step advantage。This work proposes DRACO: Distributing Rubric-based Advantage for Credit Optimization, which generates rubrics dynamically during training to track the policy's evolving capability, scores those rubrics once per completed trajectory, and redistributes that judgment over the steps responsible for annotated rubrics to produce differentiated per-step advantages in GRPO.
提出 VeriPhy,一种可审计的物理验证系统,其中纯文本 planner 在观察任何帧之前将 prompt 编译为类型化的物理义务与静态验证的执行计划。VeriPhy, an auditable physical-verification system in which a text-only planner compiles the prompt into typed physical obligations and a statically validated execution plan before any frame is observed, is presented.
推导出 open-vocabulary mutual information (OVMI),一种衡量 decoder 相对于用户可能希望传达词汇的参考分布所传达信息的信息论量度,为语音 BCI 社区提供了一种原则性方法以比较异构系统、改进词汇设计并衡量领域进展。Deriving open-vocabulary mutual information (OVMI), an information-theoretic quantity that measures the information conveyed by a decoder relative to a reference distribution over the words a user may wish to communicate, provides the speech BCI community with a principled way to compare heterogeneous systems, improve vocabulary design, and measure progress in the field.
表面 prompting 未能恢复多样性,而针对入口的干预则成功:使用早期 checkpoint 的 late-layer parameter interpolation 在不损失 pass@1 的情况下将解的覆盖度提高了 37%。While surface prompting fails to recover diversity, entrance-targeted interventions succeed: late-layer parameter interpolation with early checkpoints increases solution coverage by 37% at no loss in pass@1 and late-layer parameter interpolation with early checkpoints increases solution coverage by 37% at no loss in pass@1.
本工作提出 RoboTok,一个可扩展的数据引擎:利用人类操作视频作为查询,从互联网检索与操作相关的演示以训练灵巧机器人策略,并从以演员为中心的参考系下估计的 3D 手部轨迹中学习一个潜在运动空间This work introduces RoboTok, a scalable data engine that uses a query human manipulation video to retrieve manipulation-relevant internet demonstrations for training dexterous robot policies and learns a latent motion space from 3D hand trajectories expressed in estimated actor-centered reference frames.
研究结果强调了方向特定的迁移测试、严格的 embedding space 隔离,以及在 memory migrations 中为 memory repair 保留源历史的必要性。Findings highlight the necessity of direction-specific migration testing, strict embedding space isolation, and the retention of source histories for memory repair in memory migrations.
Tangram 是一种 serving 框架,将先前系统动态处理的内容静态解析,可作为现有 non-uniform 压缩方法的即插即用底座,在匹配其精度的同时,端到端吞吐量较 full-KV 基线最高提升 2.6×。Tangram is a serving framework that statically resolves what prior systems handle dynamically, and serves as a drop-in substrate for existing non-uniform compression methods, matching their accuracy while improving end-to-end throughput by up to $2.6\times over the full-KV baseline.
涵盖数学推理、多领域 STEM、代码生成以及多轮 Agent 任务的实验表明,RISE 在所有设置下均优于仅使用 RLVR 的训练以及 on-policy self-distillation。Experiments spanning mathematical reasoning, multi-domain STEM, code generation, and multi-turn agentic tasks show that RISE outperforms RLVR-only training and on-policy self-distillation across all settings.
该工作提出了 MaxKernel,一个多 Agent 系统,实现了 TPU kernel 开发的 three distinct paradigms:Human-in-the-Loop (HITL) Agent,用于协作式分步设计;Autonomous (Auto) Agent,执行全自动、由指标和 trace 驱动的优化循环;以及 Graph-Based Autonomous Search,将 Auto Agent 扩展以对设计空间进行全局探索。This work presents MaxKernel, a multi-agent system that implements three distinct paradigms for TPU kernel development: a Human-in-the-Loop (HITL) agent for collaborative, step-by-step design; an Autonomous (Auto) agent that executes a fully automated, metric/trace-driven optimization loop; and a Graph-Based Autonomous Search that scales the Auto agent for global exploration of the design space.
该工作通过向参考解中注入受控的 AST 级 corruption,从 400 个 BigCodeBench 问题构建了一个评估框架,为每个修复任务赋予已知的最小 patch,并将 edit fidelity 定位为 code-repair 质量的一个独立维度,表明其可被度量与学习。This work constructs an evaluation framework from 400 BigCodeBench problems by injecting controlled AST-level corruptions into reference solutions, giving each repair task a known minimal patch, and positions edit fidelity as a distinct axis of code-repair quality and shows that it can be measured and learned.
循环状态写回被确立为低精度循环动态的关键决定因素,state-storage interface 被识别为量化循环推理的核心设计考量。R recurrent-state write-back is established as a key determinant of low-precision recurrent dynamics and the state-storage interface is identified as a central design consideration for quantized recurrent inference.
提出了 Motion-Omni,一个端到端框架,其中的 spoken dialogue model 原生输出显式的 facial expression 以及手部、上半身和下半身运动,这些输出直接由生成语音的 hidden states 生成。Motion-Omni is presented, an end-to-end framework in which a spoken dialogue model natively outputs explicit facial expression together with hand, upper-body and lower-body motion, generated directly from the hidden states that produce the speech.
该工作发现 reasoning operations 在 held-out 表征中是可分的,且 separability 在中间层达到峰值,并验证该结构无法由词汇或位置混淆因素解释。This work finds that reasoning operations are separable in held-out representations, with separability peaking in middle layers, and verify that this structure is not explained by lexical or positional confounds.
研究揭示沿音视频边的路由是双向的:音频可影响视频生成,视频也可影响音频生成;模型参数中编码的偏差是泄漏的主要来源之一。It is revealed that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation, and biases encoded in the model's parameters and emerges as a major contributor to leakage.
本文认为,AI agent——即以大语言模型作为主要推理引擎、动态生成与丢弃代码作为工具性资源的系统——的出现构成了对"软件"本身的根本性重构,而非渐进式的工具改进。This paper argues that the emergence of AI agents -- systems where large language models serve as the primary reasoning engine, dynamically generating and discarding code as an instrumental resource -- constitutes a fundamental restructuring of what software is, not an incremental tool improvement.
HoloWorld 是首个在统一连贯的 3D 城市世界中同时支持室内与室外生成的框架,构建于持续更新的跨尺度世界上下文之上。HoloWorld is the first framework to unify indoor and outdoor generation within a coherent 3D urban world, built on a continuously updated cross-scale world context.
提出 UniMate,一个统一的 foundation model,可从绑定骨骼的 3D 资产与文本提示合成任意骨骼的关节运动,无需测试时优化或针对每个骨骼的重新训练,在质量、泛化性与效率上均超越 SOTA 基线。UniMate is presented, a unified foundation model that synthesizes articulated motion for arbitrary skeletons from a rigged 3D asset and a text prompt, with no test-time optimization or per-skeleton retraining, and outperforms state-of-the-art baselines in quality, generalization, and efficiency.
本文将安全微调数据集中的回复拆分为两个独立部分:模板化的拒答声明与解释拒答的理由,并表明拒答声明会诱导模型依赖表层线索,从而妨碍对有害与良性查询的准确区分。This paper decomposes a response in the safety-tuning dataset into two distinct components: a boilerplate refusal statement and a rationale explaining the refusal, and shows that refusal statements impede accurate discrimination between harmful and benign queries by inducing reliance on superficial cues.
视频编辑涵盖多种编辑范式,但在单一统一框架中同时实现高质量的指令引导与主体引导编辑仍具挑战性。我们提出 EditVid,一个免训练框架,结合用于局部一致性的稀疏因果记忆、用于长程身份保持的基于对应关系的后注意力 token 注入,以及用于编辑局部性的软潜在融合。同一框架支持指令引导和参考引导的编辑,包括风格迁移、属性修改、对象插入、部分级编辑和主体替换。在Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The same framework supports instruction-guided and reference-guided edits, including style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On
为推进任务课程聚焦于大规模词分类,本竞赛设置两条互补赛道:Deep 赛道面向同一被试内部的大规模词分类,目标是追求最佳性能;Broad 赛道面向跨被试泛化。Advancing the curriculum of tasks to focus on word classification to focus on word classification at scale, two complementary tracks are presented in this competition: the Deep track targets within-subject word classification at scale, aiming at the best possible performance; the Broad track targets cross-subject generalisation.
提出 FactoSR,一个因子化强化学习框架,显式解读视觉投影所塌缩的维度,并指出强化显式的、因子化的 4D 一致性是将 VLM 演化为稳健、具有世界感知能力的推理器的关键一步。FactoSR is presented, a factorized reinforcement learning framework that explicitly interpret the dimensions collapsed by visual projection, and suggests that reinforcing explicit, factorized 4D consistency is a critical step toward evolving VLMs into robust, world-aware reasoners.