本文提出 MatryoshkaLoRA,一种受 Matryoshka 启发、面向 LoRA 的通用训练框架,通过在已有 LoRA adapter 之间插入一个固定的、经精心设计的对角矩阵来按比例缩放其子秩,从而学习到准确的层次化低秩表示。MatryoshkaLoRA is proposed, a general, Matryoshka-inspired training framework for LoRA that learns accurate hierarchical low-rank representations by inserting a fixed, carefully crafted diagonal matrix between the existing LoRA adapters to scale their sub-ranks accordingly.
论文
1753 张论文卡片
因果基础模型是经过预训练的神经网络,可在全新的数据集上通过上下文学习估计因果量(如平均处理效应),无需模型更新。Causal foundation models are pretrained neural networks that estimate causal quantities, such as the average treatment effect, on entirely new datasets using in-context learning without requiring model updates.
本文提出 Conformal Relevance 框架,利用上下文学习的示例筛选与集成构造打分函数,在保持覆盖的同时以极低人工成本提升简洁性。The Conformal Relevance framework is introduced which uses in-context learning example curation and ensembling to create a score function which maintains coverage while improving conciseness with minimal manual input.
提出了 Gander,一个原生多模态双工交互模型,基于 MiniCPM-o 4.5 构建,并通过异步 Agent 循环进一步适配实时交互,同时开源其模型、代码和数据,以推动社区进一步研究与开发。Gander is presented, a native multimodal duplex interaction model that builds on MiniCPM-o 4.5 and is further adapted for realtime interaction with an asynchronous agent loop and is released together with its models, code, and data to facilitate further research and development in the community.
Mask Forcing 是一种双噪声掩码展开策略,通过扰动自回归学生的自展开来缓解由 reverse-KL 模式寻求引发的模式坍缩,能以更高视觉质量高效改进多种自回归视频扩散蒸馏方法,且无需引入真实视频数据或额外后训练阶段。Mask Forcing, a Dual-Noise Masking Rollout strategy that perturbs the AR student self-rollout to mitigate mode collapse induced by reverse-KL mode seeking, improves multiple AR video diffusion distillation methods with higher visual quality efficiently, without incorporating real video data or additional post-training stages.
本文提出 AuK,一个开源基础模型,通过自然语言指令与音频上下文的统一接口整合语音生成与编辑,在零样本与指令控制的语音生成以及通用指令引导编辑上取得领先性能。AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context, is introduced and leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing is demonstrated.
本文证明空间覆盖与三维推理性能相关,并提出 CoVeR,一种仅使用 token 坐标、不依赖学习信号的确定性、无训练选择器,在三项三维推理基准上均超越既有 SOTA,并可作为即插即用模块泛化到四种 VLM。It is shown that spatial coverage is associated with 3D reasoning performance and CoVeR, a deterministic, training-free selector that uses only token coordinates, with no learned signals is introduced, which outperforms prior SOTAs on all three 3D reasoning benchmarks and generalizes as a plug-and-play module tested across four VLMs.
本文提出 TANGO,首个面向语言条件人形机器人在杂乱环境中通行的全身视觉语言导航框架,在视觉语言导航任务中达到 SOTA,并在需要避障的困难场景中超越强模块化基线。This work introduces TANGO, the first whole-body vision-language navigation framework for language-conditioned humanoid traversal in cluttered environments, and demonstrates state-of-the-art performance in vision-language navigation, while outperforming strong modular baselines in navigating challenging scenes requiring obstacle negotiation.
TransNormal-2 是基于 FLUX.2 的整流流框架,采用单步确定性推理,针对 VAE 解码器两侧的重建退化问题,施加 RGB 引导的残差修正以降低局部于边界的解码误差,且不自由改写粗预测。TransNormal-2, a FLUX.2-based rectified-flow framework with single-step deterministic inference that addresses VAE reconstruction degradation on both sides of the VAE decoder, and applies an RGB-guided residual correction to reduce boundary-localized decoding errors without freely rewriting the coarse prediction.
本文提出 BeaconKV,一种免训练的 KV cache 压缩方法,通过为每个全局查询簇维护紧凑代表性 beacon query 来预测哪些 KV 对将被重访,无需存储完整查询历史。BeaconKV is proposed, a training-free KV cache compression method that maintains beacon queries, compact representatives for each global query cluster, to anticipate which KV pairs will be revisited without storing the entire query history.
本文发布大规模机器人操控数据集与基准,用于诊断 VLA 模型的具身推理能力,并将 RoboSPA 确立为面向更强、更可靠、更具泛化性具身 Agent 的挑战性诊断基准。A large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models, and establishes RoboSPA as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents.
本文提出通过让 evaluator 与解决方案协同进化来自动化 evaluator 的设计,并证明突破 evaluation 瓶颈可释放 ADRS 的潜力,为下一代数据系统生成高度优化、可部署的代码。This work proposes automating the design of evaluators by co-evolving them with the solutions, demonstrating that addressing the evaluation bottleneck unlocks the potential of ADRS to generate highly optimized, deployable code for next-generation data systems.
Marigold V2 在应用于其他稠密回归任务(如表面法向量估计与本征图像分解)时取得 SOTA 结果。Marigold V2 achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition, and achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition.
提出 A*-Thought-V2,一个由 LLM 引导的几何动力学框架,将 CoT 建模为隐状态轨迹,并以显式-隐式交错潜在架构替代硬删除,引入更广义的软目标以促进更丰富的步骤级特征学习。A*-Thought-V2 is presented, a geometric dynamics of LLM guided framework that models CoT as a hidden-state trajectory and replaces hard deletion with an explicit-implicit interleaved latent architecture that reflects broader soft targets that encourage richer step-level feature learning.
发现空间步态参数对视觉域偏移更敏感,且更好的 HMR 重建并不一定带来下游步态估计的改进;提出 GaitXFormer,作为直接基于 RGB 的参考模型用于步态参数估计。It is found that spatial gait parameters are more sensitive to visual domain shift and that improved HMR reconstruction alone does not necessarily translate to improved downstream gait estimation, and GaitXFormer is introduced as a direct RGB reference model for estimating gait parameters.
提出 OpenWAM,一个开源研究栈,将世界-动作预训练转化为可控的实验项目,并发布完整栈,包括基础设施、评估协议、预训练模型和数据配方,以促进未来研究。This work introduces OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program, and releases the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.
发现合作者专业能力在模型早期层最易被解码,到网络中点前降至接近随机水平,并基于合成语料以单个模型作为初步验证。It is found that partner expertise is most decodable in the early layers and falls to near chance before the midpoint of the network, and one model on a synthetic corpus is used as an initial demonstration.
提出 UCF-Net,一个不确定性感知级联融合网络,利用 CLIP 语言对齐的语义先验与 DINO 自监督视觉结构先验,在域内与跨域评估中均取得最优平均 AUC。UCF-Net is proposed, an uncertainty-aware cascaded fusion network that harnesses CLIP's language-aligned semantic priors and DINO's self-supervised visual-structure priors and achieves the best mean AUC among the evaluated methods in both in-domain and cross-domain evaluations.
提出一个前馈式生成 Transformer,用于直接的单视图与多视图图像重光照,完全绕过显式本征属性估计,并采用置换不变的位置编码对称处理无序多视图输入,避免序列偏差。This work introduces a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation and employs permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias.
我们提出 Cadence,一种面向数值时序的误差有界有损压缩器,将 3.3 亿参数的时序基础模型(Google TimesFM-3)与自适应算术编码器相结合,保证每个样本满足 |x_t - x̂_t| ≤ τ。一项负面结论限定了设计空间:在无损编码场景下,基础模型毫无价值,因为节省的比特数仅与预测器精度呈对数关系 Δb = log_2(MAE_old / MAE_new)。因此 TimesFM-3 相对 32 阶线性预测器 1.51 倍的精度优势,在 20.28 比特中仅换取 0.60 比特,中位数增益仅 +0.03%。误差有界编码仅在一点上突破了这一限制。We present Cadence, an error-bounded lossy compressor for numeric time series pairing a 330M-parameter time-series foundation model (Google TimesFM-3) with an adaptive arithmetic coder, guaranteeing |x_t-x_t|leτ on every sample. One negative result constrains the design space: for lossless coding a foundation model is worth nothing, because bits saved are logarithmic in predictor accuracy, Δb=log_2(MAE_{old}/MAE_{new}). So the 1.51times advantage TimesFM-3 holds over a 32-tap linear predictor buys 0.60 bits of 20.28, a median gain of +0.03%. Error-bounded coding escapes this at one point: once
提出一个自演化 Agent 框架,通过识别 Agent 轨迹中不稳定、低一致性的步骤并将其转化为情节记忆,以供后续运行调用,从而缩小一致性 gap。This work presents a self-evolving agent framework that reduces the consistency gap by identifying unstable, low-consistency steps in agent trajectories and converting them into episodic memory the agent can draw on in future runs.
Noah 是一个时间感知、任务无关的生成式 Transformer 模型,对完整多模态患者旅程进行表征与预测,是该领域首个真正整体化的生成模型,支持自回归预测,并具备可选的时间控制、零样本分类与反事实干预模拟能力。Noah is a time-aware, task-agnostic, generative transformer model representing and forecasting the full multimodal patient journey, and is the first truly holistic generative model in its field, enabling autoregressive forecasting with optional time control, zero-shot classification, and counterfactual intervention simulation.
本文为 LLM serving 中的 SD(投机解码)提出了一种简单且可解释的 latency 模型,能够准确刻画实际观测到的 latency,解释为何加速比常随服务器负载上升而下降,并系统刻画了 draft length、acceptance rate 以及 verifier 与 drafter 规模在不同 serving 条件下对 latency 的影响。A simple and interpretable latency model for SD in LLM serving is developed that accurately describes observed latency, explains why speedups often diminish as server load increases, and characterizes how draft length, acceptance rate, and verifier-drafter size shape latency across serving conditions are characterized.
VDiff-Bench 提供一个针对性诊断基准,用于评估 MLLM 的比较视觉理解能力,揭示标准单图视觉-语言任务无法捕捉的失败模式。VDiff-Bench provides a targeted diagnostic for evaluating comparative visual understanding in MLLMs, exposing failures that are not captured by standard single-image vision-language tasks.
主张防御的运行单元应是可修订的协同 episodes,将观察到的迁移、任务权限与响应历史关联起来,并提出跨执行监控建议具有可测试性,但并不声称提出新的检测器或测得具体的遏制收益。It is argued that the operational unit of defence should be a revisable coordination episode linking observed transfers, task authority, and response history, and it makes the recommendation to monitor across executions testable without claiming a new detector or a measured containment benefit.
我们提出 RenderFormer-V2,一个统一的基于 transformer 的学习型神经渲染模型,可与现代基于物理的渲染系统互补,无需逐场景训练或专用代码,即可处理焦散、体积散射、环境光照、带纹理与置换的表面以及分布外材质等多种光传输效果。RenderFormer-V2 将全局光传输建模为序列到序列变换。沿袭前作,它仍采用两阶段流程:先是与视图无关的阶段,解析场景内基元到……We present 'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to modern physics-based rendering systems, that can handle diverse light-transport effects such as caustics, volumetric scattering, environment lighting, textured and displaced surfaces and out-of-distribution materials without per-scene training or specialized code. RenderFormer-V2 models global light transport as a sequence-to-sequence transformation. Following its predecessor, RenderFormer-V2 also employs a two stage process: a view-independent stage that resolves intra-scene primitive to
提出 EvoHarnessBench,一个在工具、技能与 Agent 三个维度上对可控 harness 演化条件下的 Agent 进行评测的 benchmark,并将 harness 演化确立为一项独立挑战:Agent 需要在持续演化的 harness 下保持原有有效行为EvoHarnessBench is introduced, a benchmark for evaluating agents under controlled harness evolution across three axes (tools, skills, and agents), and establishes harness evolution as a distinct challenge for building agents that can keep pace with an evolving harness while preserving previously effective behavior.
提出 Graph Machine,一种保持 O(n) 规模状态并通过稀疏动态路由访问的架构,使用边——由类似指针追逐的引用机制以可微分方式更新的指针类对象。The Graph Machine is introduced, an architecture that maintains an O(n)-sized state and accesses it through sparse, dynamic routing and uses edges - pointer-like objects updated differentiably by a referral mechanism resembling pointer chasing.
我们研究一种受治理的企业分析方法:语言模型负责解读问题,确定性 policy 负责选取并运行预先批准的分析程序,返回结果与证据。我们证明,在限定的分析类内(包含关系运算,以及聚合、比较、窗口、排序和相似度),这种限制仍可保持表达力。固定的语义、policy、数据和执行规则也使结果可复现。在 440 次运行中,三个 8B 模型生成 SQL 并在运行时选取工具,而 Qwen3-8B 仅解读意图,policyWe study a governed approach to enterprise analytics: a language model interprets the question, while deterministic policy selects and runs a pre-approved analytical program that returns both results and evidence. We show that this restriction can remain expressive within a defined analytical class, using relational operations plus aggregation, comparison, windows, ranking, and similarity. Fixed meaning, policy, data, and execution rules also make results replayable. Across 440 runs, three 8B models generated SQL and selected tools at runtime, while Qwen3-8B interpreted intent only and policy
AgentAudit 沿能力、接地性、安全与行为四个方面的十个维度评估完整执行轨迹——即指令完整性、规划器、记忆、工具选择、工具调用、工具正确性、对齐、工具忠实性、安全性与执行完整性——以精确定位导致观察到的失败的具体阶段。AgentAudit evaluates the entire execution trace across ten capability, grounding, security and behavioural dimensions, namely instruction integrity, planner, memory, tool selection, tool invocation, tool correctness, alignment, tool faithfulness, security and execution integrity, to pinpoint the exact stage responsible for an observed failure.
Glyph 是一个生产系统,将列描述生成与列类型标注这两个耦合问题建模为协同工作的 LLM Agent,并以有状态图形式编排,使多 Agent LLM 目录编制可审计且可作为生产服务运行。Glyph, a production system that frames two coupled problems, column description generation and column type annotation for data classification, as cooperating LLM agents orchestrated as stateful graphs, makes multi-agent LLM cataloging auditable and operable as a production service.
评估 AI Agent 能否作为科学家利用 SAE 工具开展自主机理发现,旨在将实验性模型理解确立为可测量的能力,推动闭环自主 AI R&D。Whether AI agents can act as scientists utilizing SAE tools for autonomous mechanistic discovery is evaluated to establish experimental model understanding as a measurable capability for closed-loop autonomous AI R&D.
开发一条 on-policy 专家修正流水线,由元层级 MLE Agent 自动化,在弱模型自身的 rollout 中定位失败回合,并请专家仅重写该回合,从而保留模型的规划风格,融合 Harness 进化与模型适配带来的收益。An on-policy expert-correction pipeline is developed, automated by a meta-level MLE agent, that localizes the failing turn in the weaker model's own rollout and asks the expert to rewrite only that turn, which preserves the model's planning style and combines the gains of harness evolution and model adaptation.
Tutti 是一种高效的 SSD-backed KV caching 方案,将 CPU 从 HBM 与 SSD 之间的关键数据与 I/O 控制路径中彻底移除,在提供近乎无限容量的同时,实现了与 DRAM-backed LMCache 几乎相当的 inference 性能。Tutti is an efficient SSD-backed KV caching solution that eliminates CPU intervention from the critical data and I/O control paths between HBM and SSDs, and achieves nearly the same inference performance as DRAM-backed LMCache, while providing almost infinite capacity.
Discovery Certification Protocol(DCP)将结果声明转化为在已注册模型、信息边界与预算下的可执行审计,并由确定性验证器基于冻结记录复现本地决策The Discovery Certification Protocol (DCP) turns an outcome claim into an executable audit under a registered model, information boundary, and budget, and a deterministic verifier reproduces these local decisions from frozen records.
提出 DianShi-RxnDB,一个通过全自动抽取与归一化流水线(整合专利文本、图像和反应路线图)构建的大规模细粒度有机反应数据平台。D DianShi-RxnDB is presented, a large-scale, fine-grained organic reaction data platform built via a fully automated extraction and normalization pipeline integrating patent text, images, and reaction schemes integrating patent text, images, and reaction schemes.