结果表明,通过结合基础 LLM 改进(PaLM 2)、医学领域微调以及包括新颖集成精化方法在内的提示策略,医学问答正快速接近医生水平的表现。Results highlight rapid progress towards physician-level performance in medical question answering by leveraging a combination of base LLM improvements (PaLM 2), medical domain finetuning, and prompting strategies including a novel ensemble refinement approach.
论文
1686 张论文卡片 · OA 绿色
深入回顾文献中 GAN 相关研究,并从两个视角阐述针对三大挑战所提出的架构变体与损失变体。An in-depth review of GAN-related research in the literature is provided, and an account of the architecture-variant and loss-variants, which have been proposed to handle these three challenges from two perspectives are provided.
总结了利用人工智能以及语音与语言处理预测阿尔茨海默病相关认知衰退的现有研究发现。The existing findings on the use of artificial intelligence, speech, and language processing to predict cognitive decline in the context of Alzheimer’s disease are summarized.
本文介绍 Gemma,一族基于 Gemini 模型所使用的研究与技术构建的轻量级 SOTA 开源模型,并全面评估模型的安全性与责任性,同时详细描述模型开发过程。This work introduces Gemma, a family of lightweight, state-of-the art open models built from the research and technology used to create Gemini models, and presents comprehensive evaluations of safety and responsibility aspects of the models, alongside a detailed description of model development.
提出一种全可微、样本级影响力估计器,可在不进行耗时的逆 Hessian 乘积或凸性假设的前提下,量化因添加或移除数据而产生的参数变化。A fully differentiable, sample-level influence estimator that quantifies parameter shifts from adding or removing data, without time-consuming inverse-Hessian products or convexity assumptions is introduced.
本实证研究通过建模数据组成对计算定律及分配指数的影响,推导出指定模型规模、token 数和数据混合精确配置的效率前沿,为可预测地扩展多模态基础模型奠定了必要基础。This empirical research establishes the essential groundwork for predictably scaling multimodal foundation models by modeling the influence of data composition on compute laws and allocation exponents and derive an efficiency frontier specifying precise configurations of model size, token count, and data mixture.
该工作主张将研究构想视为多目标的联合问题,并将其建模为 Quality-Diversity (QD) 搜索;同时提出 IDEAgent,一个通过 lineage 管理思路演化的 multi-agent 框架。This work argues that research ideation should be treated as a conjunction of both objectives and framed as a Quality-Diversity (QD) search, and introduces IDEAgent, a multi-agent framework that manages the evolution of ideas through lineages.
该工作发布了 LAMAR,一种具备语言感知能力的多语种 cross encoder,在训练中兼顾语义相关性与语言连贯性,在通用多语种 reranking 基准上整体以及各语言单独评估中均达到最佳性能。This work releases LAMAR, a language aware multilingual cross encoder trained to account for both semantic relevance and language coherence, which achieves the best performance overall and across all languages examined individually on general multilingual reranking benchmarks.
近期条件视频生成模型已展现出将 3D 引擎渲染(如深度图与无纹理几何体)转化为照片级真实视频的潜力,可应用于游戏与沉浸式内容创作。此类应用要求长时程自回归生成,在持续合成新帧的同时维持持久的 3D 世界。自回归生成器以有界 KV cache 逐 chunk 合成视频,因此当相机再次访问已从上下文中驱逐的位置时,模型常会重新生成不一致的外观,尽管该Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for gaming and immersive content creation. These applications require long-horizon auto-regressive generation that continuously synthesizes new frames while preserving a persistent 3D world. Auto-regressive generators synthesize video chunk by chunk with a bounded KV cache, so when the camera revisits a location after its context has been evicted, the model often regenerates inconsistent appearance, even though the
在万亿参数策略上完整 RL 训练流水线得到实验验证,并展示了 30B 混合专家智能体的持续学习,确立了 MOLT 作为大规模智能体 RL 研究的轻量基础。The complete RL training pipeline on a one-trillion-parameter policy is experimentally validated and sustained learning with a 30B mixture-of-experts agent is demonstrated, establishing MOLT as a lightweight foundation for large-scale agentic RL research.
提出 Multi-Head Latent Control,一种轻量级层,读取冻结 LLM 或 VLM 的隐状态轨迹以生成部署时的控制信号,从而支持从部分生成的提前交接,并在多模型系统中实现更准确的干预决策。Multi-Head Latent Control is introduced, a lightweight layer that reads hidden-state trajectories from a frozen LLM or VLM to produce deployment-time control signals, enabling early handoff from partial generations and more accurate intervention decisions in multi-model systems.
VisCo 是一个训练高效的自压缩框架,复用预训练 VLM 本身作为内在压缩器,使用少量 memory token 压缩视觉信息,并将层次化信息从编码传递到解码。VisCo is a training-efficient self-compression framework that reuses the pretrained VLM itself as an intrinsic compressor that compresses visual information using a small set of memory tokens and transfers hierarchical information from encoding to decoding.
研究表明,当冻结模型与外部记忆配合、且该记忆将每个 episode 提炼为可检索的自然语言规则时,反馈信号足以支撑持续学习。It is shown that feedback is a sufficient signal for continual learning when the frozen model is paired with an external memory that distils each episode into retrievable natural-language rules when the frozen model is paired with an external memory.
本研究在 Reachy Mini 机器人平台上,通过系统提示、检索增强生成 (RAG) 和有状态的提示编排,将选定的 KBD 需求落地实现,表明 KBD 可以塑造负责任的机器人行为,并有望提升机器人辅助学习中的学习效果。This study operationalized selected KBD requirements in the Reachy Mini robot platform through system prompting, retrieval-augmented generation, and stateful prompt orchestration, indicating that KBD can shape responsible robot behavior and potentially increase learning effectiveness in robot-supported learning.
提出 Spectral Alignment,一种基于 guidance 的轻量级方法,将中间预测的功率谱校准到预先计算的先验,并与 Classifier-Free Guidance (CFG) 互补。Spectral Alignment is proposed, a lightweight, guidance-based method that calibrates the power spectrum of intermediate predictions to a pre-computed prior and is complementary to Classifier-Free Guidance (CFG).
该工作在多语句、多模态设定下研究 ASV,考察跨匿名语音聚合信息是否影响隐私,并发现帧级聚合得到的 EER 最低。This work investigates ASV in a multi-utterance, multimodal setting and examines whether aggregating information across anonymized speech impacts privacy, and finds that frame-level aggregation yields the lowest EERs.
这些结果将 tracked scattering 确立为通往高维 one-step generation 的路径,并给出一张设计图,将 diffusion 相关监督、Drift-like 动力学与 GAN-like 目标联系起来。These results establish tracked scattering as a route to high-dimensional one-step generation and provide a design map relating diffusion-related supervision, Drift-like dynamics, and GAN-like objectives.
该工作提出一个面向异常检测的免训练 agentic 框架,无需领域特定知识,旨在追踪被检测对象随时间变化的空间-时间动态与底层变换,然后基于逐对象的时间状态轨迹进行推理,在 grounding 帧中识别异常对象。This work introduces a training-free agentic framework for anomaly detection free of domain-specific knowledge, designed to track spatial-temporal dynamics and underlying transformations of detected objects over time, and then reason over the object-wise temporal state trajectories to identify abnormal objects in grounded frames.
实验跟踪器能够展示训练的进展,但更改正在运行的训练通常仍需编写特定于训练器的代码。我们提出了 Interactive Training 2,一个通过共享协议引导训练的开源控制平面。训练应用声明其暴露的设置和操作,人类与自动化控制器通过同一接口提交请求,训练循环在安全的控制点对请求进行验证并应用。定制的 Aim 工作区将实时指标与控制面板以及请求和结果的时序记录整合在一起。我们在 a[……]上展示了该系统。Experiment trackers show how training is progressing, but changing a live run still usually requires trainer-specific code. We present Interactive Training 2, an open-source control plane for steering training through a shared protocol. Training applications declare which settings and actions they expose, humans and automated controllers submit requests through the same interface, and the training loop validates and applies them at safe control points. A customized Aim workspace combines live metrics and controls with a chronological record of requests and outcomes. We demonstrate the system a
提出 DataPrep-Bench,首个统一基准,在共享的下游任务 grounding 协议下,对 LLM 驱动的数据准备在六个领域、多种 base model 上的两类能力进行联合评估。DataPrep-Bench is introduced, the first unified benchmark that jointly evaluates both capabilities under a shared downstream-grounded protocol over six domains and multiple base models of LLM-driven data preparation.
已部署的平台与其面向运维的评估共同构成了一条可信赖、统计上可靠的 AI 辅助工作流,适用于设施运维,并可推广到其他大型科学仪器。Together, the deployed platform and its operations-grounded evaluation present a promising workflow for trustworthy, statistically grounded AI assistance in facility operations, transferable to other large scientific instruments.
DeCoRAG 是一个多模态 Graph RAG pipeline,将知识处理从耦合的视觉-语义推理转向认知层面的 Decoupling,进而把推理空间从稠密、带噪的背景推向纯净、意图驱动的语义簇。DeCoRAG is a multimodal Graph RAG pipeline that shifts knowledge processing from coupled visual-semantic reasoning to cognitive Decoupling, and subsequently drives the reasoning space from dense, noisy backgrounds to purified, intent-driven semantic clusters.
结果表明,RAG 增强的 LLM 能显著提升生成文本的事实一致性、领域专属性与规范精度,同时降低产生无支持内容的风险;本地部署的 RAG 增强 LLM 不应仅被视为文本生成工具,而应作为认知计算基础设施中的语义处理模块,在法律和信息高度动态的环境中支撑合规与组织决策。The results demonstrate that augmenting LLMs with RAG significantly improves the factual consistency, domain specificity and normative precision of generated texts while reducing the risk of unsupported content generation and indicate that locally deployed LLMs enhanced with RAG should be regarded not merely as text generation tools but as semantic processing modules within cognitive computing infrastructures supporting regulatory compliance and organizational decision-making in environments characterized by high legal and informational volatility.
将 Positive--Direction Matching (PDM)——一种分支感知的 OPD 目标,分别约束正预测方向与 CFG 条件方向——引入 dense-to-sparse 视频控制;由于朴素的 guided matching 对推理 guidance 尺度极为敏感,分支感知监督可实现更鲁棒、更有效的知识迁移。Positive--Direction Matching (PDM), a branch-aware OPD objective that separately constrains the positive prediction and the CFG conditional direction, is introduced to dense-to-sparse video control, where naive guided matching is highly sensitive to inference guidance scales, while branch-aware supervision enables more robust and effective knowledge transfer.
提出一个统一且受控的多轮环境,可在三阶段精确控制长程规划;研究表明,多教师 on-policy 蒸馏 (MOPD) 通过收敛到跨环境共享的规划模式来整合多种能力。A unified and controlled multi-turn environment that enables precise control of long-horizon planning across three stages is introduced and it is shown that multi-teacher on-policy distillation (MOPD) integrates capabilities by converging to shared planning-pattern across environments.
本研究考察是否存在一条实用路径,在没有坐标界面、且无需高成本区域级监督的条件下提升归因效果,并指出了这样一条可行路径。A study investigates whether there is a practical path to improve attribution without a coordinate interface and without costly region-level supervision, and indicates a practical path to improve attribution without a coordinate interface and without costly region-level supervision.
提出一种面向历史文档修复的新框架,利用搭载 RAG 的大语言模型,有效缓解了推断上下文相关专有名词的难题。A novel framework for historical document restoration that leverages large language models with retrieval-augmented generation (RAG) and effectively mitigates the challenge of inferring context-dependent proper nouns is introduced.
本文提出无需训练的 Sol-Attn(Sparsifying online attention),在单次 online-softmax pass 中统一动态路由、稀疏计算与近似修正,在稀疏注意力中取得更好的精度–效率权衡。This paper introduces training-free Sol-Attn (Sparsifying online attention), which unifies dynamic routing, sparse computation, and approximation correction in a single online-softmax pass, achieving a better accuracy-efficiency trade-off in sparse attention.
本文提出 Chamaileon,通过将问题建模为跨上下文结合景观(cross-context binding landscape modeling),统一多目标与多态 binder 设计,有效生成可适配多样构象景观与多目标需求的序列。Chamaileon is introduced, which unifies multi-target and multi-state binder design by formulating the problem as cross-context binding landscape modeling and effectively generates sequences adaptable to diverse conformational landscapes and multi-target requirements.
即使 NVIDIA 控制流 ISA 与重汇聚(reconvergence)机制持续演进,Divergence 仍能保持稳定且可预期的性能开销。Divergence retains a stable and predictable performance cost even as NVIDIA's control-flow ISA and reconvergence mechanisms continue to evolve.
本文提出 IndicTalk,目前最大的多语言印度语码混合(code-mixed)会话语料库,包含超过 13,28,604 段事件驱动的多轮对话,覆盖 9 种印度语言的 18 种语言变体,将公开发布以支持低资源印度语多语言会话 AI 的开发与评估。IndicTalk is presented, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages and will be released to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages.
RedE 利用 final-answer attention 作为自动监督信号来塑造 step-level 表征空间,使其中的噪声步骤可被可靠识别与过滤,并在检测性能上超越有竞争力的基线。RedE leverages final-answer attention as an automatic supervision signal to shape the step-level representation space, yielding refined embeddings in which noisy steps can be reliably identified and filtered and improves detection performance over competitive baselines.
本文提出 PAJAMA,该系统将程序合成为 judge,将其决策聚合为联合裁决(joint verdict),并通过 fallback 机制选择性地将低置信度用例升级交由 LLM 处理。PAJAMA is introduced, a system that synthesizes programs as judges, aggregates their decisions into a joint verdict, and incorporates a fallback mechanism to selectively escalate low-confidence cases to an LLM.
WorldDiT 是一种统一的 diffusion Transformer 架构,将动作生成与视觉世界建模耦合,无需大型预训练 VLM 动作主干即取得强性能,在报告全部四个 suite 的方法中,其总模型参数量与平均成功率处于已报告的 Pareto 前沿上。WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves strong performance without a large pretrained VLM action backbone, lies on the reported Pareto frontier for total model parameters and mean success among methods reporting all four suites.
大量实验表明,结合端上延迟感知设计与定制化训练策略,建立了高效 LVLM 编码的新 SOTA,在端上以近 1.7 倍速度运行的同时显著优于现有以编码器为中心的基线。Extensive experiments demonstrate that the on-device latency-informed design combined with the tailored training strategy establishes a new state-of-the-art for efficient LVLM encoding, significantly outperforming existing encoder-centric baselines while operating on-device at nearly 1.7xthe speed.
本文提出一种面向由一个或多个 agent 构建的自动化研究系统的词汇表,使其设计选择更易于描述与比较,从而将结构性设计问题——例如 agent 应在何时通信、获得或失去某项能力,或在多次运行间传递信息——转化为可测试的选择。A vocabulary for automated research systems built from one or more agents to make their design choices easier to describe and compare, which turns structural design questions, such as when agents should communicate, gain or lose a capability, or carry information across runs, into testable choices.