本文提出 DiffGate,一种将 GRPO 与选择性、有界教师指导相结合的结果门控目标,在四种模型–领域设置下均提升了 pass@8,表明在该评估协议下解的覆盖度得到改善。DiffGate is introduced, an outcome-gated objective that combines GRPO with selective, bounded teacher guidance, and improves pass@8 across all four model--domain settings, demonstrating improved solution coverage under the evaluation protocol.
论文
1686 张论文卡片 · OA 绿色
场地特异性施肥推荐系统会根据地点、土壤属性、作物类型和生产目标调整养分建议,但当推荐功能主要通过交互式界面访问、输出未做版本管理、训练后的近似模型无法独立加载或基准测试时,其科学复用性受到限制。本技术报告介绍 Turba 施肥机器学习技术栈——面向摩洛哥场地特异性施肥推荐的可复现三层开源实现,其中 turba-client 提供对公开……Site-specific fertilizer recommendation systems adapt nutrient advice to location, soil properties, crop type, and production targets, but scientific reuse is constrained when recommendation functions remain accessible mainly through interactive interfaces, outputs are not versioned, and trained approximations cannot be independently loaded or benchmarked. This technical report presents the Turba fertilizer machine learning stack, a three-layer open-source implementation for reproducible site-specific fertilizer recommendation in Morocco. turba-client provides programmatic access to publicly a
传感器不仅有助于理解世界,还有助于决定下一步动作。然而现有传感器模型大多止于感知层面:识别状态或预测结果,而动作则通过任务特定且通常封闭的标签空间单独建模。本文提出感知-语言-动作(SLA)建模,一个在统一模型中连接多模态传感器观测、自然语言与动作的框架。SLA 将语言用作感知与动作之间的语义接口,使异构动作得以表示、预测与解释……Sensors are useful not only for understanding the world but also for deciding what to do next. Existing sensor models however largely stop at perception: they recognize states or predict outcomes, leaving actions modeled separately through task-specific and often closed label spaces. We introduce Sensor-Language-Action (SLA) modeling, a framework that connects multimodal sensor observations, natural language, and actions within a unified model. SLA uses language as a semantic interface between sensing and acting, allowing heterogeneous actions to be represented, predicted, and explained while
长程 LLM agent 常借助稀疏 mixture-of-experts(MoE)模型实现,但 agentic 行为与 MoE 结构的协同设计仍缺乏充分探索。本文系统研究 agentic 后训练与 MoE 专家选择之间的关联。在现成 MoE 模型中,我们观察到专家选择呈现出与 agentic 轨迹自然对齐的专门结构:agent 执行语义相似操作(如 READ、UPDATE)的回合之间,其专家路由的重叠度高于执行不同操作……Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connections between agentic post-training and MoE expert selection. In off-the-shelf MoE models, we observe expert selection exhibits a specialized structure that naturally aligns with agentic trajectories. Specifically, expert routing overlaps more between turns where the agent performs semantically similar operations (e.g., READ, UPDATE) than between turns with differing oper
自回归视频扩散支持流式生成与交互控制,但其 KV cache 随生成历史持续增长。现有压缩策略要么采用固定窗口丢弃历史,要么依据局部注意力与相似度信号选择 token,均无法直接衡量当前块是否贡献了超出已保留上下文的信息。本文提出 DeCoPrune,一种将 cache 压缩视为去噪一致性问题的免训练方法。经验上发现,去噪难度可作为 token 价值的有用代理……Autoregressive video diffusion supports streaming generation and interactive control, but its KV cache grows continuously with the generated history. Existing compression strategies either discard history using fixed windows or select tokens through local attention and similarity signals, which do not directly measure whether the current chunk contributes information beyond the retained context. We introduce DeCoPrune, a training-free method that treats cache compression as a denoising-consistency problem. We find empirically that denoising difficulty provides a useful proxy for a token's valu
分布匹配蒸馏(DMD)通过分别估计的目标分数与学生分数之差训练少步学生模型,因而必须维护一个拟合学生动态分布的辅助扩散模型,带来额外的显存与计算开销。本文提出 DMAD——分布匹配对抗蒸馏,将分布匹配重构为分类任务,直接学习所需的 log-density 比。共享主干上的两个判别头分别区分真实数据与教师样本是否来自学生模型,并对其 logits 施加线性损失以训练……Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distinguish real data and teacher samples from the student's, and linear losses on their logits train the
知识密集型任务需要通过推理共享的 artifact 语料库(如案件或科学文献)来回答多个问题。随着人类与这些语料库交互,他们自然地积累关于 artifact 的经验知识,从而能够快速识别每个新任务所需的完整相关 artifact 集合。然而,现有 AI agent 缺乏构建或复用此类 artifact-grounded 经验的合适记忆方案,导致答案质量降低且在线成本升高。现有记忆方案从先前任务求解过程中提取并复用信息……(原文截断)Knowledge-intensive tasks require answering many questions by reasoning about a shared corpus of artifacts (e.g., court cases, or scientific literature). As humans interact with these corpora, they naturally accumulate experiential knowledge about artifacts, enabling them to quickly identify the complete set of relevant artifacts for each new task. However, existing AI agents lack appropriate memory solutions to build or reuse such artifact-grounded experience, leading to lower answer quality and higher online cost. Existing memory solutions extract and reuse information from prior task-solvin
大语言模型正日益应用于基于长且异构信息源的任务。传统 RAG 依赖固定的相似度检索,而 agentic 变体虽能自适应查询与工具使用,但仍以检索为中心。然而在许多任务中,解答所需的证据并未显式存在于任何单一来源条目中,而必须通过对多个来源条目进行过滤、聚合或计算来推导。在本工作中,我们提出 RECAST(通过计算、访问与综合进行证据路由)……(原文截断)Large language models are increasingly applied to tasks grounded in long, heterogeneous information sources. Conventional Retrieval-Augmented Generation (RAG) relies on fixed similarity-based retrieval, while agentic variants adapt queries and tool use but remain largely retrieval-centric. However, in many tasks, the evidence required for a solution is not explicitly present in any single source item. Instead, it must be derived through filtering, aggregation, or computation across multiple source items. In this work, we introduce RECAST (Routing Evidence through Computation, Access, and Synth
测试时扩展(TTS)通过分配额外的推理算力来提升大语言模型的推理能力。现有提升 TTS 效率的方法通常一次针对单一资源维度优化准确率,推进 accuracy–cost 或 accuracy–latency 的 Pareto 前沿。然而用户需求是多维的:用户可能同时指定准确率、延迟与推理成本需求,而不同需求可能偏好不同的控制器。我们将个性化测试时扩展形式化为发现可执行控制器,以最大化……(原文截断)Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontier. Yet user requirements are multidimensional: users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controllers. We formulate Personalized Test-Time Scaling as discovering executable controllers that maximize
记忆增强强化学习增强了 LLM agent 解决复杂长程任务的能力。技能是此类记忆的一种形式,将指令与跨任务类型的适用条件配对。然而随着策略改进,若不加区分地保留所有技能,会导致过时或有害条目累积并误导 agent。我们提出 SkillForge,一种 agentic RL 方法,通过由适应性驱动的技能生命周期(试用、活跃、稳定、退役状态)来编译与演化技能库,使技能与模型在整个训练过程中协同演化。预 RL……(原文截断)Memory-augmented reinforcement learning strengthens LLM agents' ability to solve complex long-horizon tasks. Skills are one such form of memory, pairing instructions with an applicability condition over task types. However, retaining every skill indiscriminately as the policy improves lets obsolete or harmful entries accumulate and mislead the agent. We propose SkillForge, an agentic RL method that compiles and evolves the skill library through a fitness-driven skill lifecycle of trial, active, stable, and retired states, so that the skills and the model co-evolve throughout training. A pre-RL
提出 RoboDojo,一个面向通用机器人操作策略综合评估的仿真-真机统一基准,将 30 种策略集成到 XPolicyLab 并在 RoboDojo 上进行评测,建立了公开的排行榜与系统性的策略性能分析。RoboDojo is introduced, a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies that integrates 30 policies into XPolicyLab and evaluates them on RoboDojo, establishing a public leaderboard and systematic analysis of current policy performance.
DeepSeek Engram 等条件记忆架构利用输入 n-gram 查询已学习的 embedding,以有限的额外计算扩展大语言模型(LLM)容量。除模型扩展外,该架构还展现出将事实知识存储与通用计算解耦的潜力,为在固定 Transformer 主干的同时更新事实知识提供了一条可行路径。然而实现这一目标颇具挑战:同一事实的不同表达可能激活不同的 n-gram embedding,而更新共享 embedding 又会……Conditional memory architectures such as DeepSeek Engram use input n-grams to look up learned embeddings, expanding the capacity of large language models (LLMs) with limited additional computation. Beyond model scaling, this architecture has demonstrated the potential to decouple factual knowledge storage from general-purpose computation, offering a promising route to updating factual knowledge while keeping the Transformer backbone fixed. Realizing this potential is challenging because different expressions of a fact may activate different n-gram embeddings, while updating shared embeddings c
Astra 具备行动能力,但可靠的操作取决于其观察和控制世界的系统。本文提出 PhysEvo,一个围绕单个冻结模型构建的物理递归自我改进(RSI)框架。任务 Agent 执行机器人任务,元 Agent 利用产生的轨迹诊断失败、修订工具与技能,并测试修正效果。元 Agent 还能改进自身的诊断工具,使保留的修订同时支持后续行动与后续自我改进。该过程发展出关节级控制、循证观察以及可复用的操作……Astra can act, yet reliable manipulation depends on the system through which it observes and controls the world. We introduce PhysEvo, a framework for physical recursive self-improvement (RSI) around a single frozen model. A task agent executes robot tasks; a meta-agent uses the resulting trajectories to diagnose failures, revise tools and skills, and test corrections. The meta-agent can also improve its own diagnostic tools, so retained revisions support both later action and later self-improvement. This process develops joint-level control, evidence-seeking observation, and reusable manipula
在大规模数据集上训练的视觉-语言-动作(VLA)策略在其训练域内表现良好,但仍难以泛化到真实部署中机器人遭遇的各种场景。Agentic 机器人系统通过视觉-语言模型(VLM)编排器来弥补策略的不足,该编排器学习何时调用策略、如何下达指令、以及何时改用脚本化技能。然而,由于整个系统围绕一个语言可操控性有限的冻结策略构建,编排器只能规避策略的失败却无法真正克服它们。策略成为瓶颈……Vision-language-action (VLA) policies trained on large datasets are capable within their training domains, yet they still fail to generalize to the variety of situations a robot meets in real-world deployment. Agentic robot systems complement the policy with a vision-language model (VLM) orchestrator that learns when to call the policy, how to instruct it, and when to use scripted skills instead. However, because the harness is built around a frozen policy that has limited language steerability, the orchestrator can avoid the policy's failures but never overcome them. The policy becomes the bo
将图像生成模型包裹在 Agent 框架中可有效提升 Text-to-Image 任务性能:该框架能利用记忆、技能、工作流编排、结果验证与迭代优化来持续构建和修订 prompt,从而生成更好的图像。然而这些增益对扩散模型而言是外部的,只有运行完整框架时才能实现。本文提出扩散在线上下文蒸馏(D-OPCD),将 Agent 改进后的 prompt 视为特权上下文,并将编码在 Agent 框架中的知识蒸馏进……Wrapping an image generation model in an agentic harness can effectively boost Text-to-Image task performance: the harness can leverage memory, skills, workflow orchestration, result verification, and iterative refinement to continually construct and revise prompts, thereby eliciting better images. These gains, however, remain external to the diffusion model and are realized only while the full harness runs. We propose Diffusion On-Policy Context Distillation (D-OPCD), which treats the agent-improved prompt as privileged context and distills the knowledge encoded in the agent harness into the
大语言模型 Agent 可通过保留从先前交互中提炼的可复用技能来跨任务提升能力。已有研究联合优化任务执行与技能提取,使策略与技能库协同进化。然而,随着执行者持续学习,通过技能在后续训练步骤中的复用对其进行奖励,可能将技能收益与执行者自身的改进混淆;而直接测试每个候选技能又需要代价高昂的额外执行者 rollout。本文提出 UniSkill,利用共享策略与环境交互并提出技能……Large language model agents can improve across tasks by retaining reusable skills distilled from prior interactions. Recent work jointly optimizes task execution and skill extraction, enabling the policy and skillbank to co-evolve. However, as the actor continues learning, rewarding skill proposals through their reuse in subsequent training steps may conflate skill benefits with actor improvement, while directly testing each proposed skill requires costly additional actor rollouts. In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose ski
科学 Agent 的核心挑战在于将分析经验转化为基于物理证据的可复用专业知识。本文介绍 Gan Jiang,一个基于我们自主开发的衍射分析生态系统(XMatcher、XQueryer、XDecomposer 与 WPEM)构建的面向粉末 X 射线衍射的自学习 Agent。这些引擎共同涵盖物相鉴定、多相分解以及物理约束的全谱建模。Gan Jiang 通过诊断失败、修订技能指令与代码,并在复用前验证修订,将分析经验转化为可执行的技能,伴随……A central challenge for scientific agents is to turn analytical experience into reusable expertise grounded in physical evidence. Here we introduce Gan Jiang, a self-learning agent for powder X-ray diffraction built on a diffraction-analysis ecosystem we developed: XMatcher, XQueryer, XDecomposer and WPEM. Together, these engines span phase identification, multiphase decomposition and physics-constrained whole-pattern modelling. Gan Jiang converts analytical experience into executable skills by diagnosing failures, revising skill instructions and code, and validating revisions before reuse, wi
从三维形状重建可编辑 CAD 模型仍是一项具有挑战性的工程任务。现有方法可以提出 CAD 操作,但没有任何单一提案源能对不同零件几何与不同重建阶段同样有效。本文介绍 CADFather,一个自主 Agent 系统,通过协调互补工具从三维网格恢复参数化 CAD 程序。视觉-语言助手检查目标与中间重建结果的渲染图像,再决定扩展哪些候选 CAD 程序、调用哪些工具、生成多少提案,以及……Reconstructing an editable CAD model from a 3D shape remains a challenging engineering task. Existing methods can propose CAD operations, but no single source of proposals works equally well across different part geometries and stages of reconstruction. We introduce CADFather, an autonomous agentic system that coordinates complementary tools to recover parametric CAD programs from 3D meshes. A vision-language assistant inspects renders of the target and intermediate reconstructions, then decides which candidate CAD programs to extend, which tools to invoke, how many proposals to generate, and
将音乐转录为人类可读的乐谱需要对节奏、和声、旋律与曲式的整体理解。两大障碍限制了这一目标的实现:标注录音稀缺,以及精确的局部预测仍可能产生不一致的音乐序列。本文提出 SheetSage2,一个统一的音乐转录框架,结合合成数据、任务特定的结构化解码与自回归蒸馏。自动标注的 MIDI 渲染为音频后,可为各类音乐理解任务提供可扩展的监督。任务特定的结构化解码器整合互补的音乐……Transcribing music into a human-readable score requires a coherent understanding of rhythm, harmony, melody, and form. Two obstacles limit this goal: annotated recordings are scarce, and accurate local predictions can still produce inconsistent musical sequences. We present SheetSage2, a unified music transcription framework that combines synthetic data, task-specific structured decoding, and autoregressive distillation. Automatically annotated MIDI, rendered into audio, provides scalable supervision across music understanding tasks. Task-specific structured decoders integrate complementary mu
稀疏视图新视角合成是三维内容创作中的核心问题,但基于扩散的方法受迭代去噪限制,多视图生成在推理时开销高昂。本文提出 NAMVIS,一种无扩散框架,将多视图图像合成重新表述为几何条件下的下一尺度自回归过程。NAMVIS 不通过反复去噪生成目标视图,而是通过少量由粗到细的尺度步预测离散视觉 token,并在同一尺度内以及跨目标视图间并行采样所有 token。为锚定此过程……Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time. We introduce NAMVIS, a diffusion-free framework that reformulates multi-view image synthesis as geometry-conditioned next-scale autoregression. Instead of generating target views through repeated denoising, NAMVIS predicts discrete visual tokens through a small number of coarse-to-fine scale steps, while sampling all tokens within each scale and across target views in parallel. To anchor thi
视觉-语言-动作(VLA)基础模型规模迅速扩大以提升操作性能与泛化能力,但这种规模化带来了高昂的计算成本,使真实世界部署日益困难。现有方法通常通过设计更小的架构或减少基于流(flow-based)策略中的迭代去噪步数来缓解该问题。本文提出 FastOPD,一个从基础到轻量的 VLA 框架,通过高效的在线蒸馏实现大规模 VLA 的实际部署。具体而言,FastOPD 适配流映射(flow map)……Vision-Language-Action (VLA) foundation models have scaled rapidly to enhance manipulation performance and generalizability, but this scaling incurs high computational costs that render real-world deployment increasingly challenging. Existing approaches typically mitigate this issue by designing smaller architectures or reducing the iterative denoising steps in flow-based policies. In this work, we propose FastOPD, a foundation-to-lightweight VLA framework that enables the practical deployment of large-scale VLAs through efficient on-policy distillation. Specifically, FastOPD adapts a flow map
WildCity 是一个由自动驾驶车队在复杂城市环境中采集的真实多模态数据集,旨在推动城市级渲染的进展,并更广泛地推动 AI 在空间感知、记忆与推理方面达到与人类认知相当规模的能力。WildCity, a real-world multimodal dataset collected by autonomous fleets traversing complex urban environments, aims to catalyze progress not only in city-scale rendering, but more broadly in the pursuit of AI that can perceive, remember, and reason across space at a scale comparable to human cognition.
检索增强生成系统日益依赖文档结构处理:结构对齐分块、LLM 生成的块上下文、标题路径元数据,以及分层两阶段检索。各自研究在不同语料、嵌入器和指标上支持各自方法,但没有一项控制了共同混杂因素:在块前添加任何文本都会扰动其嵌入。我们提出机制隔离的消融方法,在同一协议下测试全部四种处理,跨条件匹配块大小,并加入一个语义为空的安慰剂——结构上有效但被打乱Retrieval-augmented generation systems increasingly rely on document-structure treatments: structure-aligned chunking, LLM-generated chunk contexts, heading-path metadata, and hierarchical two-stage retrieval. Separate studies support each on different corpora, embedders, and metrics, and none control for a shared confound: any text prepended to a chunk perturbs its embedding. We present a mechanism-isolating ablation testing all four treatments under one protocol, matching chunk sizes across conditions and adding a semantically null placebo---heading paths that are structurally valid but shuf
大语言模型只能使用其上下文窗口中容纳的文本,并且每次发送 prompt 时都会重新计算其内部的 key-value (KV) 状态。我们测试了一个记忆层,即公开包 galahad-kv,它将每个约 16000 token 块的 KV 状态保存到加密的本地 NVMe 磁盘,并在之后按字节精确、无重计算地加载回来。我们在 5000 万 token 的真实公共文本上运行,通过单块 NVIDIA H100 上的 vLLM 提供服务,配合 Gemma 4 12B 和 Gemma 4 31B。我们探测的每个块都从加密存储中无重计算地加载回来(100 个中的 100 个A large language model can only use the text that fits in its context window, and it recomputes its internal key-value (KV) state for a prompt every time the prompt is sent. We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it. We ran it on 50,000,000 tokens of real public text, served through vLLM on one NVIDIA H100, with Gemma 4 12B and Gemma 4 31B. Every block we probed was loaded back from the encrypted store with no recompute (100 of 100,
潜变量世界模型在预测未来状态和在真实世界中规划方面表现出色。然而在实践中,我们缺乏一种原则性的方法来估计其能力如何随模型规模、数据和算力扩展,这一开放问题减缓了该领域进展。本工作提出 RoboJEPA,一种基于联合嵌入预测架构 (JEPA) 的世界模型,在涵盖 12 种机器人形态的大规模数据集上训练。我们表明 RoboJEPA 的想象误差——其潜变量推演的误差——遵循关于算力的二阶幂律,使我们能够预先Latent world models have shown a remarkable ability to predict future states and to plan in the real world. In practice, however, we lack a principled way to estimate how their capabilities scale with model size, data, and compute, an open problem that slows progress in the field. In this work we present RoboJEPA, a world model based on the Joint Embedding Predictive Architecture (JEPA) and trained on a large-scale dataset spanning 12 robotic embodiments. We show that RoboJEPA's imagination error, the error of its latent rollouts, follows a second-order power law in compute, allowing us to pre
可验证奖励的强化学习 (RLVR) 主要通过交互后的标量结果奖励将 Agent 经验转化为学习信号。然而对于组相对目标,当所有推演获得相同奖励时,该信号即消失,尽管这些轨迹可能包含关于任务内容及 Agent 失败方式的有用信息。我们提出一个互补问题:事后回顾能否教会 Agent 在行动前本可预见的内容?我们引入前瞻学习,利用事后经验从行动前的Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails. We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting? We introduce prospective learning, which uses post-hoc experience to supervise foresight predictions from the pre-
AI Agent 日益在能够诊断失败并通过经验改进的环境中运行,然而现有评估主要衡量 Agent 在固定时间点能做什么,而非其学习效果如何。评估自我改进需要回答三个问题:未来性能是否改进并泛化至学习交互之外;新能力获取效率如何;自我改进过程在哪里失效?为回答这些问题,我们在受控环境中研究自我改进,其中 Agent 分摊AI agents increasingly operate in environments where they can diagnose failures and improve through experience, yet existing evaluations largely measure what an agent can do at a fixed point in time rather than how effectively it learns. Evaluating self-improvement requires answering three questions: does future performance improve and generalize beyond the interactions that enabled learning; how efficiently are new capabilities acquired; and where does the self-improvement process break down? To answer these questions, we study self-improvement in a controlled setting where agents amortize pa
由大语言模型构建的多 Agent 系统 (MAS) 通过协调专业化 Agent 来处理复杂任务,但有效的工作流难以预先设计。测试时演化利用执行反馈来优化工作流,然而广泛的修订可能扰动有用组件,而重执行未变更的请求会带来冗余计算。受生物进化中继承与选择相互作用的启发,我们引入 Inherit-MAS,在工作流和执行层面显式化继承。元模型首先综合出由 worker Agent 组成的工作流Multi-agent systems (MAS) built from large language models coordinate specialized agents to tackle complex tasks, but effective workflows are difficult to design in advance. Test-time evolution refines workflows using execution feedback, yet broad revisions can disturb useful components, while re-executing unchanged requests can incur redundant computation. Inspired by the interplay of inheritance and selection in biological evolution, we introduce Inherit-MAS, which makes inheritance explicit at the workflow and execution levels. A meta-model first synthesizes a workflow of worker agents with
恶意软件防御通常会尽快移除或隔离可疑程序。该策略虽便于遏制,却也浪费了观察攻击者行为与部署针对性反制的机会。ORCAGen 另辟蹊径:离线利用 GenAI 构建针对特定恶意软件的欺骗剧本,部署前进行校验,运行时仅强制执行已验证的逻辑。ORCAGen 将 RAG 与结构化提示工程相结合,以同时生成 PoC 恶意软件及对应的欺骗编排代码。Malware defenses often remove or isolate suspicious programs as quickly as possible. While effective for containment, this approach can also waste an opportunity to observe attacker behavior and deploy targeted countermeasures. ORCAGen takes a different approach: it uses GenAI to build malware-specific deception playbooks offline, validates them before deployment, and enforces only the verified logic at runtime. ORCAGen combines Retrieval-Augmented Generation (RAG) with structured prompt engineering to generate both proof-of-concept (PoC) malware and corresponding deception orchestration code.
大规模信息检索系统(包括 RAG 与推荐引擎)广泛采用多层分层数据结构,以在高维向量空间中实现超高速近似最近邻搜索。然而,确保贪心导航既精确又高效的几何条件仍未被充分理解。本文研究由 d 维环面 T^d 上 n 个数据点构建的邻近图分层结构上的贪心导航效率,并确定了一个确定性覆盖条件,在该条件下……Large-scale information retrieval systems, including retrieval-augmented generation (RAG) and recommendation engines, widely use multi-layered hierarchical data structures for ultra-fast approximate nearest-neighbor search in high-dimensional vector spaces. However, the geometric conditions that ensure accurate and efficient greedy navigation remain poorly understood. In this work, we study the efficiency of greedy navigation on a hierarchy of proximity graphs constructed from \(n\) data points on the \(d\)-dimensional torus~$\mathbb{T}^d$. We identify a deterministic coverage condition under
LLM 可在前缀式提取下泄露记忆化的训练序列:给定训练样本的前缀,模型可能为原始续写赋予高概率。但在部署系统中,前缀很少被单独评估,常与指令、检索文档或其他任务相关上下文一同出现,正如 RAG 所做的那样。这促使我们去考察:上下文条件化究竟是缓解了记忆化,还是仅仅改变了可被提取的记忆化样本集合。本文对这一问题展开研究……Large language models (LLMs) can expose memorized training sequences under prefix-based extraction: given a prefix from a training example, the model may assign high probability to the original continuation. In deployed systems, however, prefixes are rarely evaluated in isolation. They often appear together with instructions, retrieved documents, or other task-specific context, as in retrieval-augmented generation (RAG). This motivates examining whether contextual conditioning mitigates memorization or merely changes the set of memorized samples that become extractable. We investigate this iss
LLM 越来越多地作为组件嵌入软件系统,并以 chatbot、copilot、RAG、workflow、coding agent、AI agent 等标签推广。这些标签究竟指代真实的架构形态,还是仅作品牌包装,尚未得到系统性评估。在调研的来源中,标签确实承载架构含义,在厂商用法中体现得最为清晰:copilot 指在逐步用户确认下操控宿主应用的 router-worker 架构,而近来向 agent 标签的迁移则与 AI 自主规划相吻合……Large language models (LLMs) are increasingly embedded as components in software systems, marketed under labels such as chatbot, copilot, retrieval-augmented generation, workflow, coding agent and AI agent. Whether these labels denote genuine architectural forms or serve as branding has not been assessed systematically. In the sources surveyed, labels do carry architectural content, most clearly in vendor usage: copilot denotes a router-worker architecture operating a host application under step-by-step user confirmation, while the more recent shift to the label agent coincides with AI-planned
得益于涵盖全身自由度的扩展预训练数据,LingBot-VLA-2.0 在两个机器人平台上展现出强大的跨具身长时程移动操作能力。Benefiting from the expanded pretraining data that covers whole-body degrees of freedom, LingBot-VLA-2.0 demonstrates strong cross-embodiment long-horizon mobile manipulation capability across the two robotic platforms.
强化学习(RL)是推动大型基础模型走向自我改进的核心训练范式。本报告介绍 MiMo-V2.6 系列,这是一个通过扩展 RL 算力来推动模型智能边界的 omni-modal 模型族。在 RL 之前,我们在广泛的跨模态语料上进行 mid-training 以提供充足的探索空间,并基于预训练的 hybrid-SWA 架构构建稳健的基础设施以支持后续规模化。我们沿三个维度扩展 RL 算力:(1) 更大的 batch 和更高的吞吐量,采用异步训练来持续消费Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement. This report introduces the MiMo-V2.6 series, an omni-modal family that pushes the frontier of model intelligence by scaling RL compute. Prior to RL, we conduct mid-training on a broad multimodal corpus to provide ample exploration space, and build a solid infrastructure on the pretrained hybrid-SWA architecture to support subsequent scale-up. We scale RL compute along three dimensions: (1) larger batches and higher throughput, with an asynchronous training that consumes
现有的 3D 世界模型能够生成逼真、可探索的场景,但场景在时间维度上静止。OuroWorld 是一个无需掩码(mask-free)的框架,可将任意静态 3D Gaussian Splatting 场景转化为 3D Cinemagraph:从任意视角都能无缝循环、拥有生动多样运动的动态场景。视觉语言模型(vision-language model)推断合理的动态并引导视频模型合成参考视频,再将其提升并补全为多视图视频。为从这种不完美的监督中学习,我们提出 Inconsistency-Robust Periodic 4DGS:通过傅里叶级数形变场从构造上保证循环Recent 3D world models generate photorealistic, explorable scenes that remain frozen in time. OuroWorld is a mask-free framework that turns any static 3D Gaussian Splatting scene into a 3D cinemagraph: a dynamic scene with vivid, diverse motion looping seamlessly from any viewpoint. A vision-language model infers plausible dynamics and guides a video model to synthesize a reference video, which we lift and complete into multi-view videos. To learn from this imperfect supervision, we propose Inconsistency-Robust Periodic 4DGS: a Fourier-series deformation field guarantees looping by constructio
这是一本关于大语言模型的书。正如标题所示,本书主要聚焦于基础概念,而非全面覆盖所有前沿技术。全书共分六个主要章节,每章探讨一个关键领域:预训练、生成模型、提示工程、对齐、推理和推理能力。本书面向自然语言处理及相关领域的大学生、从业者和研究者,也可作为任何对大语言模型感兴趣的人的参考。This is a book about large language models. As indicated by the title, it primarily focuses on foundational concepts rather than comprehensive coverage of all cutting-edge technologies. The book is structured into six main chapters, each exploring a key area: pre-training, generative models, prompting, alignment, inference, and reasoning. It is intended for college students, professionals, and practitioners in natural language processing and related fields, and can serve as a reference for anyone interested in large language models.