研究库 论文知识库
Papers · organized/paper_cards

论文

1640 张论文卡片 · OA 绿色

开放获取 全部 绿色 · 1640
From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix
从生产流量到后训练:构建覆盖企业请求组合的自托管 LLM
arXiv:2609.01572 工程化 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

数据驻留约束迫使企业自托管 LLM,但不断引入新模型而不下线旧模型会扩张服务集群,分散有限的 GPU 池。我们通过沿指令遵循、函数调用和内部任务分布三个维度,针对生产错误分析所发现的质量差距,将 200 多个内部应用的流量整合到单一模型上。质量通过按生产流量分层的离线基准进行跟踪,并由确定性验证器或经过校准的 LLM 评判器打分。不同于针对Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decommissioning their predecessors expands the serving fleet, fragmenting a finite GPU pool. We consolidate traffic from over 200 internal applications onto a single model by closing quality gaps identified through production error analysis along three axes: instruction following, function-calling, and internal task distribution. Quality is tracked by offline benchmarks stratified to production traffic and scored by deterministic verifiers or calibrated LLM judges. Rather than optimi

Recursive Criticality of AI Self-Improvement
AI 自我改进的递归临界性
arXiv:2609.00137 安全与风险 观点 OA · 绿色 被引 0 · S2 + OpenAlex

该框架识别了 AI R&D 系统的若干可测量属性,可用于区分递归放大与其他来源驱动的快速进展,包括递归反馈的强度、改进向后续系统传播的有效性、周期时长,以及进一步取得进展的难度递增。The framework identifies measurable properties of AI R\&D systems that can help distinguish recursive amplification from rapid progress driven by other sources, including the strength of recursive feedback, how effectively improvements propagate into successor systems, cycle duration, and the increasing difficulty of further progress.

Agent Memory Is a Surface for Endogenous Authorization Laundering
Agent 内存是内生授权洗钱的表层载体
arXiv:2609.01836 Agent 智能体 方法 OA · 绿色 被引 3 · S2

本工作在采购、网络安全与金融场景下评估了 5 个 LLM 作为记忆写入者、2 个 LLM 作为执行者,并提出 EAL-Bench,用于衡量持久记忆在多大程度上准确保留不断演化的授权状态,以及错误是否会向下游传播为未授权操作。This work evaluates five LLMs as memory writers and two as executors across procurement, cybersecurity, and finance and introduces EAL-Bench, which measures how accurately persistent memory preserves evolving authorization state and whether errors propagate to downstream unauthorized actions.

DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation
DramaChain Bench:面向短剧生成的端到端基准
arXiv:2609.00646 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 DramaChain Bench,首个覆盖完整生产链各阶段的短剧基准,并验证最终剧集质量并非仅由视频生成决定。DramaChain Bench is presented, the first short-drama benchmark that evaluates every stage of the complete production chain, and confirms that final episode quality is not governed by video generation alone.

Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds
模仿学习是否保留灵巧操作的时间鲁棒性?跨任务执行速度的专家-学习者对比
arXiv:2609.01453 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

任务名义成功率相同并不意味着跨执行速度保留专家性能;本文在相同任务条件、初始条件采样与加速倍率下对比了专家与学习者。Equal nominal task success does not imply preservation of expert performance across execution speeds, and an expert and learner under the same task conditions, initial-condition draws, and speedup factors is compared.

6. LLM Research Papers: The 2026 List (Jan–May) — Sebastian Raschka
6. LLM Research Papers:2026 清单(1月–5月) — Sebastian Raschka
arXiv:2604.12374 LLM 基础设施 方法 OA · 绿色 被引 37 · S2

Nemotron 3 Super 是 Nemotron 3 系列中首个采用 NVFP4 进行预训练的模型,借助 LatentMoE(一种同时优化精度 per FLOP 与精度 per parameter 的新型 Mixture-of-Experts 架构),并集成 MTP 层以通过原生 speculative decoding 加速推理。Nemotron 3 Super is the first model in the Nemotron 3 family to be pre-trained in NVFP4, leverage LatentMoE, a new Mixture-of-Experts architecture that optimizes for both accuracy per FLOP and accuracy per parameter, and include MTP layers for inference acceleration through native speculative decoding.

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
中期训练阶段的知识蒸馏更偏向推理而非事实记忆
arXiv:2609.01532 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Switch Distillation,一种简单的 mid-training 目标:以教师预测熵作为轻量路由信号,仅在教师置信的 token 上蒸馏,其余回退到交叉熵;在不同教师规模下均稳定优于现有蒸馏目标。Switch Distillation is proposed, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy, which consistently outperforms existing distillation objectives across teacher sizes.

Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data
通过对非结构化数据的自适应结构化实现 token 高效的数据推理 Agent
arXiv:2608.31082 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 agentic data cracking,将非结构化数据自适应、推测性地结构化为推理自身的副产品,是面向非结构化数据上 Agentic reasoning 的下一代数据基础设施的第一步。This work proposes agentic data cracking, a method that structures unstructured data adaptively and speculatively as a byproduct of reasoning itself, a first step toward next-generation data infrastructure for agentic reasoning over unstructured data.

Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry
学习结果变化的位置:面向多模态几何的可寻址信用推理
arXiv:2608.30457 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 credit-addressable reasoning:推理时暴露的语义单元同时定义学习阶段比较候选与分配 credit 的位置;并实例化为 Code-CoT,保留图示、将视觉关系表示为行可寻址的可执行代码,并将推理组织为类型化事件。This work introduces credit-addressable reasoning, in which the semantic units exposed during inference also define where learning compares alternatives and assigns credit, and instantiates Code-CoT, which retains the diagram, represents visual relations as line-addressable executable code, and organizes reasoning into typed events.

AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
AgentJudgeBench:一个用于评估 LLM 法官在智能体工具调用方面表现的多难度基准
arXiv:2608.26623 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

首个系统性研究 LLM-as-a-judge 在工作流 DAG 上对 Agentic 工具调用评判可靠性的基准,区别于面向开放式文本或偏好的通用 LLM-as-a-judge 任务;揭示了当前 LLM judge 的根本局限,并给出面向 Agentic 系统可靠评估的实践指南。The first benchmark to systematically study LLM-as-a-judge reliability for agentic tool-calling over workflow DAGs, as distinct from the broader LLM-as-a-judge task of open-ended text or preference evaluation, exposes fundamental limitations of current LLM judges and yields practical guidelines for reliable evaluation in agentic systems.

VibeVoice-ASR-Streaming Technical Report
VibeVoice-ASR-Streaming 技术报告
arXiv:2609.02812 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

VibeVoice-ASR-Streaming 是首批基于 LLM 的端到端流式说话人归属 ASR 方法之一,无需独立 diarization 阶段即可在语音到达时输出"谁说了什么"。VibeVoice-ASR-Streaming is one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR, allowing the model to produce''who said what''as speech arrives, without a separate diarization stage.

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
HarnessDev:LLM 能创建并演进自身的 Agent Harness 吗?
arXiv:2609.01437 评测基准 评测集 OA · 绿色 被引 5 · S2

研究发现,生成的 harness 在代码与搜索/研究任务上仍显著落后于成熟的人工参考方案,而在写作与机器学习实验任务上达到或超过所选参考方案,且执行成本差异巨大。It is found that generated harnesses remain substantially behind mature human-engineered references on code and on search and research, while matching or exceeding the selected references on writing and machine-learning experimentation, with large variation in execution cost.

CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing
CRISP:面向结构-质量驱动的悬崖感知输入自适应稀疏预填充
arXiv:2609.01925 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文用 C_struct 替代 Jensen-Shannon Divergence 路由:C_struct 是一种结构化代理,通过度量 Vertical-Slash 兼容位置上的 mass 来复现 JSD 的路由决策,同时消除池化 matmul 与后续 KL 散度开销。This work replaces the Jensen-Shannon Divergence routing with C_struct, a structural proxy that measures mass at Vertical-Slash compatible positions and reproduces JSD's routing decisions while eliminating both the pooled matmul and subsequent KL divergence overhead.

A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss
一瞥即可:基于 SimLoss 的单遍细粒度图像描述生成
arXiv:2609.00591 RAG 检索增强 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 SimLoss,一种面向单轮细粒度图像描述的无参考 embedding 空间目标;结果表明 embedding 空间监督能在单轮描述器延迟下恢复多阶段验证的质量。SimLoss is proposed, a reference-free embedding-space objective for single-pass fine-grained image captioning, and results show that embedding-space supervision can recover the quality of multi-stage verification at the latency of a single-pass captioner.

Aspire: Can Models Self-Evolve from Vague Goals?
Aspire:模型能否从模糊目标中自我演进?
arXiv:2608.31111 评测基准 评测集 OA · 绿色 被引 1 · S2

本文提出 ASPIRE,面向模糊目标驱动自我演化的基准,表明模糊目标会将搜索资源导向目标解释阶段;并在涵盖 6 类目标的 520 题隐藏专家评测集上评估所得系统。This work introduces ASPIRE, a benchmark for vague-goal-driven self-evolution and shows that vague goals redirect search effort toward goal interpretation, and evaluates the resulting systems on a hidden, expert-authored set of 520 items spanning six goals.

5️⃣ arXiv · Is Agentic RAG Worth It? An Experimental Comparison of RAG Approaches(⭐⭐⭐⭐ 高优先级)
5️⃣ arXiv · Agentic RAG 是否值得?RAG 方法的实验对比(⭐⭐⭐⭐ 高优先级)
arXiv:2601.07711 RAG 检索增强 评测集 OA · 绿色 被引 6 · S2

基于实证对 "Enhanced" 与 "Agentic" RAG 范式进行评估,为真实场景中选取最有效的 RAG 设计(兼顾性能与成本)提供指导。An empirically driven evaluation of the "Enhanced" and "Agentic" RAG paradigms is conducted, offering guidance on selecting the most effective RAG design for real-world applications, considering both performance and costs.

S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
S3Gym:LLM 能将自测试与自评判转化为自我提升吗?
arXiv:2608.31100 评测基准 评测集 OA · 绿色 被引 1 · S2

这些发现表明,仅识别成功动作并不足够;Agent 还需将反馈转化为可执行且可迁移的策略。本文给出统一框架以诊断该过程,并定位阻碍 Agent 将交互经验转化为可靠自我提升的瓶颈。These findings show that recognizing successful actions is insufficient; agents must also transform feedback into executable and transferable policies, and provide a unified framework for diagnosing this process and identifying the bottlenecks that prevent agents from translating interaction experience into reliable self-improvement.

SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions
SnapBench:面向移动端交互的即拍即问多模态检索 benchmark
arXiv:2608.29607 RAG 检索增强 评测集 OA · 绿色 被引 1 · S2

本文提出 SnapBench——首个面向鲁棒"拍照即问"多模态检索的配对基准,以及一种简单的自适应融合方法 MOOR(Modality-anchored, Outlier-aware, Optimal Reweighting),并指出在"拍照即问"检索中需要具备可靠性感知的模态校准。SnapBench is introduced, the first paired benchmark for robust snap-and-ask multimodal retrieval, and MOOR (Modality-anchored, Outlier-aware, Optimal Reweighting), a simple adaptive fusion approach, highlighting the need for reliability-aware modality calibration in snap-and-ask retrieval.

Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers
机构报纸流水线:从历史报纸中提炼数十亿高质量 token
arXiv:2608.18972 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Institutional Newspapers Pipeline,一个模块化系统,旨在从历史报纸扫描件中提取高质量、结构化的数据集;其架构设计使每个步骤都保持可解释和可定制,并使整个 pipeline 在计算上足够精简,可在工作站级硬件上运行。The Institutional Newspapers Pipeline is presented, a modular system designed to extract high-quality, structured datasets from historical newspaper scans that was architected so that each step remains interpretable and customizable, and so that the pipeline as a whole remains computationally frugal enough to run on workstation-level hardware.

ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering
ViSAR:面向视觉文档问答的无训练自适应 $k$ 检索
arXiv:2609.02486 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ViSAR(Visual Semantic Activation Retrieval),一种面向 late-interaction 视觉文档检索的无训练自适应 k 检索方法,并表明相似度矩阵结构与答案准确率相关,为面向检索质量感知的文档理解指明了未来方向。ViSAR (Visual Semantic Activation Retrieval), a training-free adaptive-$k$ retrieval method for late-interaction visual document retrieval, is introduced and it is shown that the similarity matrix structure correlates with answer accuracy, suggesting future directions for retrieval quality-aware document understanding.

Improving Health Literacy through Lay Summarization of Radiological Reports: An Evaluation of BioNER and Retrieval-Augmented Generation
通过放射学报告的大众化摘要提升健康素养:BioNER 与检索增强生成的评估
arXiv:2609.02396 RAG 检索增强 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本研究探讨了相较于标准 LLM 生成方式,RAG 与命名实体识别在多大程度上能提升自动生成通俗摘要的质量、事实一致性及可读性。This study investigates the extent to which Retrieval-Augmented Generation and Named Entity Recognition improve the quality, factual consistency, and readability of automatically generated lay summaries compared with standard LLM-based generation.

NE-R1: Enhancing Named Entity Recognition Model via Reinforcement Learning
NE-R1:通过强化学习增强命名实体识别模型
arXiv:2609.02366 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 NE-R1,一种面向自适应检索增强 NER 的新框架,在多个基准上达到 SOTA 性能,域内评估平均 F1 提升 2.52%,零样本跨域评估平均 F1 提升 1.18%。This paper proposes NE-R1, a novel framework for adaptive retrieval-augmented NER, which achieves state-of-the-art performance on various benchmarks, with an average F1 score gain of 2.52% in in-domain evaluation and 1.18% in zero-shot cross-domain evaluation.

Exploring Collaboration between a language and a non-language agent
探索语言智能体与非语言智能体之间的协作
arXiv:2609.00474 Agent 智能体 应用落地 OA · 绿色 被引 1 · S2

为解决 LLM 与非语言 Agent 的协作问题,本文提出 latent state internalization,将子 Agent 的连续表示直接投射到 LLM 的 token 流中作为习得的状态 token,并随着动作推进环境状态而进行动态重编码。To solve LLM collaboration with non-language agents, latent state internalization is introduced, which projects the subagent's continuous representations directly into the LLM's token stream as learned state tokens, with dynamic re-encoding as actions advance the environment state.

NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference
NeoMME:用于高效微调与推理的单塔多模态原生多语言基础编码器
arXiv:2609.01657 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 NeoMME,一个 260M 与 800M 参数的多模态多语言双向编码器系列,可在单个双向 Transformer encoder 中处理多语言文本与原始图像 patch。This work introduces NeoMME, a family of 260M and 800M-parameter Multimodal and Multilingual bidirectional Encoders that process multilingual text and raw image patches in a single bidirectional Transformer encoder.

FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos
FoldingAgent:从演示视频中推断参数化折纸过程
arXiv:2609.00377 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 FoldingAgent,一个从折纸演示视频中推断显式参数化折纸程序的 Agent 框架,利用预训练 Vision-Language Model 的推理能力,并配备可模拟几何变换、验证物理合理性、检索与比较视觉内容以及评估自身预测的专用工具集。FoldingAgent is presented, an agentic framework for inferring explicit parametric folding programs directly from origami demonstration videos that leverages the reasoning power of a pre-trained Vision-Language Model equipped with a suite of specialized tools that enable the agent to simulate geometric transitions, verify physical plausibility, retrieve and compare visual content, and evaluate its own predictions.

Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents
无知还是无能?为 LLM 智能体构建知识门控的可验证任务
arXiv:2608.30322 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种知识门控的任务构建协议,将任务指令与一个包含私有约定、参考表与效用算子的紧凑工件分离,并证明被保留的任务能够改善 post-training。A knowledge-gated task-construction protocol is introduced that separates a task instruction from a compact artefact containing private conventions, reference tables, and utility operators, and it is shown that the retained tasks improve post-training.

5. GraphRAG / LLMs+Graphs 综合研究
arXiv:2606.11560 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本教程综合了推动这些汇聚方向的算法、系统与设计原则,为数据科学与数据挖掘研究者提供统一视角,涵盖将 LLM、图数据管理、图挖掘、图 ML 与 agentic 计算融合到下一代 graph-native AI 系统中。This tutorial synthesizes the algorithms, systems, and design principles driving these converging directions, offering data science and data mining researchers a unified perspective on integrating LLMs, graph data management, graph mining, graph ML, and agentic computation into next-generation graph-native AI systems.

Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering
超越视觉相似性:面向知识库视觉问答的实体对齐检索
arXiv:2608.21450 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 KBMR,首个面向 KB-VQA 的基于 MLLM 的 embedding retriever,并引入一个基于 MLLM 的语义判别器以生成连续的实体一致性权重,应对维基百科规模检索中的噪声监督挑战。KBMR is proposed, the first MLLM-based embedding retriever tailored for KB-VQA, and an MLLM-based semantic discriminator that generates continuous entity-consistency weights is introduced to tackle the challenge of noisy supervision in Wikipedia-scale retrieval.

Debias-SparseGPT: Bias-Aware Pruning for Large Language Models
Debias-SparseGPT:面向 LLM 的偏置感知剪枝
arXiv:2609.02496 工程化 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Debias-SparseGPT,一种 post-training 剪枝方法,通过在人口统计对比输入上定义的二阶项引入表示去偏好,在保持模型困惑度与零样本准确率的前提下,一致地降低剪枝带来的偏差,效果优于 SparseGPT。Debias-SparseGPT is introduced, a post-training pruning method incorporating representational debiasing using a second-order term defined over demographically contrasting inputs that consistently reduces pruning-induced bias compared to SparseGPT while preserving model perplexity and zero-shot accuracy.

Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens
Sparse Readout Prism:用特征而非 token 解释 Logit-Lens 分数
arXiv:2609.01936 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Sparse Readout Prism (SRP) 仅使用 readout 的权重对其进行分解,将任意 token logit 或 logit 差表示为来自稀疏 readout 特征贡献之和,揭示了 readout 特征作为 lens 解读新单元的价值,暴露出 token 身份可能掩盖的结构,并支持跨 token、上下文、层与 lens 的比较。Sarse Readout Prism (SRP), which decomposes the readout using only its weights and expresses any token logit or logit difference as a sum of contributions from sparse readout features, reveals readout features as a new unit of analysis for lens readings, exposing structure that token identities can obscure and enabling comparisons across tokens, contexts, layers, and lenses.

WHALE: A Simple Recipe for Joint Harness-Weight Optimization
WHALE:联合 Harness-权重优化的简洁方案
arXiv:2609.00196 评测基准 方法 OA · 绿色 被引 8 · S2

本文提出 Weight-Harness Alternating LEarning (WHALE),一种简单的方法,交替进行两个阶段:先在当前 harness 下更新模型,再在更新后的模型下通过在线拒绝采样微调与 Meta-Harness 搜索更优的 harness。Weight-Harness Alternating LEarning (WHALE), a simple recipe that alternates two phases: updating the model under the current harness, then searching for a better harness under the updated model with online rejection-sampling fine-tuning and Meta-Harness, is proposed.

Small Language Models as Judges for Rubric-Based Reinforcement Learning
小语言模型作为基于评分标准的强化学习评判器
arXiv:2608.30005 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文研究较小的语言模型能否作为高效且可靠的 rubric 评分器,并比较了从小模型中提取逐项判断的三种方式:生成式判定、Yes/No Logprob 边际以及探针判别器。This work studies whether smaller language models can serve as efficient and reliable rubric-based judges, and compares three ways of extracting criterion-level judgments from small models: Generative verdicts, Yes/No Logprob margins, and Probe judges.

An Empirical Study on Zero-Data Bootstrapping for Conversational Recommender Systems
面向对话推荐系统的零数据引导的实证研究
arXiv:2504.15476 多模态 方法 OA · 绿色 被引 7 · S2

结果表明:领域驱动的合成数据一致优于零样本提示与朴素合成基线;主动选择相比随机采样提升了数据效率;元数据与协同过滤信号各自提升选择质量;在低资源场景下,合成数据可优于稀缺的真实对话,并进一步对真实对话形成补充。The results show that domain-grounded synthetic data consistently outperforms zero-shot prompting and naive synthetic baselines; active selection improves data efficiency over random sampling; metadata and collaborative filtering signals each improve selection quality; and, in low-resource settings, synthetic data can outperform scarce real dialogues while further complementing them.

Extending concurrent separation logic to the hardware level to verify the xv6 OS kernel on RISC-V with AI agents
将并发分离逻辑扩展到硬件层面,借助 AI Agent 在 RISC-V 上验证 xv6 OS 内核
arXiv:2609.04043 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

证明了一个应用层定理:若用户在 UART 控制台输入 echo hello world,系统唯一能产生的输出即为 hello world;这证明了基于 LLM 的 Agent 能够对如此底层的细节进行推理。An application-level theorem is proved: if the user types echo hello world as input on the UART console, the only output the system can produce is hello world, which proves LLM-based agents are capable of reasoning about such low-level details.

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States
Puffin-World:以原生 3D 世界状态扩展统一多模态模型
arXiv:2609.04196 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Puffin-World,一种统一的多模态架构,集成物理理解、空间仿真与 3D 世界生成重建,无需依赖外部离线模块,可支持需要多任务协同的交错式闭环应用。Puffin-World is proposed, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules and enables interleaved closed-loop applications requiring synergy across multiple tasks.

4️⃣ arXiv · Securing the Agent: Vendor-Neutral, Multitenant Enterprise Retrieval and Tool Use(⭐⭐⭐⭐ 高优先级)
4️⃣ arXiv · 守护 Agent:厂商中立的多租户企业级检索与工具调用(⭐⭐⭐⭐ 高优先级)
arXiv:2605.05287 Agent 智能体 观点 OA · 绿色 被引 1 · S2

本文提出一种分层隔离架构,结合策略感知的 ingestion、retrieval-time gating 与共享推理,并通过服务端 agentic 编排加以执行,在为多租户隔离提供天然强制点的同时,允许客户端框架保留对 agent 组合与延迟敏感操作的控制权。A layered isolation architecture combining policy-aware ingestion, retrieval-time gating, and shared inference, enforced through server-side agentic orchestration is introduced, creating natural enforcement points for multitenant isolation while allowing client-side frameworks to retain control over agent composition and latency-sensitive operations.