研究库 论文知识库
Papers · organized/paper_cards

论文

144 张论文卡片 · 应用落地

开放获取 全部 绿色 · 1640
Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems
Emergence World:长周期多 Agent 系统的对抗性压力测试。
arXiv:2609.17320 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,模型级对齐并不具备可组合性:单独能力强且看似安全的 agent 在组成系统后,可能随着 AI 的持续性与互联化而产生性质上不同的失效模式。The results suggest that model-level alignment is not compositional: individually capable and apparently safe agents can form systems with qualitatively different failure modes as AI becomes persistent and interconnected.

Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents
置信来自经验:从推理到 Agent 的经验性置信估计
arXiv:2609.17708 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 XConf (eXperiential Confidence):与模型累积经验一起估计置信度,并将经验式置信度估计视为未来通用置信度估计的新范式。This work proposes XConf (eXperiential Confidence): estimating confidence together with the model's accumulated experience, and sees experiential confidence estimation as a new paradigm for future general-purpose confidence estimation.

In-Context Robot Learning with VLM Agents
基于 VLM Agent 的机器人上下文学习
arXiv:2609.19138 Agent 智能体 应用落地 OA · 绿色 被引 5 · S2

本文提出 GPT-Policy,一个用于上下文机器人学习的通用 Agent 框架,集成了一个保留任务相关视觉过渡的 context compiler、一个提出机器人-工具动作的 VLM,以及一个验证并执行每个动作并报告结果的 constrained controller。GPT-Policy is introduced, a general-agent framework for in-context robot learning that integrates a context compiler that preserves task-relevant visual transitions, a VLM that proposes robot-tool actions, and a constrained controller that verifies and executes each action and reports its outcome.

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
DeepSeek-V4.1-Flash:突破 KV Cache 压缩的极限
arXiv:2609.19969 LLM 基础设施 应用落地 OA · 绿色 被引 43 · S2

推出 DeepSeek-V4.1-Flash 模型,这是一个具有 552B 骨干参数、支持最长一百万 token 上下文的多模态 Mixture-of-Experts 模型,显著提升了 agent 工作负载的成本效率,并突破了 KV cache 压缩的极限。The DeepSeek-V4.1-Flash model, a multimodal Mixture-of-Experts model with 552B backbone parameters and support for contexts of up to one million tokens, is introduced, substantially improving cost efficiency for agentic workloads and pushing the limits of KV cache compression.

Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents
反思、修订、复用:面向 GUI 智能体的免训练技能进化
arXiv:2609.17653 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出 EvoSkill-GUI,一个免训练框架,其中每个 skill 都是一个结构化的多文件包,包含检索元数据、可执行计划、备份定位、故障恢复规则、可访问性工具以及失败案例。This work proposes EvoSkill-GUI, a training-free framework in which each skill is a structured multi-file package containing retrieval metadata, executable plans, backup localization, failure-recovery rules, accessibility utilities, and failure cases.

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
不要遮蔽环境:观测监督会改变 RL 下智能体的探索行为
arXiv:2609.20715 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ActObs,对每条轨迹中已有的观测 token 也进行监督,并将这一差异归因于 SFT:动作与观测梯度迅速趋于正交,而仅训练动作会留下较大的残留观测梯度,并将环境预测能力拉低至基座模型之下。This work introduces ActObs, which also supervises the observation tokens already present in each trajectory, and traces this difference to SFT: action and observation gradients rapidly become orthogonal, while action-only training leaves a large residual observation gradient and degrades environment prediction below the base model.

Srijika: OpenType-Layout-Reusing Font Restyling for Nine Indic Scripts
Srijika:面向九种印度文字的 OpenType 布局复用字体再设计
arXiv:2609.05661 工程化 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Srijika,一个为九种 Brahmic 文字(Devanagari、Tamil、Bengali、Telugu、Kannada、Malayalam、Gujarati、Gurmukhi、Odia)生成可安装 OpenType 字体的系统,并附带一份负面结果目录,覆盖参考引导重风格化中失败的 conditioning、目标函数选择和数据凸包限制。Srijika is presented, a system for producing installable OpenType fonts for nine Brahmic scripts: Devanagari, Tamil, Bengali, Telugu, Kannada, Malayalam, Gujarati, Gurmukhi, and Odia, and a negative-results catalogue covering failed conditioning, objective choices, and data-hull limits of reference-guided restyling.

FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection
FRAUDSkill:面向音频反诈骗检测的结构化冻结权重 Skill 优化
arXiv:2609.18766 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出 FRAUDSkill——一种结构化的 frozen-weight 适配框架:底层音频-语言模型保持不变,转而优化外部的 skill program、路由策略和决策规则,并将结构化输出控制与验证引导的多路径推理相结合,以确保符合协议规范的预测。FRAUDSkill is proposed, a structured frozen-weight adaptation framework that leaves the underlying audio-language model unchanged while optimizing an external layer of skill programs, route-specific policies, and decision rules and combines structured output control with validation-guided multi-path inference to ensure protocol-compliant predictions.

MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads
MLLMs 在 Synergy Heads 中信息分布漂移时产生幻觉
arXiv:2609.09206 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

HEAL 将动态信息校准因子注入协同注意力头的 value 向量中,主动调节视觉-语言依赖,引导输出分布趋向事实证据,提供了一条简单且可解释的增强模型可信度的路径。HEAL injects dynamic information calibration factors into the value vectors of synergy heads and actively regulates visual-language dependencies, steering the output distribution towards factual evidence, offering a simple and interpretable pathway to enhance model trustworthiness.

Harness-Zero: Harness Distillation via Agent-as-Harness
Harness-Zero:通过 Agent-as-Harness 进行 Harness 蒸馏。
arXiv:2609.24974 评测基准 应用落地 被引 3 · S2

本文提出 Harness-Zero,通过 agent-as-harness 实现 harness 蒸馏,并证明在使用相同演化 harness 的前沿 LLM 上,agent-as-harness 优于 code-as-harness。This work introduces Harness-Zero, which enables harness distillation through agent-as-harness, and shows that for frontier LLMs using the same evolved harness, agent-as-harness outperforms code-as-harness.

ACLArena: Agent Continue Learning in Multi-stage Post-training
ACLArena:多阶段后训练中的 Agent 持续学习
arXiv:2609.23989 工程化 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ACLArena,一个全面研究、分析与评估 Agent 持续学习(ACL)的框架,并提出新的 ACL 方案:将高质量轨迹的离线回放与多个由 RL 专精化的 LoRA 专家路由网络相结合,显著提升 Agent 跨多领域学习的能力。This work introduces ACLArena, a framework for comprehensively studying, analyzing, and evaluating Agent Continual Learning, and proposes a new ACL recipe that combines offline replay over high-quality trajectories with a routed network of multiple LoRA experts each specialized via RL, substantially improving the agent's ability to learn across multiple domains.

All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
脚本感知的 Mixture-of-Experts 统一多语言场景文本识别
arXiv:2609.24058 多模态 应用落地 OA · 绿色 被引 1 · S2

本文构建了 TextMuSS-10M,一个涵盖 10 种文字、229 种语言的大规模合成场景文本数据集,并提出了 ScriptMoE,一种具备文字感知能力的 Mixture-of-Experts (MoE) 架构。该架构在精度上达到最高,且比 per-language experts 更简单、比 VLM 更轻量,同时精度优于两者。This work constructs TextMuSS-10M, a large-scale synthetic scene text dataset spanning 10 scripts and 229 languages and proposes ScriptMoE, a script-aware Mixture-of-Experts (MoE) architecture that achieves the highest accuracy and is simpler than per-language experts, lighter than VLMs, and more accurate than both.

Emergent Collusion in Long-Horizon LLM Agent Interaction
长程 LLM Agent 交互中的涌现合谋
arXiv:2609.24967 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

总体而言,研究结果表明,长期交互会以产生安全风险的方式重塑 Agent 的协作模式;限制 Agent 可获取的交互历史数量与范围能够减少串通行为。Overall, the findings show that long-horizon interaction can reshape how agents coordinate in ways that create safety risks, and restricting the amount and scope of interaction history available to agents reduces collusion.

Blaming Across the Aisle: Political Contrasting and Blame Attribution in the Danish Parliament
跨越党派指责:丹麦议会中的政治对比与归责
arXiv:2609.26346 LLM 基础设施 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

政治话语日益呈现更强的敌对感是普遍观感,但稳健证据仍稀缺。本文研究 1997 至 2026 年丹麦议会的归责行为,结合专门构建的分类器 BlameBERT(F1: 0.80)与多层统计建模。该分类器采用面向低至中等资源语言的标注高效流程。结果显示出一条香蕉形轨迹:归责水平约在 2016 年前下降,随后在近年(2019–2026)进入显著且持续的上升阶段。执政地位显著Political discourse is widely perceived to be growing more hostile, yet robust evidence remains scarce. This study examines blame attribution in the Danish Parliament from 1997 to 2026, combining a purpose-built classifier, BlameBERT (F1: 0.80), with multilevel statistical modeling. The classifier is constructed using an annotation-efficient pipeline for blame attribution in low-to-mid resource languages. The results reveal a banana-shaped trajectory, with blame declining until around 2016 before entering a significant and sustained increase in recent years (2019-2026). Government status consi

AgentKernel: The Trust-Native Agentic Operating System
AgentKernel:原生可信的智能体操作系统
arXiv:2609.29647 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文论证代理需要一个操作系统级基底,为身份、输入中介、内存治理与执行控制提供强制且不可绕过的服务,并提出 AgentKernel,一个以安全为一流设计约束为前提、面向信任的代理操作系统。This work argues that agents need an operating-system substrate providing mandatory, non-bypassable services for identity, input mediation, memory governance, and execution control, and introduces AgentKernel, a trust-native agent operating system built around the premise that security must be a first-class design constraint.

Jev in the Wild: A Data-Driven Analysis of the Jev Model's Functionality, Applications and Ecosystem
野外 Jev:Jev 模型功能、应用与生态的数据驱动分析
arXiv:2609.30216 Agent 智能体 应用落地 OA · 绿色 被引 5 · S2

本文对从 GitHub 收集的 Jev 项目进行了大规模、数据驱动的分析,发现其公共生态在早期增长迅速,新项目不断涌现并被集成到既有仓库中;研究提示 Jev 充当一种可复用的决策组件,其功能随所嵌入的工作流而变化。A large-scale, data-driven analysis of Jev projects collected from GitHub finds rapid early growth in Jev's public ecosystem, with both new projects and integration into existing repositories, and suggests that Jev serves as a reusable decision component whose functionality varies with the surrounding workflow.

CARD: Cluster-level Adaptation with Reward-guided Decoding for Personalized Text Generation
[标题中文] CARD:面向个性化文本生成的基于聚类级自适应与奖励引导解码
arXiv:2601.06352 工程化 应用落地 OA · 绿色 被引 2 · S2

本文提出 CARD,一种通过渐进式细化实现有效个性化的层级框架:先按共享风格模式对用户聚类,再为各组学习专用的 LoRA adapter,从而在低资源场景下也能实现稳健的泛化与强劲的性能。This work presents CARD, a hierarchical framework that achieves effective personalization through progressive refinement that first clusters users according to shared stylistic patterns and learns group-specific LoRA adapters, enabling robust generalization and strong low-resource performance.

EngramRAG: Dynamic Usage-Weighted Topology and Synaptic Consolidation for Multi-Hop Agentic Memory
EngramRAG:面向多跳 Agentic 记忆的动态使用加权拓扑与突触巩固。
arXiv:2609.32049 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

EngramRAG 是一种自适应记忆架构,将低延迟的 Waking State 反射与异步后台 Dreaming State 整合周期耦合,并引入融合密集向量、BM25 与 U-PPR 的三源混合检索,通过动态 Reciprocal Rank Fusion (RRF) 实现。The proposed EngramRAG is an adaptive memory architecture coupling a low-latency Waking State reflex with an asynchronous background Dreaming State consolidation cycle, and introduces triple-source hybrid retrieval fusing dense vectors, BM25, and U-PPR via dynamic Reciprocal Rank Fusion (RRF).

ASCT: Attentive Search over Counterfactual Trees for Credit Assignment in Agentic Reinforcement Learning
ASCT: 在反事实树上进行注意力搜索,用于 Agentic 强化学习中的信用分配
arXiv:2609.35215 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出反事实树上的注意力搜索(ASCT),将训练时的多步搜索转化为局部动作信用,将反事实评估与策略学习相连接,同时部署时只需使用 actor。Attentive Search over Counterfactual Trees (ASCT) is introduced, a framework that turns training-time multi-step search into local action credit and connects counterfactual evaluation to policy learning while deploying the actor alone.

WISE-ATTA: When to Ask for Labels in Budgeted Active Test-Time Adaptation
[标题中文] WISE-ATTA:预算受限的主动测试时自适应中的标签请求时机
arXiv:2609.37687 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Budgeted ATTA:测试批次中仅有一小部分可获得标签,且监督时机是 active test-time adaptation 中一个关键但尚未充分探索的方面。Budgeted ATTA is introduced in which labels are available for only a fraction of test batches, and the timing of supervision is a key, yet underexplored, aspect of active test-time adaptation.

EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making
EVOKE:在 Agent 中引出世界知识以实现可迁移的决策
arXiv:2609.38334 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 EVOKE,一种后训练方法,通过在固定状态下以目标多样性对直接决策施加监督,来施加压力以激发模型内化的、可迁移动作的世界知识,并提供了一种通过直接决策监督激发内化世界知识以获得可迁移动作的新视角。EVOKE is introduced, a post-training method that supplies pressure on eliciting internalized world knowledge for transferable action through direct decision supervision through goal diversity at fixed states, and offers a new perspective on eliciting internalized world knowledge for transferable action through direct decision supervision.

RAGScope: A Leakage-Controlled, Cost-Aware Evidence-Gating Protocol for RAG Hallucination Triage
RAGScope:面向 RAG 幻觉分诊的泄漏可控、成本感知的证据门控协议
arXiv:2609.39075 RAG 检索增强 应用落地 OA · 绿色 被引 1 · S2

RAGScope 是一个泄漏受控的评估协议,用于评估仅使用任务输入、检索上下文和答案文本的本地证据门,结合了上下文分组划分、折范围预处理、组自举区间、部署工作点、端到端运行时以及显式的源偏移压力测试。RAGScope, a leakage-controlled protocol for evaluating local evidence gates that use only the task input, retrieved context, and answer text is presented, which combines context-grouped splits, fold-scoped preprocessing, group bootstrap intervals, deployment operating points, end-to-end runtime, and explicit source-shift stress tests.

Replacing Large Language Models with Jev Decision Models for Low-Latency Edge Service Orchestration
用 Jev 决策模型替代大语言模型以实现低延迟边缘服务编排
arXiv:2609.22753 Agent 智能体 应用落地 OA · 绿色 被引 10 · S2

将 Jev 面向决策的 API 集成到边缘服务编排中,在保持服务完成度的同时降低开销,并支持在有界契约下针对延迟受限的准入进行决策模型替换。Jev's decision-oriented application programming interface (API) is integrated into edge service orchestration to reduce overhead while retaining service completion, and decision-model substitution for latency-bound admission on bounded contracts is supported.

"I think this is the most disruptive technology": Exploring Sentiments of ChatGPT Early Adopters using Twitter Data
"我认为这是最具颠覆性的技术":基于 Twitter 数据探索 ChatGPT 早期采用者的情感
arXiv:2212.05856 评测基准 应用落地 OA · 绿色 被引 284 · S2

基于 10,732 条早期 ChatGPT 用户推文的混合方法研究,对每个主题进行深入定性情感分析,结果显示大多数早期采用者在软件开发颠覆性、娱乐与创意发挥等主题上表达了压倒性的积极情感。A mixed-method study using 10,732 tweets from early ChatGPT users to conduct an in-depth qualitative sentiment analysis of each topic, showing that the majority of the early adopters have expressed overwhelmingly positive sentiments related to topics such as Disruptions to software development, Entertainment and exercising creativity.

Tracking State Footprints: How Agents Can Transact
追踪状态足迹:Agent 如何进行事务处理
arXiv:2610.03140 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文将多智能体系统(MAS)协调建模为数据管理问题,提出用智能体的状态足迹来刻画它们:即在其自身局部上下文与状态、以及编排器和外部系统状态上的读写行为。This work frames MAS coordination as a data management problem and proposes to describe agents by their state footprint: the state they read and write across their own local context and state, as well as the state of the orchestrator and external systems.

Collective Bias Mitigation via Model Routing and Collaboration
Collective Bias Mitigation:通过模型路由与协作的集体偏见缓解
arXiv:2610.03240 工程化 应用落地 OA · 绿色 被引 1 · S2

本文首次系统地探索了不同 LLM 的有效选择与组织,以培育更公平的 LLM 回答,并展示了 CBM 显著优于独立基线。This work is the first to systematically explore the effective selection and organization of distinct LLMs to cultivate fairer LLM responses and show CBM substantially outperforms standalone baselines.

Self-Supervised Scaling of Terminal Environments for Scientific Domains
科学领域终端环境的自监督规模化
arXiv:2610.02710 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明现有科学软件能够为终端 Agent 提供可扩展且经过行为验证的监督,并提出 software-in-the-loop reconstruction,一种自监督框架,从现有软件工作流(即把结构化输入映射为输出的可执行程序)中获取参考输出与验证目标。Results indicate that existing scientific software can provide scalable, behaviorally verified supervision for terminal agents, and introduces software-in-the-loop reconstruction, a self-supervised framework that obtains reference outputs and verification targets from existing software workflows, executable programs mapping structured inputs to outputs.

Taming VLAs under Robot Execution Errors: Self-Compensation and Stress Testing
在机器人执行误差下驯服 VLA:自补偿与压力测试
arXiv:2609.37334 多模态 应用落地 OA · 绿色 被引 1 · S2

论文提出 Self-compensating VLA,一种部署阶段的自适应方法,使 VLA 策略在生成指令时能够预补偿机器人的执行误差,并在平均任务成功率上高于基线策略以及在训练阶段增强鲁棒性的方法。Self-compensating VLA is proposed, a deployment-time adaptation method that enables a VLA policy to pre-compensate for the robot's execution errors when generating commands, and achieves higher average task success than both the base policies and methods that build in robustness during training.

Co-Evolving Robot Orchestrators and Policies through Deployment
通过部署协同进化机器人编排器与策略
arXiv:2610.09228 多模态 应用落地

在大规模数据集上训练的视觉-语言-动作(VLA)策略在其训练域内表现良好,但仍难以泛化到真实部署中机器人遭遇的各种场景。Agentic 机器人系统通过视觉-语言模型(VLM)编排器来弥补策略的不足,该编排器学习何时调用策略、如何下达指令、以及何时改用脚本化技能。然而,由于整个系统围绕一个语言可操控性有限的冻结策略构建,编排器只能规避策略的失败却无法真正克服它们。策略成为瓶颈……Vision-language-action (VLA) policies trained on large datasets are capable within their training domains, yet they still fail to generalize to the variety of situations a robot meets in real-world deployment. Agentic robot systems complement the policy with a vision-language model (VLM) orchestrator that learns when to call the policy, how to instruct it, and when to use scripted skills instead. However, because the harness is built around a frozen policy that has limited language steerability, the orchestrator can avoid the policy's failures but never overcome them. The policy becomes the bo

FastOPD: On-Policy Distillation for Lightweight VLA Deployment
FastOPD:面向轻量 VLA 部署的在线蒸馏
arXiv:2610.02832 多模态 应用落地 OA · 绿色 被引 0 · OpenAlex

视觉-语言-动作(VLA)基础模型规模迅速扩大以提升操作性能与泛化能力,但这种规模化带来了高昂的计算成本,使真实世界部署日益困难。现有方法通常通过设计更小的架构或减少基于流(flow-based)策略中的迭代去噪步数来缓解该问题。本文提出 FastOPD,一个从基础到轻量的 VLA 框架,通过高效的在线蒸馏实现大规模 VLA 的实际部署。具体而言,FastOPD 适配流映射(flow map)……Vision-Language-Action (VLA) foundation models have scaled rapidly to enhance manipulation performance and generalizability, but this scaling incurs high computational costs that render real-world deployment increasingly challenging. Existing approaches typically mitigate this issue by designing smaller architectures or reducing the iterative denoising steps in flow-based policies. In this work, we propose FastOPD, a foundation-to-lightweight VLA framework that enables the practical deployment of large-scale VLAs through efficient on-policy distillation. Specifically, FastOPD adapts a flow map

ORCAGen: Orchestrating Context-Aware Malware Deception with RAG-Guided Generative AI
ORCAGen:基于 RAG 引导生成式 AI 的上下文感知恶意软件欺骗编排
arXiv:2610.12415 RAG 检索增强 应用落地

恶意软件防御通常会尽快移除或隔离可疑程序。该策略虽便于遏制,却也浪费了观察攻击者行为与部署针对性反制的机会。ORCAGen 另辟蹊径:离线利用 GenAI 构建针对特定恶意软件的欺骗剧本,部署前进行校验,运行时仅强制执行已验证的逻辑。ORCAGen 将 RAG 与结构化提示工程相结合,以同时生成 PoC 恶意软件及对应的欺骗编排代码。Malware defenses often remove or isolate suspicious programs as quickly as possible. While effective for containment, this approach can also waste an opportunity to observe attacker behavior and deploy targeted countermeasures. ORCAGen takes a different approach: it uses GenAI to build malware-specific deception playbooks offline, validates them before deployment, and enforces only the verified logic at runtime. ORCAGen combines Retrieval-Augmented Generation (RAG) with structured prompt engineering to generate both proof-of-concept (PoC) malware and corresponding deception orchestration code.

Is Memorization Context-Sensitive? Prefix-Based Extraction Beyond Isolated Prefixes
记忆化是否具有上下文敏感性?超越孤立前缀的前缀式提取
arXiv:2610.12085 RAG 检索增强 应用落地

LLM 可在前缀式提取下泄露记忆化的训练序列:给定训练样本的前缀,模型可能为原始续写赋予高概率。但在部署系统中,前缀很少被单独评估,常与指令、检索文档或其他任务相关上下文一同出现,正如 RAG 所做的那样。这促使我们去考察:上下文条件化究竟是缓解了记忆化,还是仅仅改变了可被提取的记忆化样本集合。本文对这一问题展开研究……Large language models (LLMs) can expose memorized training sequences under prefix-based extraction: given a prefix from a training example, the model may assign high probability to the original continuation. In deployed systems, however, prefixes are rarely evaluated in isolation. They often appear together with instructions, retrieved documents, or other task-specific context, as in retrieval-augmented generation (RAG). This motivates examining whether contextual conditioning mitigates memorization or merely changes the set of memorized samples that become extractable. We investigate this iss

From Foundation to Application: Improving VLA Models in Practice
从基础到应用:实践中改进 VLA 模型
arXiv:2607.06403 多模态 应用落地 OA · 绿色 被引 28 · S2

得益于涵盖全身自由度的扩展预训练数据,LingBot-VLA-2.0 在两个机器人平台上展现出强大的跨具身长时程移动操作能力。Benefiting from the expanded pretraining data that covers whole-body degrees of freedom, LingBot-VLA-2.0 demonstrates strong cross-embodiment long-horizon mobile manipulation capability across the two robotic platforms.

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering
MultAttnAttrib:长文档问答中的免训练多模态归因
arXiv:2607.01420 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本工作提出 MultAttnAttrib,一种免训练的归因生成方法,利用模型的预填充过程、选定的注意力头以及校准阈值在文档中定位源证据,且在多种归因生成方法上一致地表现更优。This work introduces MultAttnAttrib, a training-free attribution-generation method that leverages a model's prefill pass, selected attention heads, and calibrated thresholds to locate source evidence within a document, and consistently outperforms a variety of attribution-generation methods.

Teaching LLMs to Recommend and Defer in Underrepresented Epilepsy Care
教导 LLM 在欠发达的癫痫诊疗中进行推荐与转诊
arXiv:2606.31036 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

在资源受限环境中,专业癫痫专家稀缺,使基于 LLM 的决策支持对管理纵向治疗的一线临床医生具有吸引力。此类系统必须适应当地处方实践并知道何时转诊。我们在乌干达儿科癫痫诊疗中研究该问题,基于纵向非结构化门诊记录预测抗癫痫用药方案。标准提示与医生处方取得了一定程度的一致性,但神经科医生审查显示许多错误反映的是分布失校的处方默认值而非失败。Specialist epilepsy expertise is scarce in resource-constrained settings, making LLM-based decision support attractive for frontline clinicians managing longitudinal treatment. Such systems must adapt to local prescribing practice and know when to defer. We study this problem in Ugandan pediatric epilepsy care, predicting anti-seizure medication regimens from longitudinal unstructured clinic notes. Standard prompting achieves non-trivial agreement with physician prescriptions, but neurologist review shows that many errors reflect distribution-miscalibrated prescribing defaults rather than fail

JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications
京东 Oxygen AI 商品中心(Oxygen AIIC)V1:以 LLM/VLM 为核心的工业级商品理解、管理与应用解决方案
arXiv:2606.28070 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

京东 Oxygen AI 商品中心(Oxygen AIIC):基于 LLM/VLM 的工业级商品知识生产与服务平台,已在大规模场景下取得可量化的收益。The JD Oxygen AI Item Center (Oxygen AIIC), an industrial-scale platform built on LLMs/VLMs for item-knowledge production and service, has delivered measurable gains at scale.