研究库 论文知识库
Papers · organized/paper_cards

论文

1686 张论文卡片 · OA 绿色

开放获取 全部 绿色 · 1686
OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning
OmniCapBench:面向细粒度音视频描述的深度结构化评估框架
arXiv:2610.12458 多模态 评测集 OA · 绿色 被引 0 · OpenAlex

多模态大语言模型(MLLMs)正快速发展以实现持续的音视频推理,这迫切需要能够揭示其能力上限的评估。音视频描述是一项理想的诊断任务,但现有 benchmark 面临耦合的权衡:整句评分覆盖全面但缺乏定位,局部探针定位精确但缺乏覆盖,且无约束的 LLM 评判器带来不稳定性。我们提出 OmniCapBench(Omni-Video Caption Benchmark),将音视频描述评估重构为深度结构化的诊断框架Multimodal large language models (MLLMs) are rapidly evolving toward continuous audio--visual reasoning, creating an urgent need for evaluations that expose their capability limits. Audio--visual captioning is an ideal diagnostic task, yet current benchmarks face a coupled trade-off: whole-caption scores provide coverage without localization, local probes provide localization without coverage, and unconstrained LLM judges introduce instability. We introduce OmniCapBench (Omni-Video Caption Benchmark), a benchmark that reframes audio--visual caption evaluation as a deep-structured diagnostic fr

Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
Memento 3:通过反思式规则手册实现基于模型的递归自我改进
arXiv:2610.11794 Agent 智能体 方法 OA · 绿色 被引 0 · OpenAlex

在未知环境中学会行动需要智能体推断世界运作方式,并在新证据出现时修正理解。然而,有限的观察可能支持多个世界模型,它们都能解释过去的交互,但对未见状态有不同的预测。我们提出 Memento 3,在 Memento 系列的基础上,使冻结的 LLM 智能体能够通过外部记忆持续学习显式世界模型。智能体维护一个自然语言规则手册作为持久化语义记忆,记录可修正的环境动态假设,同时保留未知部分Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent semantic memory, recording revisable hypotheses about environment dynamics while leaving unknown aspects

WorldGuide: Goal-Directed Video World Model for Procedural Task Execution
WorldGuide:面向过程化任务执行的目标导向视频世界模型
arXiv:2610.12459 多模态 方法 OA · 绿色 被引 0 · OpenAlex

视频生成器和基于视频的世界模型能够合成合理的视觉轨迹,但长时程过程化任务要求生成过程能适应已实际生成的内容。模型必须根据其生成的状态决定下一步动作、执行该动作,并识别任务何时完成。开环(Open-loop)生成无法适应执行结果,而现有闭环(closed-loop)系统往往依赖预训练执行器或间接验证,导致动作决策与成功执行之间存在鸿沟。我们将过程化视频生成形式化为闭环Video generators and video-based world models can synthesize plausible visual trajectories, but long-horizon procedural tasks require generation to adapt to what has actually been produced. A model must determine the next action from its generated state, execute that action, and recognize when the task is complete. Open-loop generation cannot adapt to execution outcomes, while existing closed-loop systems often rely on pretrained executors or indirect verification. This leaves a gap between deciding an action and successfully realizing it. We formulate procedural video generation as closed-loo

One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
One Block, Multiple Depths:具有深度编程专家的循环视觉 Transformer
arXiv:2610.12448 多模态 方法 OA · 绿色 被引 0 · OpenAlex

在本工作中,我们证明单个 Transformer block 循环应用即可在相近推理 FLOPs 下匹配全深度视觉编码器的精度,且无需中间特征蒸馏。reViT 通过将每个循环深度的 FFN 表示为小型共享专家库的凸组合来恢复深度特定的变换。一个连续归一化的深度坐标对该混合进行编程,在 FFN 参数空间中定义一条可重采样的轨迹。我们在两种场景下评估该设计:有监督 ImageNet-1k 训练以及从 DINOv2 教师模型蒸馏。在各规模下In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank. A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space. We evaluate this design in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher. Across b

A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization
深入探究 Agentic BBO:面向黑盒优化的 LLM 智能体基准测试
arXiv:2610.12183 Agent 智能体 评测集 OA · 绿色 被引 0 · OpenAlex

黑盒优化(BBO)出现在许多目标和工程问题中,其目标函数评估代价高昂且次数有限。近期的大语言模型(LLM)智能体通过结合任务语义、计算、优化工具以及反馈驱动的决策,提供了一种新的 BBO 解决方式,因与数学严谨工具的集成而展现出巨大潜力。然而,现有的 Agentic BBO 研究使用不同的任务领域和系统配置,导致结果难以比较,且单个设计选择的影响难以孤立分析。因此,我们引入Black-box optimization (BBO) arises in many scientific and engineering problems where objective evaluations are expensive and limited. Recent large language model (LLM) agents offer a new way to approach BBO by combining task semantics, computation, optimization tools, and feedback-driven decision making, showing great potential due to the integration with mathematically rigorous tools. However, existing agentic BBO studies use different task domains and system configurations, making their results difficult to compare and the effects of individual design choices hard to isolate. We therefore i

REMORY: Learning Residual Memory for Context Compaction
REMORY:学习用于上下文压缩的残差记忆
arXiv:2610.11287 Agent 智能体 方法 OA · 绿色 被引 0 · OpenAlex

长时程智能体压缩其历史以在有限上下文窗口内继续执行,但单一的文本摘要可能无法支撑所有后续决策。我们提出 REMORY,一种神经记忆网络,通过有界序列的软记忆 token 来补充摘要。给定历史和摘要,该网络学习生成帮助冻结 LLM 近似其在完整历史下产生的续写内容的 token。这些 token 以摘要为条件并附加在其后,构成沿序列维度的残差连接类比。在 SummHay 上,REMORY 提升了Long-horizon agents compact their history to continue within a finite context window, but a textual summary alone may not support every subsequent decision. We introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens. Given the history and summary, the network learns to generate tokens that help a frozen LLM approximate the continuation it would produce with the full history. The tokens are conditioned on the summary and appended after it, forming an analogue of a residual connection along the sequence dimension. On SummHay, REMORY im

Opera: A Verbal Critic Framework for Long-horizon Coding Agents
Opera:面向长程编码 Agent 的口头评论框架
arXiv:2609.33987 Agent 智能体 方法 OA · 绿色 被引 0 · OpenAlex

长程编码 Agent 需要及时纠偏,但反馈若误判当前工作或未触及根本问题,可能无效甚至有害。现有评论方法聚焦于评估轨迹并生成反馈,却很少追踪反馈给出后的情况。我们提出 Opera —— 一种口头评论框架,将每次纠偏视为一条持续存在的备忘,跟踪至所诊断问题被解决。Opera 通过周期触发与事件驱动触发决定复盘时机,借助类型化算子诊断问题,并依据可见证据对反馈进行审计。Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is delivered. We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved. Opera decides when to review through periodic and event-driven triggers, diagnoses issues with typed operators, audits feedback against visible evi

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
PluraMath:将数学推理评估拓展至丰富资源语言之外
arXiv:2607.05992 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

确认了丰富资源语言与代表性不足语言之间在数学推理性能上存在持续差距,性能更优主要与更强的指令遵循能力相关,并提出了完全开源的数据集、数据采集流程与评估框架。A persistent gap in mathematical reasoning performance between high-resource and underrepresented languages is confirmed, with stronger results largely associated with better instruction-following ability, and a fully open-source dataset, data acquisition pipeline, and evaluation framework is introduced.

Skill Constellations: Tracing the Supply Chain of Agent Skills on GitHub
Skill Constellations:追踪 GitHub 上 Agent Skill 的供应链
arXiv:2610.11169 Agent 智能体 方法 OA · 绿色 被引 0 · OpenAlex

Agent Skill 是 SKILL.md 指令与脚本,由 Claude Code、Codex 等 AI 编码 Agent 以其用户权限运行。开发者通过在仓库间复制来共享 Skill,使其成为缺乏注册表、版本与来源信息的软件供应链。被复制 Skill 的来源、安全修复的覆盖范围以及需要审查的仓库均无从知晓。仅在单一时点记录哪些仓库持有 Skill 的研究,无法揭示彼此之间的复制关系。我们贡献了首个带时间戳的 Agent Skill 复制网络,构建于……Agent skills are SKILL.md instructions and scripts that AI coding agents such as Claude Code and Codex run with the permissions of their user. Developers share skills by copying them between repositories, which makes them a software supply chain without a registry, versions or provenance. The origin of a copied skill, the reach of a security fix and the repositories that warrant review are therefore unknown. Studies that record which repositories hold a skill at a single point in time cannot reveal who copied it from whom. We contribute the first dated copy network of agent skills, built from

A GPU-Parallel Framework for Heterogeneous Multi-Task Reinforcement Learning
面向异构多任务强化学习的 GPU 并行框架
arXiv:2606.03335 评测基准 评测集 OA · 绿色 被引 0 · OpenAlex

GPU 并行仿真带来丰富的机器人交互,但现有 benchmark 很少将这种规模与异构操作任务以及标准化的多任务 RL 评估相结合。我们提出 Hebero(异构机器人学习 benchmark)—— 基于 Isaac Lab 的 GPU 并行 benchmark,可在全部 40 项异构任务上对单一策略进行高效联合训练与评估。扩展实验表明,在固定时间预算下,增加每任务并行副本数可提升成功率。为支持稀疏奖励与有限演示条件下的学习,我们提出 Dem……GPU-parallel simulation provides abundant robot interaction, but existing benchmarks rarely combine this scale with heterogeneous manipulation tasks and standardized multi-task RL evaluation. We introduce Hebero (Heterogeneous Benchmark for Robot Learning), a GPU-parallel Isaac Lab benchmark that enables efficient joint training and evaluation of a single policy across all 40 heterogeneous tasks. Scaling experiments show that increasing parallel replicas per task improves success under a fixed wall-clock budget. To support learning with sparse rewards and limited demonstrations, we propose Dem

Mara Chain: Rethinking Failure as a Stepping Stone for AI System Auto-Evolution
Mara Chain:将失败重新视为 AI 系统自动演化的踏脚石
arXiv:2609.35855 评测基准 观点 OA · 绿色 被引 0 · OpenAlex

对已部署 AI 系统的优化日益体现为对 prompt、skill、harness 与代码的编辑,而非模型权重。现有方法通常采用 propose-evaluate-select 流程优化这些产物:评估候选配置,仅保留满足接受标准的方案。然而我们的分析表明,被丢弃的候选方案往往包含对后续优化至关重要的信息;丢弃它们会导致后续提案反复遭遇相同的失败模式。我们提出 Mara Chain —— 一种将拒绝候选转化为……Optimizing deployed AI systems increasingly amounts to editing prompts, skills, harnesses, and code rather than model weights. Existing approaches commonly optimize these artifacts through propose-evaluate-select procedures, where candidate configurations are evaluated and only those meeting an acceptance criterion are selected. Yet our analysis shows that discarded candidates often contain information critical for subsequent optimization. Discarding them causes later proposals to revisit the same failure modes. We introduce Mara Chain, a refinement procedure that turns rejected candidates int

BinaryConnect: Training Deep Neural Networks with binary weights during propagations
arXiv:1511.00363 工程化 方法 OA · 绿色 被引 1802 · OpenAlex
CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration
CanvasAgent:通过视觉工具编排实现复杂图像创建与编辑
arXiv:2607.05465 Agent 智能体 方法 OA · 绿色 被引 1 · S2

提出了用于复杂图像创建与编辑的大规模多模态工具调用数据集 CanvasCraft,以及通过多轮交互学习编排异构视觉工具的工具增强多模态 Agent CanvasAgent。CanvasCraft, a large-scale multimodal tool-use dataset for complex image creation and editing, and CanvasAgent, a tool-augmented multimodal agent that learns to orchestrate heterogeneous visual tools through multi-turn interaction are introduced.

KVpop -- Key-Value Cache Compression with Predictive Online Pruning
KVpop——基于预测性在线剪枝的键值缓存压缩
arXiv:2607.05061 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

提出 KVpop,通过直接监督保留或丢弃决策来学习固定预算的 KV 淘汰策略,并引入基于延迟记忆的打分器,在已学习的淘汰方法中独有地延迟固定步数的打分,以利用近期上下文。KVpop is introduced, which learns a fixed-budget KV eviction policy by directly supervising the keep-or-drop decision, and introduces a delayed memory-based scorer that, uniquely among learned eviction methods, defers scoring for a fixed number of steps to exploit near-future context.

MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing
MV-Forcing:通过 4D 几何基础的时空自强迫实现长多视角视频生成
arXiv:2607.05376 多模态 方法 OA · 绿色 被引 1 · S2

提出 MV-Forcing 框架,通过在顺序生成视角之间引入 4D 几何桥梁,在单一扩散模型中组合时间与视角自回归,并弥合时间与视角序列自回归中训练-推理的曝光偏差差距。MV-Forcing, a framework that composes temporal and view-wise autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views, and closes the train-inference exposure bias gap for both temporal and view-sequential autoregression.

AI Wizards at EXIST 2026: Hierarchical Soft-Label Learning for Multimodal Sexism Identification in Memes
AI Wizards 参加 EXIST 2026:用于迷因中多模态性别歧视识别的分层软标签学习
arXiv:2607.04410 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作通过轻量级 Gated MLP 将固定的 Gemini Embedding 2 视觉-语言表征映射到目标空间,使用 KL 散度与同方差不确定性加权进行训练,并展示了 AI Wizards 参加 EXIST 2026 多模态迷因性别歧视识别任务的提交方案。This work maps fixed Gemini Embedding 2 vision-language representations through a lightweight Gated MLP trained with KL divergence and homoscedastic uncertainty weighting, and presents the AI Wizards submission to EXIST 2026 for multimodal sexism identification in memes.

Improving Access to Historical Archives with Real-time RAG-based Systems
基于实时 RAG 系统提升历史档案的可访问性
arXiv:2607.03440 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出了一个端到端的档案处理与检索框架,将大语言模型(LLM)集成到档案流程中,并证明将 LLM 与成熟的文档处理与检索流程相结合,可将数字图书馆从静态存储库提升为可交互、可语义检索的档案系统。This work presents an end-to-end archival processing and retrieval framework that integrates large language models (LLMs) into the archival pipeline and demonstrates that integrating LLMs with established document processing and retrieval pipelines can elevate digital libraries from static repositories to interactive, semantically searchable archival systems.

HETERQA: Benchmarking Record Retrieval over Multiple Heterogeneous Sources
HETERQA:跨多个异构来源的记录检索基准
arXiv:2607.03028 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出 HETERQA,一个包含 857 个 QA 对的综合性基准,涵盖五个异构来源的记录检索,并表明 HETERQA 为异构来源下的记录检索提供了有效的测试平台,为未来检索方法留下了显著空间。This work introduces HETERQA, a comprehensive benchmark with 857 QA pairs for record retrieval over five heterogeneous sources and indicates that HETERQA provides an effective testbed for record retrieval over heterogeneous sources and leaves substantial room for future retrieval methods.

MANCE: Manifold Aware Concept Erasure
MANCE:流形感知的概念擦除
arXiv:2607.03973 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出流形约束假设(Manifold Constraint Hypothesis, MCH):若自然表征集中于一个结构化的低维流形上,则干预应被约束在该流形内,从而在干预过程中更好地保留表征中编码的其他信息。The Manifold Constraint Hypothesis (MCH) is proposed: if natural representations concentrate on a structured, lower-dimensional manifold, then interventions should be constrained to that manifold and better preserve other information encoded in the representation during interventions.

Taste-aware music retrieval from audio embeddings
基于音频嵌入的品味感知音乐检索
arXiv:2607.03296 RAG 检索增强 评测集 OA · 绿色 被引 2 · S2

将预测的味觉空间作为基于内容的检索索引,对 309 项条目池的排序比 CLAP-text 基线(处于随机水平)忠实得多;ridge probes 与 audio-bandstop knockout 在已记载的声-味对应关系上读出了最强表征。Operationalised as a content-based retrieval index, the predicted taste space ranks a 309-item pool far more faithfully than a CLAP-text baseline, which sits at chance; ridge probes and an audio-bandstop knockout read the strongest representations against documented sound-taste correspondences.

CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training
CONFLUX:用于胸部 CT 三维合成的潜在扩散模型与强化学习后训练
arXiv:2607.02998 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

CONFLUX 是一个面向胸部 CT 的潜在扩散模型:由 3D 变分自编码器压缩每个体数据,整流流 Transformer 在潜在空间中生成,并以分类器从生成体中恢复所请求病征的可靠性作为奖励。CONFLUX, a latent diffusion model for chest computed tomography (CT): a 3D variational autoencoder compresses each volume, and a rectified-flow transformer generates in the latent space, that rewards how reliably a classifier recovers the requested findings from each generated volume.

Gemma 4 Technical Report
Gemma 4 技术报告
arXiv:2607.02770 多模态 方法 OA · 绿色 被引 206 · S2

本工作推出 Gemma 4——Gemma 模型系列中新一代开源权重、原生多模态的语言模型,在 STEM、多模态与长上下文基准上实现性能跃升,在人类评分任务上可与更大的前沿开源模型相媲美。This work introduces Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family that establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers
OrbitQuant:面向图像与视频扩散 Transformer 的数据无关量化
arXiv:2607.02461 多模态 方法 OA · 绿色 被引 5 · S2

本文提出 OrbitQuant,一种数据无关的权重量化器,通过在归一化旋转基空间中进行量化以绕过范围估计,将图像扩散 Transformer 的 PTQ 推进到 W2A4 并保持可用生成质量。OrbitQuant is presented, a data-agnostic weight-activation quantizer that bypasses range estimation by quantizing in a normalized, rotated basis and pushes PTQ of image diffusion transformers to W2A4 with usable generation quality.

When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffers
当经典缓存策略失效时:面向语义检索缓冲区的学习增强替换
arXiv:2607.00394 RAG 检索增强 方法 OA · 绿色 被引 1 · S2

本文提出 SOLAR,一种学习增强框架,从 regret 累积中推导修改时机,并基于隐式检索反馈的贝叶斯在线学习进行内容选择,实现与缓存大小和时域无关的常数竞争比。SOLAR is proposed, a learning-augmented framework that derives modification timing from regret accumulation and content selection from Bayesian online learning over implicit retrieval feedback and achieves a constant competitive ratio, independent of cache size and horizon.

Measuring the Gap Between Human and LLM Research Ideas
衡量人类与 LLM 研究思路之间的差距
arXiv:2607.01233 评测基准 方法 OA · 绿色 被引 9 · S2

结果表明,强大的 LLM 能够产出一系列合理思路,但其范围仍比人类研究品味更窄,并存在系统性偏移。It is suggested that strong LLMs can produce a range of reasonable ideas, but that range remains narrower than, and systematically shifted relative to, human research taste.

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering
MultAttnAttrib:长文档问答中的免训练多模态归因
arXiv:2607.01420 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本工作提出 MultAttnAttrib,一种免训练的归因生成方法,利用模型的预填充过程、选定的注意力头以及校准阈值在文档中定位源证据,且在多种归因生成方法上一致地表现更优。This work introduces MultAttnAttrib, a training-free attribution-generation method that leverages a model's prefill pass, selected attention heads, and calibrated thresholds to locate source evidence within a document, and consistently outperforms a variety of attribution-generation methods.

Multi-Turn Agentic Scientific Literature Search via Workflow Induction
基于工作流归纳的多轮智能体科学文献搜索
arXiv:2607.00597 Agent 智能体 方法 OA · 绿色 被引 2 · S2

结果表明,显式且可编辑的搜索工作流为将文献搜索智能体与复杂科学意图对齐提供了有效且可控的接口。The results show that explicit, editable search workflows provide an effective and controllable interface for aligning literature search agents with complex scientific intent.

Teaching LLMs to Recommend and Defer in Underrepresented Epilepsy Care
教导 LLM 在欠发达的癫痫诊疗中进行推荐与转诊
arXiv:2606.31036 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

在资源受限环境中,专业癫痫专家稀缺,使基于 LLM 的决策支持对管理纵向治疗的一线临床医生具有吸引力。此类系统必须适应当地处方实践并知道何时转诊。我们在乌干达儿科癫痫诊疗中研究该问题,基于纵向非结构化门诊记录预测抗癫痫用药方案。标准提示与医生处方取得了一定程度的一致性,但神经科医生审查显示许多错误反映的是分布失校的处方默认值而非失败。Specialist epilepsy expertise is scarce in resource-constrained settings, making LLM-based decision support attractive for frontline clinicians managing longitudinal treatment. Such systems must adapt to local prescribing practice and know when to defer. We study this problem in Ugandan pediatric epilepsy care, predicting anti-seizure medication regimens from longitudinal unstructured clinic notes. Standard prompting achieves non-trivial agreement with physician prescriptions, but neurologist review shows that many errors reflect distribution-miscalibrated prescribing defaults rather than fail

Generated Contents Enrichment
生成内容增强
arXiv:2405.03650 多模态 方法 OA · 绿色 被引 1 · S2

本文提出一个联合训练的对抗框架,通过建模对象语义和对象间关系来增强场景图,并在 Visual Genome 数据集上以代理场景图增强指标、图像质量比较、定性示例与用户研究进行评估。A jointly trained adversarial framework is proposed that enriches scene graphs by modeling object semantics and inter-object relations and is evaluated with proxy scene graph enrichment metrics, image-quality comparisons, qualitative examples, and user studies on the Visual Genome dataset.

JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications
京东 Oxygen AI 商品中心(Oxygen AIIC)V1:以 LLM/VLM 为核心的工业级商品理解、管理与应用解决方案
arXiv:2606.28070 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

京东 Oxygen AI 商品中心(Oxygen AIIC):基于 LLM/VLM 的工业级商品知识生产与服务平台,已在大规模场景下取得可量化的收益。The JD Oxygen AI Item Center (Oxygen AIIC), an industrial-scale platform built on LLMs/VLMs for item-knowledge production and service, has delivered measurable gains at scale.

DataComp-VLM: Improved Open Datasets for Vision-Language Models
DataComp-VLM:面向视觉-语言模型的改进开源数据集
arXiv:2606.28551 多模态 评测集 MPG.PuRe (Max Planck Society) OA · 绿色 被引 1 · S2

数据混合(而非过滤)是构建高质量训练数据集的关键:以指令型数据为主的混合在扩展时优于以描述型数据为主的混合,且规模越大优势越明显。It is found that data mixing, not filtering, is key to a high-quality training dataset: instruction-heavy mixtures scale better than caption-heavy ones, with gains widening at larger scales.

Transition-Aware best-of-N sampling for Longitudinal Chest X-ray Reports
面向纵向胸部 X 光报告的过渡感知 best-of-N 采样
arXiv:2606.28393 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

首个面向预训练胸部 X 光报告生成器的无需训练 best-of-N 采样方案,显式建模"既往-当前-过渡"纵向先验,全面优于随机选择。This work presents the first training-free best-of-N sampling scheme for pre-trained chest X-ray report generators that is explicitly aware of this longitudinal prior to current transition, and outperforms random selection across the board.

Know Your Source: A Public Knowledge Store for Media Background Checks
Know Your Source:面向媒体事实核查的公共知识库
arXiv:2607.02383 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 MEDIAREF:源自网络文档的公共知识库,支持跨 200 个媒体来源、可复现且低成本的 MBC 生成评估;给出可复现的构建与更新方法,并系统评测主流 LLM 在 MBC 生成任务上的表现。MEDIAREF, a publicly available knowledge store of web-sourced documents that enables reproducible, low-cost evaluation of MBC generation across 200 media sources, is introduced, describing a reproducible methodology for constructing and updating the collection, and assessing widely used LLMs on the MBC generation task.

A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory
A-TMA:解耦长时 Agent 记忆中的状态感知失效
arXiv:2607.01935 Agent 智能体 方法 OA · 绿色 被引 4 · S2

提出 ATMA:在现有记忆系统之上的状态感知叠加层,保留被替换记录与过渡记录,为查询所需的"目标状态视图"构建证据包,并向问答模块暴露当前、历史与过渡三类标签。This work proposes ATMA, a state aware overlay for existing memory systems, which keeps superseded and transition records in the bank, builds evidence packets for the query's requested state view, and exposes current, historical, and transition labels to QA.

CheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented Reasoning
CheckRLM:基于检索增强推理的知识-思维一致性校验
arXiv:2607.02262 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

CheckRLM:通过 RAG 及时校验并修正事实错误的框架,有效提升推理过程的可靠性,大幅超越现有基线。CheckRLM is a framework that improves the reliability of the reasoning process through Retrieval-Augmented Generation (RAG) by timely checking and correcting factual errors, and substantially outperforms existing baselines.

VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon
VLA-Corrector:面向自适应动作时域的轻量级检测-修正推理
arXiv:2607.01804 多模态 方法 OA · 绿色 被引 15 · S2

VLA-Corrector:面向动作分块 VLA 策略的轻量级修正推理框架;引入轻量的潜空间视觉监控器,持续比对预测与实际视觉特征演化,可在线检测视觉动态偏差,缓解静态时域在执行鲁棒性与策略调用频率之间的权衡。VLA-Corrector, a lightweight corrective inference framework for action-chunked VLA policies, introduces a lightweight Latent-space Vision Monitor that continuously compares predicted and actual visual feature evolution, enabling online detection of visual dynamics deviations and mitigates the trade-off imposed by static horizons between execution robustness and policy-call frequency.