研究库 论文知识库
Papers · organized/paper_cards

论文

1640 张论文卡片 · OA 绿色

开放获取 全部 绿色 · 1640
VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
VākQA:泰卢固语口语事实型问答基准与评估研究
arXiv:2609.19879 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 V\={a}kQA,一个覆盖六个领域、包含 2,001 对事实型问答的泰卢固语 SQA 基准;并观察到:泰卢固语措辞保留了翻译中会丢失的文化特异性,语音输入引入的音近混淆会改变问题含义,级联 ASR-MT 误差会逐步叠加放大。V\={a}kQA, a Telugu SQA benchmark of 2,001 factoid question-answer pairs across six domains, is introduced and it is observed that Telugu phrasing retains cultural specificity that is lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively.

Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model
MiniMax-H3 能否对物理世界进行推理?一种全模态生成模型的评估
arXiv:2609.18323 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个围绕物理世界推理四个互补维度构建的综合评估框架,并构建了一系列多样化新任务,要求模型跨模态整合互补信息。This work introduces a comprehensive evaluation framework organized around four complementary dimensions of physical world reasoning, and constructs a diverse set of novel tasks that require models to integrate complementary information across modalities.

1️⃣ arXiv · Learning Rate Matters: Vanilla LoRA May Suffice(⭐⭐⭐⭐⭐ 必读)
学习率至关重要:Vanilla LoRA 可能已足够
arXiv:2602.04998 评测基准 方法 Open MIND OA · 绿色 被引 11 · S2

本文通过大规模超参数搜索,系统地重新评估了 Vanilla LoRA 以及九个代表性 LoRA 变体,发现不同 LoRA 方法偏好的学习率区间各异,并将最优学习率区间的差异归因于最大 Hessian 特征值的变化,与经典学习理论相吻合。This work systematically re-evaluate nine representative LoRA variants alongside vanilla LoRA through extensive hyperparameter searches, finding that different LoRA methods favor distinct learning rate ranges and attributes the differing optimal learning rate ranges to variations in the largest Hessian eigenvalue, aligning with classical learning theories.

Srijika: OpenType-Layout-Reusing Font Restyling for Nine Indic Scripts
Srijika:面向九种印度文字的 OpenType 布局复用字体再设计
arXiv:2609.05661 工程化 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Srijika,一个为九种 Brahmic 文字(Devanagari、Tamil、Bengali、Telugu、Kannada、Malayalam、Gujarati、Gurmukhi、Odia)生成可安装 OpenType 字体的系统,并附带一份负面结果目录,覆盖参考引导重风格化中失败的 conditioning、目标函数选择和数据凸包限制。Srijika is presented, a system for producing installable OpenType fonts for nine Brahmic scripts: Devanagari, Tamil, Bengali, Telugu, Kannada, Malayalam, Gujarati, Gurmukhi, and Odia, and a negative-results catalogue covering failed conditioning, objective choices, and data-hull limits of reference-guided restyling.

Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
样本数量远远不够:候选生成策略决定 LLM 测试时扩展的能耗与性能
arXiv:2609.19499 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,仅靠候选数量不足以刻画多候选 test-time scaling 的系统成本;评测应同时报告候选数量与准确率,以及生成调度和 GPU 层级的系统指标。The results show that candidate count alone is not enough to describe the systems cost of multi-candidate test-time scaling and Evaluations should report not only candidate count and accuracy, but also generation schedule and GPU-level systems metrics.

Verifiable Social Reasoning for LLM Assistants
面向 LLM 助手的可验证社会推理
arXiv:2609.17496 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出 Fuse——一个用于研究用户中介社会推理的多智能体仿真框架,并将其应用于 12 个 LLM,通过系统性隔离关键因素来展示其分析效用。This work introduces Fuse, a multi-agent simulation framework for studying user-mediated social reasoning, and applies Fuse to 12 LLMs and demonstrates its analytical utility by systematically isolating key factors.

Calibrating Teacher--Student Discrepancy for On-Policy Distillation
用于 On-Policy Distillation 的教师-学生差异校准
arXiv:2609.21619 工程化 方法 OA · 绿色 被引 2 · S2

提出 Calibrated On-Policy Distillation:通过正、负特权干预估计教师的自偏离区间,并仅保留超出该区间的部分以校准原始的 teacher–student 差异。Calibrated On-Policy Distillation is introduced, which estimates the teacher's self-deviation region through positive and negative privileged interventions and calibrates the original teacher--student discrepancy by retaining only the component that lies beyond this region.

Paint-Anything: Unified Any-Color Control for Image Generation and Editing
Paint-Anything:面向图像生成与编辑的统一任意颜色控制
arXiv:2609.20816 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Paint-Anything,通过物体级颜色监督学习一个统一的 hex-prompt 界面以同时支持生成与编辑;并引入 Any Color Benchmark (ACBench),包含 ACBench-T2I 和 ACBench-Edit,以衡量两项任务中的物体级 hex 颜色保真度。Paint-Anything is presented, which learns a shared hex-prompt interface for generation and editing through object-level color supervision, and introduces Any Color Benchmark (ACBench), comprising ACBench-T2I and ACBench-Edit, to measure object-level hex color fidelity across both tasks.

TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection
TeleAntiFraud 2.0:面向电信诈骗检测的可刷新、基于画像的音频基准
arXiv:2609.18748 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 TeleAntiFraud 2.0,采用 Mixed-Tree Anti-Fraud Generation Pipeline 构建,并基于月度冻结评测协议进行评估,确立了近域构造与 collapse-aware 报告作为在现实易混淆条件下评测音频电信诈骗模型的核心要求。This work presents TeleAntiFraud 2.0, constructed with the Mixed-Tree Anti-Fraud Generation Pipeline and evaluated under a monthly frozen evaluation protocol, establishing near-domain construction and collapse-aware reporting as core requirements for evaluating audio-based telecom-fraud models under realistic confusable conditions.

DeformSmith: Physics Harness-Guided Hierarchical Generation of Deformable Assets for Robot Manipulation
DeformSmith:基于物理 Harness 引导的可形变资产分层生成,用于机器人操作
arXiv:2609.18620 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,DeformSmith 生成的资产在视觉质量与物理合理性上均优于 SOTA 基线,包括 PhysGen3D、PhysGM 和 PhysX-Omni,同时支持为可形变物体的机器人操作合成训练数据。Results show that DeformSmith generates assets with better visual quality and physical plausibility than state-of-the-art baselines, including PhysGen3D, PhysGM, and PhysX-Omni, while supporting the synthesis of data for robotic manipulation of deformable objects.

Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation
基于信息瓶颈的训练自适应卷积稀疏编码,用于鲁棒视觉表征
arXiv:2609.19122 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文使用 Fast Iterative Shrinkage-Thresholding Algorithm 展开 CSC 优化,并将稀疏系数视为可微变量与网络参数联合学习;同时提出一种 label-free 的后训练策略,在固定主网络参数的情况下,根据被损坏输入自适应调整压缩强度。This work unfolds the CSC optimization with the Fast Iterative Shrinkage-Thresholding Algorithm and treats the sparsity coefficient as a differentiable variable jointly learned with the network parameters and introduces a label-free post-training strategy that adjusts the compression strength for corrupted inputs with the main network parameters fixed.

FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection
FRAUDSkill:面向音频反诈骗检测的结构化冻结权重 Skill 优化
arXiv:2609.18766 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出 FRAUDSkill——一种结构化的 frozen-weight 适配框架:底层音频-语言模型保持不变,转而优化外部的 skill program、路由策略和决策规则,并将结构化输出控制与验证引导的多路径推理相结合,以确保符合协议规范的预测。FRAUDSkill is proposed, a structured frozen-weight adaptation framework that leaves the underlying audio-language model unchanged while optimizing an external layer of skill programs, route-specific policies, and decision rules and combines structured output control with validation-guided multi-path inference to ensure protocol-compliant predictions.

MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup
MoME:面向上下文感知稀疏查找的 Mixture-of-Memory Embeddings
arXiv:2609.15126 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Mixture of Memory Embeddings (MoME):一种上下文感知的记忆机制,将每个 token 的单一记忆行替换为 M 个 slot 的混合,并通过隐藏状态上的可学习门控选择在每个位置读取哪些 slot;在 sub-billion 规模下展现出更优的记忆容量 scaling 趋势,且训练与推理均保持高效。Mixture of Memory Embeddings (MoME), a context-aware memory mechanism that replaces each token's single memory row with a mixture of M slots and uses a learned gate over the hidden state to choose which slots to read at each position, shows more promising memory-size scaling trend at sub-billion scale and remains efficient in training and inference.

MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads
MLLMs 在 Synergy Heads 中信息分布漂移时产生幻觉
arXiv:2609.09206 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

HEAL 将动态信息校准因子注入协同注意力头的 value 向量中,主动调节视觉-语言依赖,引导输出分布趋向事实证据,提供了一条简单且可解释的增强模型可信度的路径。HEAL injects dynamic information calibration factors into the value vectors of synergy heads and actively regulates visual-language dependencies, steering the output distribution towards factual evidence, offering a simple and interpretable pathway to enhance model trustworthiness.

IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts
IntBMoE:将块级条件化融入专家组合的全参与 Mixture-of-Experts
arXiv:2609.21346 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

IntBMoE 是一种 block-conditioned MoE,将三者解耦,通过密集专家组合与稀疏 block 执行配对实现;在图像分类任务上,相较代表性的稀疏与密集 MoE 基线均取得稳定提升。IntBMoE is a block-conditioned MoE that decouples all three by pairing dense expert composition with sparse block execution, and shows consistent gains over representative sparse and dense MoE baselines on image classification.

OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
OmniVChat:面向原生音视频对话的合成、基准与训练
arXiv:2609.21465 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

OmniVChat-Studio,一种用于合成单轮与多轮音视频对话的多 Agent 数据引擎;以及 OmniVChat-RL,一种在 OmniVChat 中同时针对回复正确性、效率与风格设计奖励的强化学习奖励方案。OmniVChat-Studio, a multi-agent data engine for synthesizing single- and multi-turn audio-visual dialogues and OmniVChat-RL, a reinforcement learning reward design that jointly targets reply correctness, efficiency, and style in OmniVChat.

APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport
APort Vault:基于 Open Agent Passport 的 AI Agent 支付授权基准
arXiv:2609.22076 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

APort Vault 是面向工具使用 AI Agent 支付授权的基准。它在公开 CTF 活动中重放了人类编写的 4,371 次针对真实支付 Agent 的攻击,覆盖 8 家实验室的 14 个模型、五种策略配置与两条重放轨道,每条轨道分别在有/无确定性的 pre-action check(实现 Open Agent Passport (OAP) 规范)下执行,共计完成 225,964 次评测。我们每次评测报告五个独立事件,因为将它们合并正是 Agent 基准产生无法经得起审查的数字的方式。请求很常见,且其速率差异巨大APort Vault is a benchmark for payment authorization in tool-using AI agents. It replays 4,371 attacks written by humans against a live payment agent during a public capture-the-flag event, across 14 models from 8 labs, five policy configurations and two replay tracks, with and without a deterministic pre-action check implementing the Open Agent Passport (OAP) specification. 225,964 evaluations completed. We report five distinct events per evaluation, because collapsing them is how an agent benchmark produces a number that does not survive review. Requests are common and their rate differs far

Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network
精化本身即可编辑:基于生成式精化网络的无训练 Prompt-to-Prompt 图像编辑
arXiv:2609.20633 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

RefineEdit 是一个基于 GRN 的 training-free prompt-to-prompt 图像编辑框架,将 bit routing 与两种稳定机制(adaptive spatial freezing 与 finite bit locking)相结合,使编辑证据能够随图像演化而被修正。RefineEdit is a training-free prompt-to-prompt image editing framework built on the GRN that combines bit routing with two stabilization mechanisms: adaptive spatial freezing and finite bit locking, allowing editing evidence to be revised as the image evolves.

Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models
价值几何:语言模型伦理偏好对齐中的任务向量组合
arXiv:2609.21094 安全与风险 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

一项基于任务向量迁移的实验:在计算了价值偏好方向对应的任务向量后,将其相对于通用指令跟随向量进行正交化处理,该方法能够有效隔离出特定价值偏好的方向,从而通过任务算术获得具有相反立场的模型。A task vector transfer based experiment where after computing the task vectors for a direction of value preference the authors orthogonalize it with respect to the general instruction following vector shows that this method is effective in isolating the direction of the specific value preference that can successfully be used to conduct task arithmetic to obtain a model with the opposite stance.

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
Designer-RSI:从用户流量中演化过程式记忆用于智能体平面设计
arXiv:2609.22086 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出一种持续自适应框架:冻结的前沿模型通过超过 230 个工具操作专业设计软件,同时外部的自然语言技能程序性记忆从经验中不断积累并精炼可复用的设计流程。A continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience is introduced.

Gricea: An Open Science Platform for Conversational AI Research
Gricea:一个面向对话式 AI 研究的开放科学平台
arXiv:2609.22039 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Gricea,一个开放科学平台,将研究表示为可配置、可部署的研究 artifact,研究者可以运行、检查、共享和复用;该平台展示了如何通过共享的研究 artifact 构建、复现和扩展 CAI 研究,从而借助开放科学实现知识的累积构建。This work presents Gricea, an open-science platform representing studies as configurable, deployable research artifacts that researchers can run, inspect, share, and reuse, demonstrating Gricea's support for constructing, reproducing, and extending CAI studies through shared research artifacts, enabling cumulative knowledge building through open science.

13. TTKV:Temporal-Tiered KV Cache(HBM+DRAM 分层)
13. TTKV:Temporal-Tiered KV Cache(HBM+DRAM 分层)
arXiv:2604.19769 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

本文提出 TTKV,一种 KV cache 管理框架,将人类记忆系统映射到具备异构容量与精度的 KV cache 上,在 128K 上下文任务上将跨层流量降低 5.94 倍。TTKV is proposed, a KV cache management framework that maps the human memory system onto the KV cache with heterogeneous capacity and precision, and reduces cross-tier traffic by 5.94x on 128K-context tasks.

From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention
从预训练到精通:面向长时程操作的真实世界子任务强化学习,最小化人工介入
arXiv:2609.21788 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 PARTS(Policy Adaptation with RL on Targeted Subtasks),一种真实世界子任务强化学习框架,能够将训练集中于瓶颈环节,同时以最小的人工干预推进训练 rollout。PARTS (Policy Adaptation with RL on Targeted Subtasks), a real-world subtask RL framework that concentrates practice at bottlenecks while allowing training rollouts to proceed with minimal human intervention, is presented.

BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence
BI-Agent 与 BI-Bench:迈向端到端商业智能自动化
arXiv:2609.20886 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

设计了一个工具增强的 BI-Agent,将 BI 工作流分解为针对结构化数据的子任务(如 search、join、transform),并在各 BI 阶段编排专门的数据管理方法;与此同时开发了一套后训练框架,从真实 BI 项目中合成训练轨迹,使 BI-Agent 能够通过监督微调(SFT)和强化学习(RL)进行进一步后训练。A tool-augmented BI-Agent is designed that decomposes BI workflows into subtasks on structured data, such as search, join, and transform, and orchestrates specialized data management methods across BI stages, and a post-training framework is developed that synthesizes training trajectories from real BI projects, enabling BI-Agent to be further post-trained using both supervised fine-tuning (SFT) and reinforcement learning (RL).

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning
GAVEL:用于验证且高效的长视野 LLM 任务规划的图世界模型
arXiv:2609.19315 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

改进表明,显式的图世界模型调度层能够显著提升长程具身规划在紧凑型和前沿托管型 LLM 能力下的可靠性和效率。Improvements show that an explicit graph world model harness can substantially improve the reliability and efficiency of long-horizon embodied planning across compact and frontier hosted LLM capabilities.

CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design
CADWorld:面向长视野计算机辅助设计的计算机使用基准
arXiv:2609.16251 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

研究发现,较弱的 Agent 常常在产出有效成果之前即告失败,而较强的 Agent 则越来越多地在结构、几何与构造过程要求上失败;CADWorld 揭示了通用 GUI 能力与可靠执行持久、可验证工程工作流之间的差距。It is found that weaker agents often fail before producing a valid artifact, whereas stronger agents increasingly fail on structural, geometric, and construction-process requirements, and CADWorld exposes a gap between general GUI competence and reliable execution of persistent, verifiable engineering workflows.

SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory Desktops
SiliconBench:统一内存桌面上 LLM 服务的速度、内存与保真度
arXiv:2609.19169 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

在更大规模 dense 模型与 MoE 模型上的比较进一步凸显了在持续生成过程中调度 prompt 处理的重要性:在并发负载下,vllm-metal 的 packed prefill-decode 路径维持了低于 omlx 的首 token 延迟。Comparisons on larger dense and MoE models reinforce the importance of scheduling prompt processing alongside ongoing generation: vllm-metal's packed prefill-decode path maintains lower first-token latency than omlx under concurrent load.

SteerDuplex: Steerable Duplex Speech Dialogue Models
SteerDuplex:可操纵的全双工语音对话模型
arXiv:2609.12623 多模态 方法 OA · 绿色 被引 2 · S2

提出 SteerDuplex,一种基于 Moshi 的全双工语音模型,在自然对话和针对指令遵循、语音表达、推理与双工交互的合成对话上进行了微调,并通过两阶段混合奖励强化学习改善时序与回复连贯性。SteerDuplex is introduced, a Moshi-based full-duplex speech model fine-tuned on natural conversations and synthetic dialogues targeting instruction following, vocal delivery, reasoning, and duplex interaction, and two-stage reinforcement learning with hybrid rewards to improve timing and response continuity.

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
GameHorizon Suite:游戏中多时间历程的数据与评估。
arXiv:2609.25001 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

GameHorizon Suite 是一套统一的数据与评估套件,可在多个时间跨度下衡量不同模型族的游戏能力,为跨水平跨度与跨模型族的游戏能力评估提供标准化标尺。The GameHorizon Suite, a unified data and evaluation suite that measures gameplay capabilities at different horizons for diverse model families, can provide a standardized yardstick for evaluating gameplay capabilities across horizons and model families.

onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction
onPanda:通过 token 级校正高效标注 LLM 与 Agent 的 on-policy 对齐数据。
arXiv:2609.24983 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出 OnPanda,一种用于高效标注 LLM 对齐数据与 Agent 轨迹的交互式工具,以 token 级修正为核心交互方式;并发布使用 onPanda 标注的数据集 Panda-CVL,以及一个面向 token 级修正的基准。OnPanda is presented, an interactive tool for efficiently annotating LLM alignment data and agent trajectories that adopts token-level correction as its core interaction and releases Panda-CVL, a dataset annotated with onPanda, together with a benchmark for token-level correction.

1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation
1% 的 token 已足够:论 On-Policy 蒸馏中的梯度估计。
arXiv:2609.24432 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

基于信噪分解与候选集近似的 IER(信息效率比),可基于 IER 及其与现有效用分数的组合进行 token 选择,同时保留采样的反向 KL 训练目标。An information-efficiency ratio (IER) based on a signal-to-noise decomposition and a candidate-set approximation enables token selection based on IER and its combination with existing usefulness scores, while retaining the sampled reverse-KL training objective.

12. SoK: Agentic RAG(arXiv 2603.07379,ACL 2026)
12. SoK:Agentic RAG(arXiv 2603.07379,ACL 2026)
arXiv:2603.07379 RAG 检索增强 观点 Open MIND OA · 绿色 被引 8 · S2

本文将 Agentic 检索-生成循环形式化为有限时域部分可观测马尔可夫决策过程,显式建模其控制策略与状态转移,并构建了全面的分类体系与模块化架构分解,按规划机制、检索编排、记忆范式与工具调用行为对系统进行分类。This paper formalizes agentic retrieval-generation loops as finite-horizon partially observable Markov decision processes, explicitly modeling their control policies and state transitions, and develops a comprehensive taxonomy and modular architectural decomposition that categorizes systems by their planning mechanisms, retrieval orchestration, memory paradigms, and tool-invocation behaviors.

CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies
CARE:面向 Vision-Language-Action 策略的经验引导原子级纠错执行。
arXiv:2609.24118 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 CARE(Corrective Atomic Robotic Execution),一种通过从执行中遇到的失败进行学习以提升恢复能力的框架,并引入 Failure State Recovery Benchmark (FSR-Bench),用于评估在局部偏差与结构异常下从中间失败态恢复的表现。This work proposes CARE (Corrective Atomic Robotic Execution), a framework that improves recovery by learning from failures encountered during execution, and introduces the Failure State Recovery Benchmark (FSR-Bench), which evaluates recovery from intermediate failure states under local deviations and structural anomalies.

Grounded Action Model: 3D Grounding as a Foundation for Robotics
Grounded Action Model: 3D Grounding as a Foundation for Robotics
arXiv:2609.23863 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Grounded Action Models (GAMs),一种基于 3D grounding 构建的机器人基础模型新范式,可自主运行,并作为底层控制器由高层规划器通过多种输入模态进行控制,支持长时序与依赖记忆的操作。Grounded Action Models (GAMs), a new paradigm of robot foundation models built with 3D grounding, are proposed, which can be run autonomously and serve as a low-level controller that a high-level planner controls using its various input modalities, allowing for long-horizon and memory-dependent manipulation.

OmniEdu: Open Foundation Models for Learning and Teaching
OmniEdu: Open Foundation Models for Learning and Teaching
arXiv:2609.23088 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

证明了对通用语言模型适配教育任务(涵盖问题求解、课程理解与教学辅助)的精选、能力均衡的监督数据的价值。The value of curated, capability-balanced supervision for adapting general language models to educational tasks spanning problem solving, curriculum understanding, and instructional support is demonstrated.

Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles
Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles
arXiv:2609.22220 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出将变异分析作为 kernel-benchmark 预言机的充分性度量:通过确定性规则向 188 个 KernelBench 问题的已验证 CUDA 实现中注入 10 个可编译故障,其中 7,384 个具备独立击杀见证;任何测试协议均按其检出比例评分。Mutation analysis as an adequacy metric for kernel-benchmark oracles is introduced: deterministic rules inject 10 compilable faults into verified CUDA implementations of 188 KernelBench problems, 7{,}384 of them with an independent kill witness; any test protocol is scored by the fraction it detects.