研究库 论文知识库
Papers · organized/paper_cards

论文

1094 张论文卡片 · 方法 · OA 绿色

开放获取 全部 绿色 · 1640
What Does Privileged Information Add to On-Policy Self-Distillation?
特权信息为 On-Policy 自蒸馏带来了什么?
arXiv:2609.20612 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

构建了 AMPLE-Math,一个包含 5,319 道数学问题、每题对应六种共享同一答案的推理视角的可复用套件,并通过对比表明 OPSD 能借助直答推理与带思维链推理所共享的参数,提升对已有推理能力的调用效率。AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, is constructed and AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, is compared, suggesting that OPSD can improve access to existing reasoning capabilities through parameters shared by direct-response and thinking-enabled inference.

UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation
UFO:面向多模态图像生成全条件对齐的链式评估
arXiv:2609.12397 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 UFO,这是首个面向全条件对齐同时评估的统一框架;并发布 UFO-Bench,一个用于整体评估现有定制化模型在文本与视觉条件多样化交互下表现的专用基准。UFO is proposed, the first unified framework for omni-condition alignment simultaneous evaluation, and UFO-Bench is presented, a dedicated benchmark designed to holistically evaluate the performance of existing customization models under the diverse mutual interactions of textual and visual conditions.

When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
When2Think:面向高效混合推理模型的难度感知长度控制学习
arXiv:2609.19671 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 When2Think,一种基于 RLVR 的后训练框架,用于实例自适应计算分配,既无需学习奖励模型,也无需学习 critic;离线参考缓存机制避免了策略更新阶段对参考模型的在线查询。This work proposes When2Think, an RLVR-based post-training framework for instance-adaptive computation allocation that requires neither a learned reward model nor a learned critic, and offline reference caching avoids online reference-model queries during policy updates.

When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
当 EOS token 不一致时:理解 On-Policy 蒸馏中的长度膨胀
arXiv:2609.20511 LLM 基础设施 方法 OA · 绿色 被引 2 · S2

我们研究 on-policy 蒸馏 (OPD) 中的长度膨胀现象,即学生回答会变得过长,甚至耗尽生成预算。我们发现基础学生模型与训练后教师模型之间的终止 token 不匹配是该行为的重要来源。在 Qwen3、Llama 和 Gemma 上,两个模型可能将停止概率分配到不同的 EOS token 上,即便它们声明的停止集合相同。这种不匹配会抑制学生偏好的终止动作,同时无法可靠地传递教师偏好的替代动作。我们证明对齐 t...We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify termination-token mismatch between base students and post-trained teachers as an important source of this behavior. Across Qwen3, Llama, and Gemma, the two models can place their stopping probability on different EOS tokens, even when their declared stopping sets are identical. This mismatch can suppress the student's preferred termination action without reliably transferring the teacher-preferred alternative. We show that aligning t

FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations
FAMOS:基于稀疏观测的前馈 3D 关节建模
arXiv:2609.20817 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 FAMOS,一个前馈模型,可从稀疏、无序的部分点云集合预测可动部件分割与关节参数;并引入一个过程式数据生成器,在训练过程中合成自标注资产,以克服现有数据集规模和多样性的局限。FAMOS, a feed-forward model that predicts movable-part segmentation and joint parameters from a sparse, unordered set of partial point clouds, is presented and a procedural data generator that synthesizes self-annotated assets during training is introduced to overcome the limited scale and diversity of existing datasets.

1️⃣ arXiv · Learning Rate Matters: Vanilla LoRA May Suffice(⭐⭐⭐⭐⭐ 必读)
学习率至关重要:Vanilla LoRA 可能已足够
arXiv:2602.04998 评测基准 方法 Open MIND OA · 绿色 被引 11 · S2

本文通过大规模超参数搜索,系统地重新评估了 Vanilla LoRA 以及九个代表性 LoRA 变体,发现不同 LoRA 方法偏好的学习率区间各异,并将最优学习率区间的差异归因于最大 Hessian 特征值的变化,与经典学习理论相吻合。This work systematically re-evaluate nine representative LoRA variants alongside vanilla LoRA through extensive hyperparameter searches, finding that different LoRA methods favor distinct learning rate ranges and attributes the differing optimal learning rate ranges to variations in the largest Hessian eigenvalue, aligning with classical learning theories.

Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
样本数量远远不够:候选生成策略决定 LLM 测试时扩展的能耗与性能
arXiv:2609.19499 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,仅靠候选数量不足以刻画多候选 test-time scaling 的系统成本;评测应同时报告候选数量与准确率,以及生成调度和 GPU 层级的系统指标。The results show that candidate count alone is not enough to describe the systems cost of multi-candidate test-time scaling and Evaluations should report not only candidate count and accuracy, but also generation schedule and GPU-level systems metrics.

Verifiable Social Reasoning for LLM Assistants
面向 LLM 助手的可验证社会推理
arXiv:2609.17496 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出 Fuse——一个用于研究用户中介社会推理的多智能体仿真框架,并将其应用于 12 个 LLM,通过系统性隔离关键因素来展示其分析效用。This work introduces Fuse, a multi-agent simulation framework for studying user-mediated social reasoning, and applies Fuse to 12 LLMs and demonstrates its analytical utility by systematically isolating key factors.

Calibrating Teacher--Student Discrepancy for On-Policy Distillation
用于 On-Policy Distillation 的教师-学生差异校准
arXiv:2609.21619 工程化 方法 OA · 绿色 被引 2 · S2

提出 Calibrated On-Policy Distillation:通过正、负特权干预估计教师的自偏离区间,并仅保留超出该区间的部分以校准原始的 teacher–student 差异。Calibrated On-Policy Distillation is introduced, which estimates the teacher's self-deviation region through positive and negative privileged interventions and calibrates the original teacher--student discrepancy by retaining only the component that lies beyond this region.

Paint-Anything: Unified Any-Color Control for Image Generation and Editing
Paint-Anything:面向图像生成与编辑的统一任意颜色控制
arXiv:2609.20816 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Paint-Anything,通过物体级颜色监督学习一个统一的 hex-prompt 界面以同时支持生成与编辑;并引入 Any Color Benchmark (ACBench),包含 ACBench-T2I 和 ACBench-Edit,以衡量两项任务中的物体级 hex 颜色保真度。Paint-Anything is presented, which learns a shared hex-prompt interface for generation and editing through object-level color supervision, and introduces Any Color Benchmark (ACBench), comprising ACBench-T2I and ACBench-Edit, to measure object-level hex color fidelity across both tasks.

DeformSmith: Physics Harness-Guided Hierarchical Generation of Deformable Assets for Robot Manipulation
DeformSmith:基于物理 Harness 引导的可形变资产分层生成,用于机器人操作
arXiv:2609.18620 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,DeformSmith 生成的资产在视觉质量与物理合理性上均优于 SOTA 基线,包括 PhysGen3D、PhysGM 和 PhysX-Omni,同时支持为可形变物体的机器人操作合成训练数据。Results show that DeformSmith generates assets with better visual quality and physical plausibility than state-of-the-art baselines, including PhysGen3D, PhysGM, and PhysX-Omni, while supporting the synthesis of data for robotic manipulation of deformable objects.

Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation
基于信息瓶颈的训练自适应卷积稀疏编码,用于鲁棒视觉表征
arXiv:2609.19122 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文使用 Fast Iterative Shrinkage-Thresholding Algorithm 展开 CSC 优化,并将稀疏系数视为可微变量与网络参数联合学习;同时提出一种 label-free 的后训练策略,在固定主网络参数的情况下,根据被损坏输入自适应调整压缩强度。This work unfolds the CSC optimization with the Fast Iterative Shrinkage-Thresholding Algorithm and treats the sparsity coefficient as a differentiable variable jointly learned with the network parameters and introduces a label-free post-training strategy that adjusts the compression strength for corrupted inputs with the main network parameters fixed.

MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup
MoME:面向上下文感知稀疏查找的 Mixture-of-Memory Embeddings
arXiv:2609.15126 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Mixture of Memory Embeddings (MoME):一种上下文感知的记忆机制,将每个 token 的单一记忆行替换为 M 个 slot 的混合,并通过隐藏状态上的可学习门控选择在每个位置读取哪些 slot;在 sub-billion 规模下展现出更优的记忆容量 scaling 趋势,且训练与推理均保持高效。Mixture of Memory Embeddings (MoME), a context-aware memory mechanism that replaces each token's single memory row with a mixture of M slots and uses a learned gate over the hidden state to choose which slots to read at each position, shows more promising memory-size scaling trend at sub-billion scale and remains efficient in training and inference.

IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts
IntBMoE:将块级条件化融入专家组合的全参与 Mixture-of-Experts
arXiv:2609.21346 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

IntBMoE 是一种 block-conditioned MoE,将三者解耦,通过密集专家组合与稀疏 block 执行配对实现;在图像分类任务上,相较代表性的稀疏与密集 MoE 基线均取得稳定提升。IntBMoE is a block-conditioned MoE that decouples all three by pairing dense expert composition with sparse block execution, and shows consistent gains over representative sparse and dense MoE baselines on image classification.

Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network
精化本身即可编辑:基于生成式精化网络的无训练 Prompt-to-Prompt 图像编辑
arXiv:2609.20633 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

RefineEdit 是一个基于 GRN 的 training-free prompt-to-prompt 图像编辑框架,将 bit routing 与两种稳定机制(adaptive spatial freezing 与 finite bit locking)相结合,使编辑证据能够随图像演化而被修正。RefineEdit is a training-free prompt-to-prompt image editing framework built on the GRN that combines bit routing with two stabilization mechanisms: adaptive spatial freezing and finite bit locking, allowing editing evidence to be revised as the image evolves.

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
Designer-RSI:从用户流量中演化过程式记忆用于智能体平面设计
arXiv:2609.22086 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出一种持续自适应框架:冻结的前沿模型通过超过 230 个工具操作专业设计软件,同时外部的自然语言技能程序性记忆从经验中不断积累并精炼可复用的设计流程。A continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience is introduced.

13. TTKV:Temporal-Tiered KV Cache(HBM+DRAM 分层)
13. TTKV:Temporal-Tiered KV Cache(HBM+DRAM 分层)
arXiv:2604.19769 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

本文提出 TTKV,一种 KV cache 管理框架,将人类记忆系统映射到具备异构容量与精度的 KV cache 上,在 128K 上下文任务上将跨层流量降低 5.94 倍。TTKV is proposed, a KV cache management framework that maps the human memory system onto the KV cache with heterogeneous capacity and precision, and reduces cross-tier traffic by 5.94x on 128K-context tasks.

From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention
从预训练到精通:面向长时程操作的真实世界子任务强化学习,最小化人工介入
arXiv:2609.21788 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 PARTS(Policy Adaptation with RL on Targeted Subtasks),一种真实世界子任务强化学习框架,能够将训练集中于瓶颈环节,同时以最小的人工干预推进训练 rollout。PARTS (Policy Adaptation with RL on Targeted Subtasks), a real-world subtask RL framework that concentrates practice at bottlenecks while allowing training rollouts to proceed with minimal human intervention, is presented.

BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence
BI-Agent 与 BI-Bench:迈向端到端商业智能自动化
arXiv:2609.20886 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

设计了一个工具增强的 BI-Agent,将 BI 工作流分解为针对结构化数据的子任务(如 search、join、transform),并在各 BI 阶段编排专门的数据管理方法;与此同时开发了一套后训练框架,从真实 BI 项目中合成训练轨迹,使 BI-Agent 能够通过监督微调(SFT)和强化学习(RL)进行进一步后训练。A tool-augmented BI-Agent is designed that decomposes BI workflows into subtasks on structured data, such as search, join, and transform, and orchestrates specialized data management methods across BI stages, and a post-training framework is developed that synthesizes training trajectories from real BI projects, enabling BI-Agent to be further post-trained using both supervised fine-tuning (SFT) and reinforcement learning (RL).

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning
GAVEL:用于验证且高效的长视野 LLM 任务规划的图世界模型
arXiv:2609.19315 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

改进表明,显式的图世界模型调度层能够显著提升长程具身规划在紧凑型和前沿托管型 LLM 能力下的可靠性和效率。Improvements show that an explicit graph world model harness can substantially improve the reliability and efficiency of long-horizon embodied planning across compact and frontier hosted LLM capabilities.

SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory Desktops
SiliconBench:统一内存桌面上 LLM 服务的速度、内存与保真度
arXiv:2609.19169 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

在更大规模 dense 模型与 MoE 模型上的比较进一步凸显了在持续生成过程中调度 prompt 处理的重要性:在并发负载下,vllm-metal 的 packed prefill-decode 路径维持了低于 omlx 的首 token 延迟。Comparisons on larger dense and MoE models reinforce the importance of scheduling prompt processing alongside ongoing generation: vllm-metal's packed prefill-decode path maintains lower first-token latency than omlx under concurrent load.

SteerDuplex: Steerable Duplex Speech Dialogue Models
SteerDuplex:可操纵的全双工语音对话模型
arXiv:2609.12623 多模态 方法 OA · 绿色 被引 2 · S2

提出 SteerDuplex,一种基于 Moshi 的全双工语音模型,在自然对话和针对指令遵循、语音表达、推理与双工交互的合成对话上进行了微调,并通过两阶段混合奖励强化学习改善时序与回复连贯性。SteerDuplex is introduced, a Moshi-based full-duplex speech model fine-tuned on natural conversations and synthetic dialogues targeting instruction following, vocal delivery, reasoning, and duplex interaction, and two-stage reinforcement learning with hybrid rewards to improve timing and response continuity.

CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies
CARE:面向 Vision-Language-Action 策略的经验引导原子级纠错执行。
arXiv:2609.24118 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 CARE(Corrective Atomic Robotic Execution),一种通过从执行中遇到的失败进行学习以提升恢复能力的框架,并引入 Failure State Recovery Benchmark (FSR-Bench),用于评估在局部偏差与结构异常下从中间失败态恢复的表现。This work proposes CARE (Corrective Atomic Robotic Execution), a framework that improves recovery by learning from failures encountered during execution, and introduces the Failure State Recovery Benchmark (FSR-Bench), which evaluates recovery from intermediate failure states under local deviations and structural anomalies.

Grounded Action Model: 3D Grounding as a Foundation for Robotics
Grounded Action Model: 3D Grounding as a Foundation for Robotics
arXiv:2609.23863 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Grounded Action Models (GAMs),一种基于 3D grounding 构建的机器人基础模型新范式,可自主运行,并作为底层控制器由高层规划器通过多种输入模态进行控制,支持长时序与依赖记忆的操作。Grounded Action Models (GAMs), a new paradigm of robot foundation models built with 3D grounding, are proposed, which can be run autonomously and serve as a low-level controller that a high-level planner controls using its various input modalities, allowing for long-horizon and memory-dependent manipulation.

OmniEdu: Open Foundation Models for Learning and Teaching
OmniEdu: Open Foundation Models for Learning and Teaching
arXiv:2609.23088 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

证明了对通用语言模型适配教育任务(涵盖问题求解、课程理解与教学辅助)的精选、能力均衡的监督数据的价值。The value of curated, capability-balanced supervision for adapting general language models to educational tasks spanning problem solving, curriculum understanding, and instructional support is demonstrated.

Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion
Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion
arXiv:2609.24220 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Document Retrieval-Aware Chunking (D-RAC),将 Web Retrieval-Aware Chunking (W-RAC) 框架扩展至任意文档格式,保留 W-RAC 的成本、确定性与可观测性优势,同时将每种可渲染格式解锁为一类输入。Document Retrieval-Aware Chunking (D-RAC), an extension of the Web Retrieval-Aware Chunking (W-RAC) framework to arbitrary document formats, is presented, preserving W-RAC's cost, determinism, and observability benefits while unlocking every renderable format as a first-class input.

Streaming Video Editing with Easy Adaptation
Streaming Video Editing with Easy Adaptation
arXiv:2609.24788 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 SVEET 框架,仅需在预训练的双向视频扩散模型上训练,即可支持高质量自回归流式视频编辑,并提出解耦训练方案,显式强制视频可控性与模型因果性优化方向之间的正交性。This paper proposes SVEET, a framework that requires merely training on a pretrained bidirectional video diffusion model but supports high-quality streaming video editing in an auto-regressive fashion and proposes a decoupled training scheme that explicitly enforces the orthogonality between the optimization directions of video controllability and model causality.

Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention
Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention
arXiv:2609.24797 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

表明 Kimi Delta Attention (KDA) 可通过单次 delta-rule 变换与其通道门控提供的反射组合实现 2D 旋转,并刻画 CKDA 的表达能力,证明每个正交的对角加秩-1矩阵恰为一个 CKDA 转移矩阵。This work shows that Kimi Delta Attention (KDA) can realize 2D rotations by combining a single delta-rule transformation with a second reflection supplied by its channel-wise gate, and characterize the expressivity of CKDA and prove that every orthogonal diagonal-plus-rank-one matrix is exactly a CKDA transition matrix.

Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations
Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations
arXiv:2609.22255 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Deep Persona,一种心理学驱动的三层架构,将 persona 组织为可观察表达、潜在信念与核心动机驱动的层级结构,用于构建高可信度的角色扮演 Agent。Deep Persona is introduced, a psychologically grounded, three-layered architecture that organizes personas into hierarchical levels of observable expression, latent beliefs, and core motivational drives, for constructing highly convincing role-playing agents.

12. KVP:RL 驱动 KV Cache 驱逐策略
arXiv:2602.10238 LLM 基础设施 方法 Open MIND OA · 绿色 被引 12 · S2

本文提出 KV Policy (KVP),一种仅基于 key 与 value 向量、在预计算生成轨迹上训练的轻量级 per-head RL Agent 框架,证明学习预测未来 token 效用是自适应 KV cache 管理中强大且可扩展的范式。KV Policy (KVP), a framework of lightweight per-head RL agents trained on pre-computed generation traces using only key and value vectors, is introduced, demonstrating that learning to predict future token utility is a powerful and scalable paradigm for adaptive KV cache management.

Towards Full Pipeline FP8 Reinforcement Learning for LLMs
面向 LLM 的全流水线 FP8 强化学习
arXiv:2609.22870 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Calibrated Clipping,一种动态方法,通过匹配下界裁剪分位数并相应重平衡上界,将 FP8 裁剪边界与高精度 BF16 分布对齐,消除熵激增并恢复与 BF16 基线可比的性能。Calibrated Clipping is proposed, a dynamic method that aligns the FP8 clipping bounds with high-precision BF16 distributions by matching the lower-bound clipping quantile and rebalancing the upper bound accordingly, which eliminates entropy surges and restores performance comparable to the BF16 baseline.

UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing
UltraTex:释放 2K 多视角扩散模型用于 3D 纹理生成
arXiv:2609.23169 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 UltraTex,一种用于高分辨率多视图扩散 3D 纹理生成的高效端到端框架,并引入 Background Token Dropping(在 DiT 主干前移除背景 token)与 Block-Sparse Attention(降低前景序列上的注意力计算)。This paper presents UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing and introduces Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequence.

A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal
大语言模型的测谎测试:读出模型不愿透露的知识
arXiv:2609.21996 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

内部识别探针(Probe of Internal Recognition, PIR)可区分"不愿回答"与"无法回答"的模型,支持装傻审计与遗忘验证,并从选择题扩展至自由生成。Probe of Internal Recognition (PIR) separates a model that will not answer from one that cannot, which supports sandbagging audits and unlearning verification, and extends from multiple-choice questions to free-form generation.

TAPe+ML: A Compact Structured Representation for Multi-Task Computer Vision
TAPe+ML:面向多任务计算机视觉的紧凑结构化表示
arXiv:2609.20869 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,将部分建模负担从网络参数转移至结构化输入表示,可在更低的数据、内存与算力需求下支持紧凑的多任务视觉系统。The results suggest that shifting part of the modeling burden from network parameters to a structured input representation can support compact multi-task vision systems with reduced data, memory, and compute requirements.

SkillSpec: Intent-Masked Specification Reasoning for Agent Skill Correctness
SkillSpec:面向 Agent Skill 正确性的意图掩码规约推理
arXiv:2609.06052 Agent 智能体 方法 OA · 绿色 被引 1 · S2

提出 SkillSpec,一种 Hoare 风格框架,将技能正确性建模为规约推理问题,并将异构技能仓库转化为统一的图表示,对齐描述、指令与代码构件。This work proposes SkillSpec, a Hoare-style framework that formulates skill correctness as a specification reasoning problem, and transforms a heterogeneous skill repository into a unified graph representation that aligns descriptions, instructions and code artifacts.

StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training
StableVQ:稳定的 Vector-Quantized Tokenizer 训练实用指南
arXiv:2609.26774 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 StableVQ,重新审视每个模块的合适学习目标,解决各模块独立训练以承担各自角色时产生的问题;其基于共享投影 codebook 构建,轻量且不引入可学习参数。StableVQ is proposed, which revisits the proper learning objective of each module and resolves the problems that arise when each is trained to fulfill its own role independently, and Built on top of shared-projection codebooks, is lightweight and introduces no learnable parameters.