研究库 论文知识库
Papers · organized/paper_cards

论文

1640 张论文卡片 · OA 绿色

开放获取 全部 绿色 · 1640
FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation
FLEET:从 logits 熵到文本生成中的增强轨迹
arXiv:2609.27657 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

FLEET 将每次生成表示为穿越状态的稀疏轨迹(状态的熵超过预设阈值),并基于这些轨迹推断逐 token 的效用分数以调整 logits,是一种将 memory 机制融入生成过程的新方法。FLEET represents each generation as a sparse trajectory through states whose entropy exceeds a predefined threshold and uses these trajectories to infer per-token utility scores that adjust the logits, a novel method that integrates a memory mechanism into the generation process.

Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery
减少六层:面向 Whisper 的无标签恢复编码器剪枝
arXiv:2609.27980 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种方法,通过逐层剔除后 WER(Word Error Rate)的变化对编码器层进行排序,并给出剪枝后的模型——即一个层数更少的更浅编码器。This work presents an approach that ranks encoder layers by the leave-one-layer-out change in Word Error Rate (WER), and presents the pruned model, which is simply a more shallow encoder with fewer layers.

Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI
Uranus:为具身 AI 构建下一代仿真基础设施
arXiv:2609.24815 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Uranus,一个以关节轨迹为条件的自回归扩散模型为核心的数据驱动机器人仿真器,为不同机器人本体与相机配置下的同步多视图生成提供统一接口。This work presents Uranus, a data-driven robot simulator built around a joint-trajectory-conditioned autoregressive diffusion model, providing a unified interface for synchronized multi-view generation across diverse robot embodiments and camera configurations.

GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression
GeoPair:面向免训练 Transformer 压缩的几何保持跨层分解
arXiv:2609.25963 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个基于原理的、无需训练的框架,依次优化跨层权重配对与共享字典分解,并识别结构兼容的投影、学习一种能更好保留各层独立校准几何的共享表征。This work introduces a principled, training-free framework that sequentially optimizes cross-layer weight pairings and shared-dictionary factorizations, and identifies structurally compatible projections and learns a shared representation that better preserves each layer's distinct calibration geometry.

Calibration as a First-Class Criterion in LLM Evaluation
将校准作为 LLM 评估中的一等标准
arXiv:2609.26489 评测基准 评测集 OA · 绿色 被引 1 · S2

本文主张每个 NLP 子领域应将主要性能指标与一个校准分数配对,呼吁将校准视为每个模型的基本属性而非边缘话题。It is argued that each NLP subfield should pair its main performance metric with a calibration score and call for treating calibration as an essential property of every model rather than a niche topic.

MemoryAthena: Adaptive Routing over Latent and Generated Memories
MemoryAthena:基于潜在记忆与生成记忆的自适应路由
arXiv:2609.25853 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果支持将生成式记忆视为对直接检索的选择性修正,并强调何时、以何种方式、以何种强度进行路由干预是核心挑战。The results support generated memory as a selective correction to direct retrieval and highlight routing when, which, and how strongly to intervene as the central challenge.

X-Planner: Event-Structured Task Planning for Embodied Intelligence
X-Planner:面向具身智能的事件结构化任务规划
arXiv:2609.25187 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 X-Planner,一个面向具身推理的规划前端,同时解决监督与表征问题,并描述了规划文本质量与下游执行情况。This work presents X-Planner, a planning front-end that addresses both the supervision and representation of embodied reasoning, and describes planning-text quality and downstream execution.

Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models
能力出众却简洁高效:提取并刻画前沿模型中隐藏的思维链
arXiv:2609.26637 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果发现 Astra 表现出 token 高效的有向推理,能更早选择正确轨迹,在内部解决基础步骤,仅外化关键推理,为前沿模型推理提供了超越基准分数的行为视角。It is found that Astra exhibits token-efficient directed reasoning, selecting a correct trajectory earlier, while resolving elementary steps internally and externalizing only crucial reasoning, which provides a behavioral lens on frontier-model reasoning beyond benchmark scores.

The Linear Representation Hypothesis Needs a Group Action
线性表示假说需要一个群作用
arXiv:2609.27158 LLM 基础设施 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文论证线性表征假设不是单一假设,而是一组通过表征等价性加以区分的命题族;该框架厘清了假设如何随度量、读取点与分析阶段变化,并被用于审计常见表征量与近期可解释性分析。It is argued that the Linear Representation Hypothesis is not one hypothesis but a family of claims distinguished by representation equivalence, and this framework clarifies how assumptions can change across metrics, reading points, and analysis stages, and is used to audit common representation quantities and recent interpretability analyses.

Knowledge Pull Requests for Continual Document Authoring
面向持续文档撰写的知识 Pull Request
arXiv:2609.26634 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Knowledge Pull Requests 相比从来源重写或从零重新生成,能整合更多信息并更好地保留已有内容,同时每生成一个 token 增加的信息量最多。Knowledge Pull Requests integrate more information and better preserve existing content than rewriting from sources or regenerating from scratch, while adding the most information per token generated.

1. vLLM Startup Latency: Six-Step Systematic Characterization
1. vLLM 启动延迟:六步式系统化表征
arXiv:2606.07362 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

本文首次对 vLLM 启动延迟进行了详细的性能表征,并构建了一个轻量级分析模型,能够针对给定硬件配置准确预测 vLLM 的启动延迟,为大规模推理环境中的资源规划提供了可操作的指导。This paper presents the first detailed performance characterization of vLLM startup latency and develops a lightweight analytical model that accurately predicts vLLM's startup latency for a given hardware configuration, providing actionable guidance for resource planning in large-scale inference environments.

Self-Organizing Agent Teams Learn to Reason Together
自组织 Agent 团队学习协同推理
arXiv:2609.22682 Agent 智能体 观点 OA · 绿色 被引 1 · S2

本文提出 Self-Organizing Agent Teams (SAT),即从先前协作中学习可复用策略、固定成员队伍的 AI 代理,用以组织角色、对话阶段、参与方式与信息流,表明组织本身可成为代理的一项能力。Self-Organizing Agent Teams (SAT) are introduced, fixed teams of AI agents that learn reusable strategies from prior collaborations to organize roles, conversational phases, participation, and information flow, suggesting that organization itself can become an agent capability.

IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis
IterSynth:通过角色解耦的迭代合成重新思考深度搜索 Agent
arXiv:2609.29444 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 IterSynth,一种角色解耦、基于摘要的范式,在用于识别信息需求的 Planner 与用于将证据整合进演化摘要状态的 Synthesizer 之间交替,作为一种模型无关的提示范式,在前沿闭源模型上相对 ReAct 及类似提示范式取得显著的零样本增益。IterSynth is proposed, a role-decoupled and summary-based paradigm that alternates between a Planner for identifying information needs and a Synthesizer for integrating evidence into an evolving summary state, and serves as a model-agnostic prompting paradigm, delivering substantial zero-shot gains over ReAct and similar prompting paradigms on frontier proprietary models.

PUBG Ally: A Conversational Embodied Agent as an AI Teammate
PUBG Ally:作为 AI 队友的对话式具身 Agent
arXiv:2609.29837 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 PUBG Ally,一个面向 PUBG: BATTLEGROUNDS 的具身代理,能推理、自主行动并作为语音队友与玩家并肩作战,将代理式工具使用与实时游戏控制相结合。PUBG Ally is introduced, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate that combines agentic tool use with real-time game control.

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
World Action Agent:通过世界动作预演利用 VLM 实现机器人操控
arXiv:2609.29964 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出 World Action Agent,一个多代理框架,通过它 VLM 可借助基础工具操控机器人,所有决策均在可视化动作工作空间内完成,在相同骨干下优于端到端 VLA、code-as-policy 代理以及一个可视化框架基线。World Action Agent is presented, a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone.

Rufus-Air: An Open LLM Post-Training Recipe
Rufus-Air:一种开放的大语言模型后训练方案
arXiv:2609.29421 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

主要发现是:多样且高质量的 SFT 奠定坚实的能力下限,难度过滤将 RL 提示维持在有效的学习区间内,而奖励可靠性为阶段排序提供了实用原则。The main findings are that diverse, high-quality SFT establishes a strong capability floor and difficulty filtering keeps RL prompts within a productive learning range, and reward reliability provides a practical principle for ordering stages.

ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation
ViRDM:驯服表征分布匹配以实现少步长因果视频生成
arXiv:2609.28923 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ViRDM,一种无需教师与评论家网络的视频后训练方案,将三网络蒸馏转化为仅生成器的后训练,在降低 GPU 显存与训练时间的同时提升视频质量。ViRDM, a teacher- and critic-free video post-training recipe that turns three-network distillation into generator-only post-training, reducing GPU memory use and training time while improving video quality, is introduced.

OmniEcho: Spatial Audio Understanding for Embodied Agents
OmniEcho:面向具身智能体的空间音频理解
arXiv:2609.23407 多模态 评测集 OA · 绿色 被引 1 · S2

本文提出 OmniEchoBench,一个面向空间音视频感知与音-视-语言导航的统一基准,以及一个空间感知的全模态模型,该模型在预训练语义音频通路之外引入 FOA 空间编码器。OmniEchoBench, a unified benchmark for spatial audio-visual perception and audio-vision-language navigation, and a spatially aware omni-modal model, which introduces an FOA spatial encoder alongside a pretrained semantic audio pathway.

Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone
Neural Spectral Capacity:仅依据网络规格度量与设计架构
arXiv:2609.23087 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Neural Spectral Capacity (NSC),一种以每个权重矩阵的奇异值谱为依据的闭式标量,在七类 Transformer 与 CNN 系列上的排序效果优于 #Params、#FLOPs 及代表性免训练代理指标。This work proposes Neural Spectral Capacity (NSC), a closed-form scalar grounded in the singular-value spectrum of each weight matrix, which outperforms #Params, #FLOPs, and representative training-free proxies in ranking across seven Transformer and CNN families.

AgentKernel: The Trust-Native Agentic Operating System
AgentKernel:原生可信的智能体操作系统
arXiv:2609.29647 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文论证代理需要一个操作系统级基底,为身份、输入中介、内存治理与执行控制提供强制且不可绕过的服务,并提出 AgentKernel,一个以安全为一流设计约束为前提、面向信任的代理操作系统。This work argues that agents need an operating-system substrate providing mandatory, non-bypassable services for identity, input mediation, memory governance, and execution control, and introduces AgentKernel, a trust-native agent operating system built around the premise that security must be a first-class design constraint.

Parts-of-Speech as Emergent Categories in SAE Latent Space
词性作为 SAE 潜空间中的涌现类别
arXiv:2609.29362 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果显示 SAE 以分布式、类别依赖的形式定位形态句法信息,而非通过原子化的语法特征。The results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.

1. Data Flow Control(DFC):AI Agent 数据安全策略的内核级执行框架
arXiv:2606.05679 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文将数据安全形式化为 provenance monomials 上的聚合谓词,并提出 Passant——一个无需物化 provenance 即可强制执行 DFC 策略的可移植查询重写层。This paper formalizes data safety as aggregate predicates over provenance monomials and presents Passant, a portable query rewriting layer that enforces DFC policies without materializing provenance.

Coding Agents for Generalized Task and Motion Planning Problems
用于广义任务与运动规划问题的编程智能体
arXiv:2609.30233 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文发现编码代理在通用 TAMP 上表现出惊人的有效性:在平均成功率上,三种代理配置均优于人工设计的规划器、一次性生成以及基于 LLM 的通用规划基线。This work finds that coding agents are surprisingly effective at generalized TAMP: all three agent configurations outperform hand-engineered planners, one-shot generation, and an LLM-based generalized planning baseline in mean success.

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
直接问 Jev:校准决策的强化学习作为 AI 对齐失败的零样本检测器
arXiv:2609.29429 安全与风险 评测集 OA · 绿色 被引 4 · S2

本文提出 RLCDAlignBench,在十类对齐失败上对 Jev 进行基准测试:谄媚、越狱、欺骗、提示注入、幻觉、隐私侵犯、社会偏见、奖励黑客、不确定性隐瞒与权力寻求。RLCDAlignBench is presented, which benchmarks Jev on ten alignment failures: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking.

DeltaWAM: Delta World Action Models for Bimanual Manipulation
DeltaWAM:面向双手操作的 Delta 世界动作模型
arXiv:2609.28811 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 DeltaWAM,通过 dense-anchor、sparse-delta 与 action 三条流联合预测视觉 delta 与动作,并设计了三种在表征与计算共享上有所不同的架构;同时开发了 Streaming Delta Memory (SDM),使用紧凑的观测 delta 更新缓存的 anchor 上下文,从而减少繁重的 video-expert 处理。This work proposes DeltaWAM, which jointly predicts visual deltas and actions using dense-anchor, sparse-delta, and action streams, with three architectures that differ in representation and computation sharing, and develops Streaming Delta Memory (SDM), which updates cached anchor context with compact observed deltas, reducing heavy video-expert processing.

Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs
你的 Transformer 可同时容纳两个思考:LLM 中线性叠加的证据
arXiv:2609.29845 工程化 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文给出证据,说明 superposition 是 Transformer 架构的内在属性,而非训练过程中涌现的结果,并证明通过轻量级 fine-tuning 可以在很大程度上恢复线性性。This work provides evidence that superposition is an intrinsic property of the Transformer architecture rather than an emergent consequence of training, and demonstrates that linearity can be substantially restored through lightweight fine-tuning.

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation
RGBD20K:面向 RGB-D 语义分割的大规模基准
arXiv:2609.29028 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种新颖的 score-purified fusion (SPF) 方法,在所有评测 benchmark 上均达到 SOTA 性能,验证了该方法在利用高质量多模态信息进行 RGB-D 语义分割任务中的有效性。A novel score-purified fusion (SPF) method is proposed, which achieves state-of-the-art performance across all evaluated benchmarks, demonstrating the effectiveness of the approach in leveraging high-quality multimodal information for RGB-D semantic segmentation.

AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation
AV-GRPO:用于联合音视频生成的模态锚定解耦扩散强化学习
arXiv:2609.29816 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 AV-GRPO,一种以模态为锚点的在线 diffusion RL 框架,以及 5DAV,一种解耦的、难度可调节的训练数据集;在 LoRA 与全量 fine-tuning 条件下,其在生成质量、语义对齐和跨模态同步性上均优于 LTX-2.3。This work proposes AV-GRPO, a modality-anchored online diffusion RL framework, and 5DAV, a decoupled, difficulty-controllable training dataset that outperforms LTX-2.3 in generation quality, semantic alignment and cross-modal synchronization under LoRA and full fine-tuning.

Learning to Discover Interesting Mathematics
学习发现有趣的数学
arXiv:2609.28603 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

近年来,大语言模型(LLMs)解决高级数学问题的能力日益增强,包括许多悬而未决数十年的难题。这为以空前规模扩展数学知识打开了大门。然而,尽管 LLM 能够猜想并证明越来越多的定理,这些新数学知识是否有趣或有用仍属未知。我们将定理的内在有趣度定义为其证明长度与陈述长度之比。证明该指标与下载量的外在度量高度相关。Recently, Large Language Models (LLMs) have been increasingly able to solve advanced mathematical problems, including many that have been open for decades. This opens the door to expansion of mathematical knowledge at unprecedented scale. Yet, while LLMs may be able to conjecture and prove more and more theorems, it remains open whether this new mathematical knowledge is interesting or useful. We define intrinsic interestingness of a theorem as the ratio between the length of its proof and the length of its statement. We show that this correlates strongly with an extrinsic measure of the downs

Learning multiple visual domains with residual adapters
使用残差适配器学习多个视觉域
arXiv:1705.08045 多模态 方法 OA · 绿色 被引 1105 · S2

本文提出一种可调深度网络架构,借助适配器残差模块可在线切换至不同视觉域;同时引入 Visual Decathlon Challenge 基准,用于评估表征同时捕获十个差异显著视觉域的能力,并衡量其跨域均匀识别的能力。This paper develops a tunable deep network architecture that, by means of adapter residual modules, can be steered on the fly to diverse visual domains and introduces the Visual Decathlon Challenge, a benchmark that evaluates the ability of representations to capture simultaneously ten very differentVisual domains and measures their ability to recognize well uniformly.

Block Sparse Attention with Log-Linear Complexity
对数线性复杂度的块稀疏注意力
arXiv:2609.31093 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 PISA,一种采用金字塔式 Top-K 选择策略的 block-sparse attention 机制,并为训练与推理开发了硬件感知的 Triton kernel,将层级路由与 LogSumExp 评分融合,无需显式构造 query-key score 矩阵。PISA is proposed, a block-sparse attention mechanism that employs a pyramid Top-$K selection strategy, and develops hardware-aware Triton kernels for both training and inference, fusing hierarchical routing and LogSumExp scoring without materializing the query-key score matrix.

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
AgentWorld:多 Agent LLM 长期协作基准
arXiv:2609.31590 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

为在传统二元任务成功之外量化协作有效性,本文提出 Causal Collaboration Effectiveness (CCE),一种基于图的指标,用于追踪 agent 动作之间的因果依赖,并度量团队投入中实际促成最终结果的比例。To quantify collaboration effectiveness in addition to conventional binary task success, Causal Collaboration Effectiveness (CCE) is proposed, a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome.

1. AlphaEval: Evaluating Agents in Production
AlphaEval: 在生产环境中评估 Agent
arXiv:2604.12162 评测基准 评测集 OA · 绿色 被引 1 · S2

本工作提出 AlphaEval,一个基于真实生产环境的基准,包含来自七家在其核心业务中部署 AI Agent 的公司的 94 个任务,覆盖六个 O*NET (Occupational Information Network) 领域;并贡献了一套从需求到基准的构建框架,将从需求到评估的完整流程标准化。This work presents AlphaEval, a production-grounded benchmark of 94 tasks sourced from seven companies deploying AI agents in their core business, spanning six O*NET (Occupational Information Network) domains, and contributes a requirement-to-benchmark construction framework that standardizes the entire pipeline from requirement to evaluation.

ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker
ZooWork-ShopRanker:开放且偏好对齐的电商重排模型
arXiv:2609.31002 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ZooWork-ShopRanker,一族对齐到裁判标注购物偏好的电商 reranker,以及 ShopRank-Bench,一个包含约 10,000 条私有流量偏好对的低污染 benchmark,按承诺该标注的裁判家族数量分档呈现,覆盖多种文本格式。ZooWork-ShopRanker, a family of e-commerce rerankers aligned to judge-labeled shopping preference, and ShopRank-Bench, a contamination-limited benchmark of ~10,000 private-traffic preference pairs in both text formats, tiered by how many judge families committed to each label.

Jev in the Wild: A Data-Driven Analysis of the Jev Model's Functionality, Applications and Ecosystem
野外 Jev:Jev 模型功能、应用与生态的数据驱动分析
arXiv:2609.30216 Agent 智能体 应用落地 OA · 绿色 被引 5 · S2

本文对从 GitHub 收集的 Jev 项目进行了大规模、数据驱动的分析,发现其公共生态在早期增长迅速,新项目不断涌现并被集成到既有仓库中;研究提示 Jev 充当一种可复用的决策组件,其功能随所嵌入的工作流而变化。A large-scale, data-driven analysis of Jev projects collected from GitHub finds rapid early growth in Jev's public ecosystem, with both new projects and integration into existing repositories, and suggests that Jev serves as a reusable decision component whose functionality varies with the surrounding workflow.

TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
TrackEverything:通过去重 3D 场景表示实现长时稠密跟踪
arXiv:2609.30222 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

TrackEverything 是首个在 40 GB GPU 显存内即可在超过 1000 帧的视频中跟踪所有可见点的 3D tracker;本文还提出 3D WAFT,用场景点云内的高效特征采样取代了显存开销巨大的 4D correlation volumes。TrackEverything is the first 3D tracker capable of tracking all visible points across videos exceeding 1000 frames within 40 GB of GPU memory, and 3D WAFT is proposed, replacing memory-prohibitive 4D correlation volumes with efficient feature sampling in the scene cloud.