研究库 论文知识库
Papers · organized/paper_cards

论文

1753 张论文卡片

开放获取 全部 绿色 · 1640
Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
广义 Agent 迭代:迭代式策略改进与递归自改进的统一形式化框架
arXiv:2609.13406 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Generalized Agent Iteration (GAI),一个形式化框架,将迭代策略改进和 RSI 描述为同一学习范式的两种情形,基于经典理论,使现有系统可比较,并为分析和设计新系统提供原则性基础。This paper proposes Generalized Agent Iteration (GAI), a formal framework that describes iterative policy improvement and RSI as two cases of a single learning paradigm that rests on the classical account, makes existing systems comparable, and provides a principled basis for analyzing and designing new ones.

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
ActionPiece: 重新思考自回归视觉-语言-动作模型中的动作 token 化
arXiv:2609.18487 多模态 观点 OA · 绿色 被引 1 · S2

本文提出物理秩一致性 (PRC) 来衡量 tokenization 在重建后保留局部物理距离排序的程度,并提出 ActionPiece,通过对表示学习和量化的联合监督来保留物理动作关系。This work introduces physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction, and presents ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization.

Agora: Git as Shared Memory for Collective AutoResearch
Agora: 以 Git 作为集体 AutoResearch 的共享记忆
arXiv:2609.18094 Agent 智能体 方法 OA · 绿色 被引 2 · S2

报告了一次持续近 12 天的运行:13 个 LLM worker 在无任务分配、无中央规划器的情况下,使用 Agora 解决了一个权重迁移问题,并记录了 agent 如何复用与验证共享工作。A run of nearly 12 days is reported in which 13 language-model workers, with no assigned tasks or central planner, used Agora to solve a weight-transfer problem and documents how agents reused and verified shared work.

Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
重新思考 PPO 中的 Critic 学习:理解与缓解 Value Flattening
arXiv:2609.18708 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

本文识别出 Value Flattening 是标准 PPO 中 critic 学习的一种重要但被忽视的失效模式,并提出一种简单稀疏监督策略可以缓解该问题;引入 SParse Proximal Policy Optimization,在每个响应中仅对少数间隔良好的状态施加 value loss,以同时缓解两种效应。Value Flattening is identified as an important yet overlooked failure mode of critic learning in standard PPO and a simple sparse supervision strategy can mitigate it; SParse Proximal Policy Optimization is introduced, which applies the value loss to only a few well-separated states in each response to mitigate both effects.

A Zeroth-Order Paradigm for LLM Preference Alignment
一种面向 LLM 偏好对齐的零阶范式
arXiv:2609.19144 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出并分析 Comparison-based Preference Optimization (ComPO),一种基于比较预言机 (oracle) 的零阶对齐方法,并在光滑性、梯度稀疏性以及 oracle 与潜在目标相容的条件下,为其基础离线方案建立了收敛性保证。This paper proposes and analyzes Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method based on comparison oracles, and establishes a convergence guarantee for its basic offline scheme under smoothness, gradient sparsity, and compatibility between the oracle and a latent objective.

Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents
置信来自经验:从推理到 Agent 的经验性置信估计
arXiv:2609.17708 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 XConf (eXperiential Confidence):与模型累积经验一起估计置信度,并将经验式置信度估计视为未来通用置信度估计的新范式。This work proposes XConf (eXperiential Confidence): estimating confidence together with the model's accumulated experience, and sees experiential confidence estimation as a new paradigm for future general-purpose confidence estimation.

Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control
Zing-0.5: 迈向具备实时联合动作与文本控制的可玩世界
arXiv:2609.17909 多模态 方法 OA · 绿色 被引 2 · S2

我们提出 Zing-0.5,一个 5B 自回归世界模型,专注于可玩性:用户可以探索生成的世界、影响事件演进,并通过键盘与在线文本的联合控制对反馈做出响应。我们的方法汇聚了三项技术贡献:(1) 统一的动作与文本条件建模,将感知幅度的键盘输入与时序对齐的文本指令、以及联合标注的视频结合,在同一序列中学习导航与事件控制;(2) 面向增量生成的事件尺度监督,使用分段级教师...We introduce Zing-0.5, a 5B autoregressive world model designed for playability: users can explore generated worlds, influence unfolding events, and respond to the resulting feedback through joint keyboard and online text control. Our approach brings together three technical contributions: (1) Unified action and text conditioning, combining magnitude-aware keyboard inputs with temporally aligned text instructions and jointly annotated videos to learn navigation and event control within the same sequence; (2) Event-scale supervision for incremental generation, using a segment-level teacher trai

HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses
HypoEvolve:遗传算法赋能多智能体 LLM 发现科学假设
arXiv:2609.15938 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出一种代际遗传算法,用以协调专门的 LLM Agent,整合机制性论证、重新审视假设并评估证据与可检验性,推动了自主科学的愿景——AI 研究团队实现超越单个模型的发现能力。This work proposes a generational genetic algorithm to coordinate specialized large language model agents that integrate mechanistic arguments, reconsider assumptions, and assess evidence and testability that advances a vision of autonomous science in which AI research teams achieve a capacity for discovery beyond that of individual models.

SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization
SpectralShift:通过谱重参数化实现 Gated DeltaNet 高效的上下文窗口扩展
arXiv:2609.14320 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 SpectralShift,一种用于 GDN 长上下文持续预训练的谱重参数化方法,通过重参数化 alpha 投影的初始化以重塑衰减谱,增强慢传播能力,并进一步引入 alpha 投影的学习率缩放以促进长上下文训练。SpectralShift is proposed, a spectral reparameterization approach for long-context continual pretraining of GDNs that reparameterizes the alpha projections initialization to reshape the decay spectrum by enhancing slow propagation capacity, and further introduces a learning-rate scaling for alpha projections to facilitate long-context training.

In-Context Robot Learning with VLM Agents
基于 VLM Agent 的机器人上下文学习
arXiv:2609.19138 Agent 智能体 应用落地 OA · 绿色 被引 5 · S2

本文提出 GPT-Policy,一个用于上下文机器人学习的通用 Agent 框架,集成了一个保留任务相关视觉过渡的 context compiler、一个提出机器人-工具动作的 VLM,以及一个验证并执行每个动作并报告结果的 constrained controller。GPT-Policy is introduced, a general-agent framework for in-context robot learning that integrates a context compiler that preserves task-relevant visual transitions, a VLM that proposes robot-tool actions, and a constrained controller that verifies and executes each action and reports its outcome.

1️⃣1️⃣ arXiv · RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic RAG Systems(⭐⭐⭐ 参考)
arXiv · RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic RAG Systems(⭐⭐⭐ 参考)
arXiv:2510.13910 评测基准 评测集 被引 5 · S2

本文提出 RAGCap-Bench,一个面向能力的 benchmark,用于对 agentic RAG workflow 中的中间任务进行细粒度评测,并构建了典型 LLM 错误的分类体系以设计针对性评测问题。This work proposes RAGCap-Bench, a capability-oriented benchmark for fine-grained evaluation of intermediate tasks in agentic RAG workflows, and constructs a taxonomy of typical LLM errors to design targeted evaluation questions.

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
PANORAMA:基于掩码提议选择的全景接地描述生成
arXiv:2609.19143 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 PanoCaps,一个由全景分割数据集构建的人工标注 benchmark,并提出 PANORAMA,一个将预训练 segmenter 条件化于上下文化短语表示以获得候选 mask、并学习选择与每个短语对应的 mask 的 VLM。This work introduces PanoCaps, a human-annotated benchmark constructed from panoptic segmentation datasets, and introduces PANORAMA, a VLM that conditions a pretrained segmenter on contextualized phrase representations to obtain candidate masks and learns to select those corresponding to each phrase.

The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction
内存墙的另一半:通过训练路由预测从 SSD 服务 35B MoE
arXiv:2609.18063 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Edge0,一个流式 MoE 推理引擎,通过 prerouter 缩小差距:每层 head 提前一个 token 预测下一层的 routing,并将该预测直接作为 routing 使用,使分阶段 expert 集合等于路由集合,无任何丢弃。Edge0 is presented, a streaming MoE inference engine that closes the gap with a prerouter: a per-layer head predicts the next layer's routing one token ahead, and the prediction is consumed as the routing itself, so the staged expert set equals the routed set and nothing is dropped.

CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents
CERA-MoA:与持续学习 LLM Agent 协同进化的路由机制
arXiv:2609.18779 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文设计了一种预测式熟悉度估计器,利用中间层隐藏状态评估 Agent 间的语义能力,避免完整 rollout 的开销,并在任务性能和效率之间实现权衡。A predictive familiarity estimator that leverages mid-layer hidden states to evaluate semantic competence among agents, avoiding the overhead of full rollouts and achieving a trade-off between task performance and efficiency is designed.

Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches
Fathom:面向卸载 KV 缓存稀疏解码的每查询读取深度
arXiv:2609.17652 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Fathom,一种 key scan,其中每个查询决定读取每个 key 通道的比特数;在真实的 coding agent 会话中,Fathom 以 92 bit 达到最准确的 136 bit scan 的步骤一致性。Fathom is presented, a key scan in which each query decides how many bits of each key channel to read, and on real coding-agent sessions Fathom reaches the step agreement of the most accurate 136-bit scan at 92 bits.

Assessing nnU-Net Generalization across Brain Tumor Populations in BraTS-GoAT 2026
在 BraTS-GoAT 2026 中评估 nnU-Net 跨脑肿瘤人群的泛化能力

BraTS-GoAT 在异质人群上使用传统 3D nnU-Net 对 1,351 个标注案例进行肿瘤分割评估,采用五折交叉验证,每折 1,000 epochs,并应用 test-time mirroring。BraTS-GoAT evaluates tumor segmentation across heterogeneous populations using a conventional 3D nnU-Net on 1,351 labeled cases using five-fold cross-validation and 1,000 epochs per fold and applied test-time mirroring.

Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training
长上下文 Mixture-of-Experts 训练中每个内存峰值的平整化
arXiv:2609.14306 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

常用的并行方案留下了四种未被约束的并行维度,各自增长方式不同:随路由矩阵增长的专家调度、随 token 数乘词表规模增长的词表投影、随深度乘序列长度增长的梯度检查点边界,以及随参数量增长的优化器状态。Four parallelism plans are left unbounded by the parallelism plans in common use, and each grows differently: expert dispatch with the routing matrix, the vocabulary projection with tokens times vocabulary, gradient checkpoint boundaries with depth times sequence length, and optimizer state with parameter count.

Fingers as Legs: Learning Self-Supported Locomotion and Manipulation with an Anthropomorphic Hand
手指作为腿:使用仿人手学习自支撑运动与操作
arXiv:2609.17172 工程化 观点 OA · 绿色 被引 0 · S2 + OpenAlex

该工作展示了一个紧凑的移动机械手,复用其手指同时完成运动与交互,无需独立的运动机构。This work demonstrates a compact mobile manipulator that reuses its fingers for locomotion and interaction, without a separate locomotion mechanism.

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
DeepSeek-V4.1-Flash:突破 KV Cache 压缩的极限
arXiv:2609.19969 LLM 基础设施 应用落地 OA · 绿色 被引 43 · S2

推出 DeepSeek-V4.1-Flash 模型,这是一个具有 552B 骨干参数、支持最长一百万 token 上下文的多模态 Mixture-of-Experts 模型,显著提升了 agent 工作负载的成本效率,并突破了 KV cache 压缩的极限。The DeepSeek-V4.1-Flash model, a multimodal Mixture-of-Experts model with 552B backbone parameters and support for contexts of up to one million tokens, is introduced, substantially improving cost efficiency for agentic workloads and pushing the limits of KV cache compression.

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
RetireOPD:用于智能体强化学习的自退役 On-Policy 蒸馏
arXiv:2609.20784 Agent 智能体 方法 OA · 绿色 被引 1 · S2

该工作提出 RetireOPD(Self-Retiring On-Policy Distillation),先用环境奖励优化一个解耦的、技能条件化的教师模型,再联合 RL 与 OPD 训练一个无技能依赖的学生模型。This work proposes RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD.

WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing
WeVisDoc:从覆盖到能力的鲁棒端到端文档解析
arXiv:2609.20423 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 WeVisDoc,一个面向鲁棒端到端文档解析的两阶段数据驱动框架,通过异构数据与保持结构的退化合成来扩展语义、结构与外观覆盖度,并指导有针对性的数据构建与目标 token 预算的重分配。WeVisDoc is presented, a two-stage data-centric framework for robust end-to-end document parsing that broadens semantic, structural, and appearance coverage through heterogeneous data and structure-preserving degradation synthesis and guides targeted data construction and reallocation of the target-token budget.

1️⃣ arXiv · Memory for Autonomous LLM Agents: Mechanisms, Evaluation, and Open Problems(⭐⭐⭐⭐⭐ 必读综述)
自主 LLM Agent 的记忆:机制、评估与开放问题
arXiv:2603.07670 Agent 智能体 综述 Open MIND OA · 绿色 被引 66 · S2

本文系统梳理了基于 LLM 的现代智能体中记忆的设计、实现与评估方法,覆盖 2022 年至 2026 年初的相关工作,并将 Agent 记忆形式化为一个涵盖时间范围、表示基底与控制策略的三维分类体系。This survey offers a structured account of how memory is designed, implemented, and evaluated in modern LLM-based agents, covering work from 2022 through early 2026, and formalizes agent memory as a three-dimensional taxonomy spanning temporal scope, representational substrate, and control policy.

What Does Privileged Information Add to On-Policy Self-Distillation?
特权信息为 On-Policy 自蒸馏带来了什么?
arXiv:2609.20612 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

构建了 AMPLE-Math,一个包含 5,319 道数学问题、每题对应六种共享同一答案的推理视角的可复用套件,并通过对比表明 OPSD 能借助直答推理与带思维链推理所共享的参数,提升对已有推理能力的调用效率。AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, is constructed and AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, is compared, suggesting that OPSD can improve access to existing reasoning capabilities through parameters shared by direct-response and thinking-enabled inference.

UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation
UFO:面向多模态图像生成全条件对齐的链式评估
arXiv:2609.12397 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 UFO,这是首个面向全条件对齐同时评估的统一框架;并发布 UFO-Bench,一个用于整体评估现有定制化模型在文本与视觉条件多样化交互下表现的专用基准。UFO is proposed, the first unified framework for omni-condition alignment simultaneous evaluation, and UFO-Bench is presented, a dedicated benchmark designed to holistically evaluate the performance of existing customization models under the diverse mutual interactions of textual and visual conditions.

When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
When2Think:面向高效混合推理模型的难度感知长度控制学习
arXiv:2609.19671 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 When2Think,一种基于 RLVR 的后训练框架,用于实例自适应计算分配,既无需学习奖励模型,也无需学习 critic;离线参考缓存机制避免了策略更新阶段对参考模型的在线查询。This work proposes When2Think, an RLVR-based post-training framework for instance-adaptive computation allocation that requires neither a learned reward model nor a learned critic, and offline reference caching avoids online reference-model queries during policy updates.

PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
PACT:压力之下企业 AI 助手值得被信任吗?
arXiv:2609.18605 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出 PACT(Pressure-Applied Compliance Testing),一个用于评估 AI agent 在压力下遵守规则的基准,覆盖十二个受监管的企业领域与四十八个场景,每个场景均为真实的多轮对话。This work introduces PACT (Pressure-Applied Compliance Testing), a benchmark for rule-following under pressure in AI agents assisting employees in daily tasks across twelve regulated enterprise domains and forty-eight scenarios, each set in a realistic multi-turn conversation.

Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents
反思、修订、复用:面向 GUI 智能体的免训练技能进化
arXiv:2609.17653 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出 EvoSkill-GUI,一个免训练框架,其中每个 skill 都是一个结构化的多文件包,包含检索元数据、可执行计划、备份定位、故障恢复规则、可访问性工具以及失败案例。This work proposes EvoSkill-GUI, a training-free framework in which each skill is a structured multi-file package containing retrieval metadata, executable plans, backup localization, failure-recovery rules, accessibility utilities, and failure cases.

When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
当 EOS token 不一致时:理解 On-Policy 蒸馏中的长度膨胀
arXiv:2609.20511 LLM 基础设施 方法 OA · 绿色 被引 2 · S2

我们研究 on-policy 蒸馏 (OPD) 中的长度膨胀现象,即学生回答会变得过长,甚至耗尽生成预算。我们发现基础学生模型与训练后教师模型之间的终止 token 不匹配是该行为的重要来源。在 Qwen3、Llama 和 Gemma 上,两个模型可能将停止概率分配到不同的 EOS token 上,即便它们声明的停止集合相同。这种不匹配会抑制学生偏好的终止动作,同时无法可靠地传递教师偏好的替代动作。我们证明对齐 t...We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify termination-token mismatch between base students and post-trained teachers as an important source of this behavior. Across Qwen3, Llama, and Gemma, the two models can place their stopping probability on different EOS tokens, even when their declared stopping sets are identical. This mismatch can suppress the student's preferred termination action without reliably transferring the teacher-preferred alternative. We show that aligning t

FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations
FAMOS:基于稀疏观测的前馈 3D 关节建模
arXiv:2609.20817 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 FAMOS,一个前馈模型,可从稀疏、无序的部分点云集合预测可动部件分割与关节参数;并引入一个过程式数据生成器,在训练过程中合成自标注资产,以克服现有数据集规模和多样性的局限。FAMOS, a feed-forward model that predicts movable-part segmentation and joint parameters from a sparse, unordered set of partial point clouds, is presented and a procedural data generator that synthesizes self-annotated assets during training is introduced to overcome the limited scale and diversity of existing datasets.

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
不要遮蔽环境:观测监督会改变 RL 下智能体的探索行为
arXiv:2609.20715 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ActObs,对每条轨迹中已有的观测 token 也进行监督,并将这一差异归因于 SFT:动作与观测梯度迅速趋于正交,而仅训练动作会留下较大的残留观测梯度,并将环境预测能力拉低至基座模型之下。This work introduces ActObs, which also supervises the observation tokens already present in each trajectory, and traces this difference to SFT: action and observation gradients rapidly become orthogonal, while action-only training leaves a large residual observation gradient and degrades environment prediction below the base model.

VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
VākQA:泰卢固语口语事实型问答基准与评估研究
arXiv:2609.19879 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 V\={a}kQA,一个覆盖六个领域、包含 2,001 对事实型问答的泰卢固语 SQA 基准;并观察到:泰卢固语措辞保留了翻译中会丢失的文化特异性,语音输入引入的音近混淆会改变问题含义,级联 ASR-MT 误差会逐步叠加放大。V\={a}kQA, a Telugu SQA benchmark of 2,001 factoid question-answer pairs across six domains, is introduced and it is observed that Telugu phrasing retains cultural specificity that is lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively.

Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model
MiniMax-H3 能否对物理世界进行推理?一种全模态生成模型的评估
arXiv:2609.18323 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个围绕物理世界推理四个互补维度构建的综合评估框架,并构建了一系列多样化新任务,要求模型跨模态整合互补信息。This work introduces a comprehensive evaluation framework organized around four complementary dimensions of physical world reasoning, and constructs a diverse set of novel tasks that require models to integrate complementary information across modalities.

1️⃣ arXiv · Learning Rate Matters: Vanilla LoRA May Suffice(⭐⭐⭐⭐⭐ 必读)
学习率至关重要:Vanilla LoRA 可能已足够
arXiv:2602.04998 评测基准 方法 Open MIND OA · 绿色 被引 11 · S2

本文通过大规模超参数搜索,系统地重新评估了 Vanilla LoRA 以及九个代表性 LoRA 变体,发现不同 LoRA 方法偏好的学习率区间各异,并将最优学习率区间的差异归因于最大 Hessian 特征值的变化,与经典学习理论相吻合。This work systematically re-evaluate nine representative LoRA variants alongside vanilla LoRA through extensive hyperparameter searches, finding that different LoRA methods favor distinct learning rate ranges and attributes the differing optimal learning rate ranges to variations in the largest Hessian eigenvalue, aligning with classical learning theories.

Srijika: OpenType-Layout-Reusing Font Restyling for Nine Indic Scripts
Srijika:面向九种印度文字的 OpenType 布局复用字体再设计
arXiv:2609.05661 工程化 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Srijika,一个为九种 Brahmic 文字(Devanagari、Tamil、Bengali、Telugu、Kannada、Malayalam、Gujarati、Gurmukhi、Odia)生成可安装 OpenType 字体的系统,并附带一份负面结果目录,覆盖参考引导重风格化中失败的 conditioning、目标函数选择和数据凸包限制。Srijika is presented, a system for producing installable OpenType fonts for nine Brahmic scripts: Devanagari, Tamil, Bengali, Telugu, Kannada, Malayalam, Gujarati, Gurmukhi, and Odia, and a negative-results catalogue covering failed conditioning, objective choices, and data-hull limits of reference-guided restyling.

Self-Evolving Search Index
自我演化的搜索索引
arXiv:2609.19656 RAG 检索增强 方法 被引 0 · S2

本文提出 SELF-INDEX,一个无需人工干预即可让索引自我演进的框架;其 Optimizer 可自主诊断检索短板,选择性修改对应的索引键,并在每次更新前对每项修订进行验证。SELF-INDEX is proposed, a framework that enables an index to self-evolve without human intervention, and its Optimizer autonomously diagnoses retrieval shortfalls, selectively revises the responsible index keys, and validates each revision before updating the index.

Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
样本数量远远不够:候选生成策略决定 LLM 测试时扩展的能耗与性能
arXiv:2609.19499 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,仅靠候选数量不足以刻画多候选 test-time scaling 的系统成本;评测应同时报告候选数量与准确率,以及生成调度和 GPU 层级的系统指标。The results show that candidate count alone is not enough to describe the systems cost of multi-candidate test-time scaling and Evaluations should report not only candidate count and accuracy, but also generation schedule and GPU-level systems metrics.