研究库 论文知识库
Papers · organized/paper_cards

论文

77 张论文卡片 · 观点

开放获取 全部 绿色 · 1640
4. Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented QA
4. 迷失在末尾:多模态检索增强问答中的首因偏差
arXiv:2606.16494 多模态 观点 OA · 绿色 被引 1 · S2

研究发现,recall@k 并非已部署 KB-VQA 的正确评价指标,且弥合差距需要 reader 侧介入;本文首次对多模态 KB-VQA 中 reader 侧位置依赖性进行了受控探查,设计了一种 gold-position 协议——在问题提示中仅改变 gold passage 所在的槽位。The first controlled probe of reader-side position dependence in multimodal KB-VQA is designed, a gold-position protocol in which only the gold passage's prompt slot varies within question, indicating that recall@k is the wrong metric for deployed KB-VQA and that the remaining headroom sits on the reader side.

5️⃣ Multi-Segment Attention · 分块位置感知KV驱逐 — arXiv:2606.02964(⭐⭐⭐ 新鲜 arXiv)
arXiv:2606.02964 LLM 基础设施 观点 OA · 绿色 被引 2 · S2

AsymCache 是一个面向 LLM 推理的计算-延迟感知 KV cache 管理系统,将 cache 驻留决策与 GPU attention kernel 性能显式对齐,包含三个关键组件:用于高效处理非连续 KV 上下文的多段注意力(MSA)、联合优化命中率与位置感知重计算代价的 cache 淘汰策略,以及面向高硬件利用率的自适应分片调度器。AsymCache is proposed, a computation-latency-aware KV cache management system for LLM inference that explicitly aligns cache residency decisions with GPU attention kernel performance, including three key components: Multi-Segment Attention (MSA) for efficient non-contiguous KV context processing, a cache eviction policy that jointly optimizes hit rate and position-aware recomputation cost, and an adaptive chunking scheduler for high hardware utilization.

7️⃣ arXiv · Position Paper:LLM Serving 需要数学优化,而非仅靠启发式 ⭐⭐⭐⭐⭐ 学术前沿
arXiv:2605.01280 LLM 基础设施 观点 OA · 绿色 被引 1 · S2

这篇立场论文认为,LLM 推理 serving 已超越通用启发式方法,如今需要数学优化与算法基础,并呼吁社区将 LLM serving 的算法设计视为一个新的研究前沿。This position paper argues that LLM inference serving has outgrown generic heuristics and now demands mathematical optimization and algorithmic foundations, and calls on the community to recognize algorithmic design for LLM serving as a research frontier.

5. Context-Fractured Decomposition Attacks on Tool-Using LLM Agents
5. 上下文碎裂分解攻击针对使用工具的 LLM Agent
arXiv:2606.09084 Agent 智能体 观点 OA · 绿色 被引 3 · S2

揭示使用工具的 LLM Agent 的一种部署失效模式——来源缺口,以及一类跨上下文多步越狱攻击,可在早期交互中保留看似无害的中间产物,并在很久以后(可能在不同 Agent 实例或工作流阶段)诱发有害行为。A deployment failure mode for tool-using LLM agents, the provenance gap, and a family of cross-context multi-step jailbreaks that preserve benign-looking intermediate artifacts from an early interaction and elicit harmful behavior much later, potentially in a different agent instance or workflow stage.

SSGM框架(Stability and Safety-Governed Memory)
3. SSGM框架(Stability and Safety-Governed Memory)
arXiv:2603.11768 安全与风险 观点 OA · 绿色 被引 21 · S2

通过形式化分析与架构分解,展示 SSGM 如何缓解拓扑引发的知识泄漏(敏感上下文被固化到长期存储),以及有助于防止语义漂移(知识在迭代摘要中退化)。Through formal analysis and architectural decomposition, it is shown how SSGM can mitigate topology-induced knowledge leakage where sensitive contexts are solidified into long-term storage, and help prevent semantic drift where knowledge degrades through iterative summarization.

条目D3:UnWeaving GraphRAG — GraphRAG vs VectorRAG 理论分析(arXiv 2603.29875v3)
条目D3:UnWeaving GraphRAG — GraphRAG vs VectorRAG 理论分析(arXiv 2603.29875v3)
arXiv:2603.29875 评测基准 观点 OA · 绿色 被引 0 · S2 + OpenAlex

文章认为基于实体的分解能形成对原始信息更精炼的表示,并有助于降低索引与生成过程中的噪声;在端到端 QA 评测中,VectorRAG 表现优于标准 GraphRAG,且接近当前 SOTA 图方法的效果。It is argued that entity-based decomposition yields a more distilled representation of original information, and additionally serves to reduce noise in the indexing, and generation process, and on end to end QA evaluation VectorRAG performs better than standard GraphRAG and almost as good as current SOTA graph-based solutions.

条目 G: 公共部门 ML Pipeline 工程教训(含性能数据表)
arXiv:2511.01545 工程化 观点 OA · 绿色 被引 1 · S2

研究表明,机器学习在公共部门的成功,将更少依赖模型准确率的突破,而更多依赖机构能否构建出透明、可复现、可问责且受公民信任的数据基础设施。It is shown that the success of machine learning in the public sector will depend less on breakthroughs in model accuracy and more on the ability of institutions to engineer transparent, reproducible, and accountable data infrastructures that citizens can trust.

COMA: A Compositional Misleading Attack Class on Security-RAG, and a Causal Counterfactual Defense
COMA:针对 Security-RAG 的组合式误导攻击类别及因果反事实防御。
arXiv:2608.17960 RAG 检索增强 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出一种审计方法,度量每个检索文档的留一因果影响,并对影响集中于低可信度文档的答案进行标记;并提出一种审计方法,度量每个检索文档的留一因果影响,并对影响集中于低可信度文档的答案进行标记。An audit is proposed that measures the leave-one-out causal influence of each retrieved document and flags answers whose influence concentrates on low-trust documents, and proposes an audit that measures the leave-one-out causal influence of each retrieved document and flags answers whose influence concentrates on low-trust documents.

MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement
MathForm:基于知识检索与验证引导优化的数学自动形式化扩展
arXiv:2608.14221 RAG 检索增强 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MathForm,一个通过 Mathlib 知识检索与验证引导的迭代优化来构建已验证训练数据的自动形式化框架,性能优于多个专用的 32B 自动形式化模型。MathForm is introduced, an autoformalization framework for constructing verified training data through Mathlib knowledge retrieval and verification-guided iterative refinement, and outperforming multiple specialized 32B autoformalizers.

A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans
一种面向 CT 扫描中空间关系验证的、可靠且可审计的模块化 Agent
arXiv:2608.21140 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个模块化的医学影像 Agent,用于轴位 CT 切片中的二元空间关系验证,采用显式模块化空间验证阶段,并表明显式模块化空间验证可作为未来面向报告的医学影像 Agent 的有前景的构建模块。This work presents a modular medical imaging agent for binary spatial relation verification in axial CT slices using explicit modular spatial verification stages, and suggests that explicit modular spatial verification can serve as a promising building block for future report-oriented medical imaging agents.

Procedura: Agentic 3D Modeling with Procedural Control
Procedura:基于过程化控制的 Agentic 三维建模
arXiv:2608.26238 Agent 智能体 观点 OA · 绿色 被引 1 · S2

本文探索"以代码表示 3D 形状"的范式,利用并放大 LLM 的编码能力进行 3D 建模,并提出 Procedura 框架——通过编写一个由命名零件构成、并通过类型化、可机器校验的连接关系装配而成的参数化程序来建模对象。The paradigm of 3D shape as code is explored, leveraging and scaling the coding ability of an LLM for 3D modeling, and Procedura, a novel 3D modeling agent framework that writes an object as a procedural assembly, a parametric program whose named parts are joined by typed, machine-checkable mates.

What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals
评估授权了什么?Inspect Evals 中声明相对推理的提交绑定普查
arXiv:2608.19269 LLM 基础设施 观点 OA · 绿色 被引 0 · S2 + OpenAlex

冻结一个大规模评测集合,并尝试对其历史断言进行 replay,以使这一原本隐式的推理步骤变得显式且可执行。A large evaluation collection is frozen and an attempt to replay its historical claims is made to make this otherwise implicit inference step explicit and executable.

Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding
在视频中定位任意目标:重新思考高效生成式时空视频定位
arXiv:2608.28192 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Decoupled Block Attention,在保留共享 video-query 上下文访问的同时消除跨 box 依赖,并结合用于时间边界与空间几何的 localization-aware policy optimization;并行 tube generation 被证明是视频中 autoregressive 定位的一种高效替代方案。Decoupled Block Attention is introduced, which preserves access to shared video-query context while eliminating cross-box dependencies, together with localization-aware policy optimization for temporal boundaries and spatial geometry, and parallel tube generation is shown to be an efficient and effective alternative to autoregressive localization in videos.

Configurable Semantic Chunking for Biomedical Information Extraction in Retrieval-Augmented Generation
RAG 中面向生物医学信息抽取的可配置语义分块
arXiv:2608.31139 RAG 检索增强 观点 OA · 绿色 被引 0 · S2 + OpenAlex

跨数据集分析表明,语义切分在具有显式关系线索的抽取数据集(如 GM-CIHT 和 DDI)上表现更优,而固定切分在密集生化抽取和二分类场景(如 ChemProt 和 ADE)下仍具竞争力甚至更强。Cross-dataset analysis shows that semantic chunking improves extraction datasets with explicit relation cues, such as GM-CIHT and DDI, while fixed chunking remains competitive or stronger for dense biochemical extraction and binary classification settings such as ChemProt and ADE.

Verification-Aware Training for Speculative Decoding
面向 Speculative Decoding 的验证感知训练
arXiv:2608.30135 工程化 观点 OA · 绿色 被引 2 · S2

提出 Verification-Aware Training,一种插件式框架,在每一步训练中模拟验证并将产生的 accept 与 reject 模式转化为监督信号,在数学、代码和聊天基准上提升了平均接受长度和实际加速比。Verification-Aware Training is introduced, a plug-in framework that simulates verification at every training step and turns the resulting accept and reject patterns into supervision and improves average acceptance length and wall-clock speedup across math, code, and chat benchmarks.

Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching
编织视觉叙事:超越原子视觉匹配的 Agentic Image Bundle Composition
arXiv:2608.28695 Agent 智能体 观点 OA · 绿色 被引 1 · S2

大量实验表明,SOTA embedding 模型和静态 decompose-and-rerank 范式存在关系盲区,而 BundleWeaver 取得了显著性能提升,凸显了从原子打分转向动态关系组合的必要性。Extensive experiments demonstrate that while state-of-the-art embedding models and static decompose-and-rerank paradigms suffer from relational blindness, BundleWeaver achieves substantial performance gains, highlighting the necessity of shifting from atomic scoring to dynamic relational composition.

Recursive Criticality of AI Self-Improvement
AI 自我改进的递归临界性
arXiv:2609.00137 安全与风险 观点 OA · 绿色 被引 0 · S2 + OpenAlex

该框架识别了 AI R&D 系统的若干可测量属性,可用于区分递归放大与其他来源驱动的快速进展,包括递归反馈的强度、改进向后续系统传播的有效性、周期时长,以及进一步取得进展的难度递增。The framework identifies measurable properties of AI R\&D systems that can help distinguish recursive amplification from rapid progress driven by other sources, including the strength of recursive feedback, how effectively improvements propagate into successor systems, cycle duration, and the increasing difficulty of further progress.

A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss
一瞥即可:基于 SimLoss 的单遍细粒度图像描述生成
arXiv:2609.00591 RAG 检索增强 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 SimLoss,一种面向单轮细粒度图像描述的无参考 embedding 空间目标;结果表明 embedding 空间监督能在单轮描述器延迟下恢复多阶段验证的质量。SimLoss is proposed, a reference-free embedding-space objective for single-pass fine-grained image captioning, and results show that embedding-space supervision can recover the quality of multi-stage verification at the latency of a single-pass captioner.

4️⃣ arXiv · Securing the Agent: Vendor-Neutral, Multitenant Enterprise Retrieval and Tool Use(⭐⭐⭐⭐ 高优先级)
4️⃣ arXiv · 守护 Agent:厂商中立的多租户企业级检索与工具调用(⭐⭐⭐⭐ 高优先级)
arXiv:2605.05287 Agent 智能体 观点 OA · 绿色 被引 1 · S2

本文提出一种分层隔离架构,结合策略感知的 ingestion、retrieval-time gating 与共享推理,并通过服务端 agentic 编排加以执行,在为多租户隔离提供天然强制点的同时,允许客户端框架保留对 agent 组合与延迟敏感操作的控制权。A layered isolation architecture combining policy-aware ingestion, retrieval-time gating, and shared inference, enforced through server-side agentic orchestration is introduced, creating natural enforcement points for multitenant isolation while allowing client-side frameworks to retain control over agent composition and latency-sensitive operations.

Substrate-Aware AI Agents: Execution Context as a First-Class Input
Substrate-Aware AI Agent:将执行上下文作为一等输入。
arXiv:2609.05232 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

最小执行契约可在生成程序中引发主动的结构适应,将计算从无约束分配中移开,在执行前显著改善观察到的资源时间分布,建立了 substrate-aware Agent 规划的受控概念验证。A minimal execution contract induces proactive structural adaptation in generated programs, shifting computation away from unconstrained allocations and substantially improving observed resource-time profiles before execution, establishing a controlled proof of concept for substrate-aware agent planning.

Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems
双层协同反思:多 Agent LLM 系统的博弈论方法
arXiv:2609.02750 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出随机反思记忆提升(SRMA),仅在 grounding 后的评估风险严格下降时才接受候选记忆,为随机评估提供置信度门控,并为分段平稳环境提供重新锚定保证。Stochastic Reflective Memory Ascent (SRMA), which accepts a candidate memory only after a grounded evaluation risk strictly decreases, is introduced and provides confidence gating for stochastic evaluation and re-anchoring guarantees for piecewise-stationary environments.

Enoki: Efficient Multi-Level Hallucination Detection
Enoki:高效多层级幻觉检测
arXiv:2609.00581 安全与风险 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Enoki,一个面向多级幻觉检测的开放信息抽取框架,支持基于 LLM、基于编码器与基于规则的三类抽取模式,通过统一接口平衡准确率与推理成本。This work proposes Enoki, an Open Information Extraction framework for multi-level hallucination detection that supports LLM-based, encoder-based, and rule-based extraction regimes, balancing accuracy and inference cost through a common interface.

KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints
KVShareArena:跨上下文与模型 checkpoint 的 KV cache 复用
arXiv:2609.10266 LLM 基础设施 观点 OA · 绿色 被引 0 · S2 + OpenAlex

揭示了复用造成的质量损失以及有效修复方法均取决于 LLM(即便在两个 8B 模型之间也是如此),并提供添加新方法的通用接口与交互式 leaderboardBoth the quality loss from reuse and which repairs help depend on the LLM, even between two 8B models are introduced, as well as a common interface for adding new methods and an interactive leaderboard.

Building a Production Greek-English Speech Recognizer
构建生产级希腊语-英语语音识别器
arXiv:2609.13498 多模态 观点 被引 0 · S2

我们报告了一项历时数月的工程项目,用于构建 Sophea,一个生产级希腊语-英语双语自动语音识别系统。我们依据九个生产门控评估该系统,涵盖希腊语和英语词错误率、语言识别及非语音音频的幻觉问题。经过二十三次训练迭代和两种模型架构,没有任何训练数据组合能同时通过全部九个门控。满足希腊语噪声环境目标需要约1,500步密集领域暴露,而保持英语语言识别仅能容忍约2%We report a multi-month engineering program to build Sophea, a production bilingual Greek-English automatic speech recognition system. We evaluate the system against nine production gates covering Greek and English word error rate, language identification, and hallucinations on non-speech audio. Across twenty-three training iterations and two model architectures, no training-data composition passed all nine gates simultaneously. Meeting the Greek noisy-environment target required about 1,500 steps of dense domain exposure, while preserving English language identification tolerated only about 2

Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction
Grouped Value Attention:通过按需 Key 重建实现高效 KV 缓存
arXiv:2609.13285 LLM 基础设施 观点 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出了分组值注意力(Grouped Value Attention, GVA),通过存储分组值并在推理时利用可学习的线性映射重构内容键,该映射可被吸收进 query 中,从而在预期解码路径中无需显式物化内容键。This work introduces Grouped Value Attention (GVA), which stores grouped values and reconstructs content keys with a learned linear map at inference, which can be absorbed into the query, eliminating the need to materialize content keys in the intended decode path.

Thought without systematicity? Evaluating reasoning models on rule induction tasks
没有系统性的思维?在规则归纳任务上评估推理模型
arXiv:2609.13948 评测基准 观点 被引 0 · S2

人类认知的一个核心原则是系统性,即理解一个概念本质上与理解该概念的相近变体相关联。推理模型能否稳健地展现这种系统性?如果可以,我们应能预期模型在其结构等价变体任务上表现一致。本文扩展了认知科学中既定的规则归纳任务,以评估当前推理模型思维的系统性。每个任务族都具有组合结构,我们借此通过任务同构创建结构等价的任务变体...A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to understanding close variations of that concept. Do reasoning models robustly exhibit such systematicity? If so, we would expect consistent performance on structurally equivalent variants of the same task. Here, we extend established rule induction tasks from cognitive science to assess the systematicity of thought in current reasoning models. Each task family has compositional structure that we use to create structurally equivalent task variations through task isomorphisms su

PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control
PhysStream:基于结构化场景记忆与细粒度运动控制的流式物理约束视频生成
arXiv:2609.17521 多模态 观点 OA · 绿色 被引 1 · S2

PhysStream 是一个用于物理驱动图像到视频合成的自回归模型,通过引入结构化场景记忆并支持基于稀疏速度增量信号的细粒度运动控制(编码物理量),使模型能够学习底层动力学。PhysStream is an autoregressive model for physics-grounded image-to-video synthesis that incorporates structured scene memory and supports fine-grained motion control via sparse velocity-increment signals that encode physical quantities, letting the model learn the underlying dynamics.

Register Tokens for Bounded-State Reasoning in Diffusion Language Models
用于扩散语言模型有界状态推理的 Register Tokens
arXiv:2609.16372 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

Register被实现为专用的固定位置 token,其连续的隐藏状态被训练用于跨生成块承载推理进度,对有界代码生成尤其有效,因为正确程序通常跨越多个块。Registers are implemented as dedicated fixed-position tokens whose continuous hidden states are trained to carry reasoning progress across generation chunks, which are especially effective for bounded code generation, where correct programs usually span several chunks.

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
ActionPiece: 重新思考自回归视觉-语言-动作模型中的动作 token 化
arXiv:2609.18487 多模态 观点 OA · 绿色 被引 1 · S2

本文提出物理秩一致性 (PRC) 来衡量 tokenization 在重建后保留局部物理距离排序的程度,并提出 ActionPiece,通过对表示学习和量化的联合监督来保留物理动作关系。This work introduces physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction, and presents ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization.

Fingers as Legs: Learning Self-Supported Locomotion and Manipulation with an Anthropomorphic Hand
手指作为腿:使用仿人手学习自支撑运动与操作
arXiv:2609.17172 工程化 观点 OA · 绿色 被引 0 · S2 + OpenAlex

该工作展示了一个紧凑的移动机械手,复用其手指同时完成运动与交互,无需独立的运动机构。This work demonstrates a compact mobile manipulator that reuses its fingers for locomotion and interaction, without a separate locomotion mechanism.

16. [arxiv:2004.05074 — Paxos vs Raft: Have we reached consensus on distributed consensus?](https://arxiv.org/abs/2004.05074)
16. Paxos vs Raft:分布式共识是否已经达成共识?
arXiv:2004.05074 工程化 观点 被引 89 · S2

本文探讨 Paxos 与 Raft 何者更优地解决分布式共识问题,通过以 Raft 的术语与实用抽象描述一个简化的 Paxos 算法,精确揭示两者的差异。This paper considers the question of which algorithm, Paxos or Raft, is the better solution to distributed consensus to determine exactly how they differ by describing a simplified Paxos algorithm using Raft's terminology and pragmatic abstractions.

onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction
onPanda:通过 token 级校正高效标注 LLM 与 Agent 的 on-policy 对齐数据。
arXiv:2609.24983 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出 OnPanda,一种用于高效标注 LLM 对齐数据与 Agent 轨迹的交互式工具,以 token 级修正为核心交互方式;并发布使用 onPanda 标注的数据集 Panda-CVL,以及一个面向 token 级修正的基准。OnPanda is presented, an interactive tool for efficiently annotating LLM alignment data and agent trajectories that adopts token-level correction as its core interaction and releases Panda-CVL, a dataset annotated with onPanda, together with a benchmark for token-level correction.

1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation
1% 的 token 已足够:论 On-Policy 蒸馏中的梯度估计。
arXiv:2609.24432 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

基于信噪分解与候选集近似的 IER(信息效率比),可基于 IER 及其与现有效用分数的组合进行 token 选择,同时保留采样的反向 KL 训练目标。An information-efficiency ratio (IER) based on a signal-to-noise decomposition and a candidate-set approximation enables token selection based on IER and its combination with existing usefulness scores, while retaining the sampled reverse-KL training objective.

12. SoK: Agentic RAG(arXiv 2603.07379,ACL 2026)
12. SoK:Agentic RAG(arXiv 2603.07379,ACL 2026)
arXiv:2603.07379 RAG 检索增强 观点 Open MIND OA · 绿色 被引 8 · S2

本文将 Agentic 检索-生成循环形式化为有限时域部分可观测马尔可夫决策过程,显式建模其控制策略与状态转移,并构建了全面的分类体系与模块化架构分解,按规划机制、检索编排、记忆范式与工具调用行为对系统进行分类。This paper formalizes agentic retrieval-generation loops as finite-horizon partially observable Markov decision processes, explicitly modeling their control policies and state transitions, and develops a comprehensive taxonomy and modular architectural decomposition that categorizes systems by their planning mechanisms, retrieval orchestration, memory paradigms, and tool-invocation behaviors.

Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene
Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene
arXiv:2609.23796 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Mira-Scene,一种组合式 3D 场景重建框架,将稀疏姿态回归替换为密集有界对应恢复,并引入多模态扩散 Transformer 联合生成物体几何与 CCM,使用模态专精的专家流配合共享注意力与位置编码以促进几何-布局一致性。Mira-Scene is presented, a compositional 3D scene reconstruction framework that replaces sparse pose regression with dense, bounded correspondence recovery and introduces a multimodal diffusion transformer that jointly generates object geometry and CCMs, using modality-specific expert streams with shared attention and positional encoding to promote geometry-layout consistency.

The Functionalizer: Lossless Functional Decomposition for Subword Tokenization
The Functionalizer:面向子词分词的无损函数分解
arXiv:2609.15991 RAG 检索增强 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Functionalizer,一种无损预分词框架,将正字与结构变体分解为分词前的可组合操作码/操作数前缀流:以 Unicode 私有使用区编码的参数化变换操作符(操作码)为前缀,连接规范基础 token(操作数)。The Functionalizer is presented, a lossless pre-tokenizer framework that factors orthographic and structural variations into a compositional opcode/operand prefix stream before tokenization: a canonical base token (operand) prefixed by parametric transformation operators (opcodes) encoded in the Unicode Private Use Area.