研究库 论文知识库
Papers · organized/paper_cards

论文

1640 张论文卡片 · OA 绿色

开放获取 全部 绿色 · 1640
Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion
Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion
arXiv:2609.24220 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Document Retrieval-Aware Chunking (D-RAC),将 Web Retrieval-Aware Chunking (W-RAC) 框架扩展至任意文档格式,保留 W-RAC 的成本、确定性与可观测性优势,同时将每种可渲染格式解锁为一类输入。Document Retrieval-Aware Chunking (D-RAC), an extension of the Web Retrieval-Aware Chunking (W-RAC) framework to arbitrary document formats, is presented, preserving W-RAC's cost, determinism, and observability benefits while unlocking every renderable format as a first-class input.

Streaming Video Editing with Easy Adaptation
Streaming Video Editing with Easy Adaptation
arXiv:2609.24788 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 SVEET 框架,仅需在预训练的双向视频扩散模型上训练,即可支持高质量自回归流式视频编辑,并提出解耦训练方案,显式强制视频可控性与模型因果性优化方向之间的正交性。This paper proposes SVEET, a framework that requires merely training on a pretrained bidirectional video diffusion model but supports high-quality streaming video editing in an auto-regressive fashion and proposes a decoupled training scheme that explicitly enforces the orthogonality between the optimization directions of video controllability and model causality.

Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention
Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention
arXiv:2609.24797 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

表明 Kimi Delta Attention (KDA) 可通过单次 delta-rule 变换与其通道门控提供的反射组合实现 2D 旋转,并刻画 CKDA 的表达能力,证明每个正交的对角加秩-1矩阵恰为一个 CKDA 转移矩阵。This work shows that Kimi Delta Attention (KDA) can realize 2D rotations by combining a single delta-rule transformation with a second reflection supplied by its channel-wise gate, and characterize the expressivity of CKDA and prove that every orthogonal diagonal-plus-rank-one matrix is exactly a CKDA transition matrix.

Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene
Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene
arXiv:2609.23796 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Mira-Scene,一种组合式 3D 场景重建框架,将稀疏姿态回归替换为密集有界对应恢复,并引入多模态扩散 Transformer 联合生成物体几何与 CCM,使用模态专精的专家流配合共享注意力与位置编码以促进几何-布局一致性。Mira-Scene is presented, a compositional 3D scene reconstruction framework that replaces sparse pose regression with dense, bounded correspondence recovery and introduces a multimodal diffusion transformer that jointly generates object geometry and CCMs, using modality-specific expert streams with shared attention and positional encoding to promote geometry-layout consistency.

Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations
Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations
arXiv:2609.22255 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Deep Persona,一种心理学驱动的三层架构,将 persona 组织为可观察表达、潜在信念与核心动机驱动的层级结构,用于构建高可信度的角色扮演 Agent。Deep Persona is introduced, a psychologically grounded, three-layered architecture that organizes personas into hierarchical levels of observable expression, latent beliefs, and core motivational drives, for constructing highly convincing role-playing agents.

ACLArena: Agent Continue Learning in Multi-stage Post-training
ACLArena:多阶段后训练中的 Agent 持续学习
arXiv:2609.23989 工程化 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ACLArena,一个全面研究、分析与评估 Agent 持续学习(ACL)的框架,并提出新的 ACL 方案:将高质量轨迹的离线回放与多个由 RL 专精化的 LoRA 专家路由网络相结合,显著提升 Agent 跨多领域学习的能力。This work introduces ACLArena, a framework for comprehensively studying, analyzing, and evaluating Agent Continual Learning, and proposes a new ACL recipe that combines offline replay over high-quality trajectories with a routed network of multiple LoRA experts each specialized via RL, substantially improving the agent's ability to learn across multiple domains.

12. KVP:RL 驱动 KV Cache 驱逐策略
arXiv:2602.10238 LLM 基础设施 方法 Open MIND OA · 绿色 被引 12 · S2

本文提出 KV Policy (KVP),一种仅基于 key 与 value 向量、在预计算生成轨迹上训练的轻量级 per-head RL Agent 框架,证明学习预测未来 token 效用是自适应 KV cache 管理中强大且可扩展的范式。KV Policy (KVP), a framework of lightweight per-head RL agents trained on pre-computed generation traces using only key and value vectors, is introduced, demonstrating that learning to predict future token utility is a powerful and scalable paradigm for adaptive KV cache management.

Towards Full Pipeline FP8 Reinforcement Learning for LLMs
面向 LLM 的全流水线 FP8 强化学习
arXiv:2609.22870 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Calibrated Clipping,一种动态方法,通过匹配下界裁剪分位数并相应重平衡上界,将 FP8 裁剪边界与高精度 BF16 分布对齐,消除熵激增并恢复与 BF16 基线可比的性能。Calibrated Clipping is proposed, a dynamic method that aligns the FP8 clipping bounds with high-precision BF16 distributions by matching the lower-bound clipping quantile and rebalancing the upper bound accordingly, which eliminates entropy surges and restores performance comparable to the BF16 baseline.

UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing
UltraTex:释放 2K 多视角扩散模型用于 3D 纹理生成
arXiv:2609.23169 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 UltraTex,一种用于高分辨率多视图扩散 3D 纹理生成的高效端到端框架,并引入 Background Token Dropping(在 DiT 主干前移除背景 token)与 Block-Sparse Attention(降低前景序列上的注意力计算)。This paper presents UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing and introduces Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequence.

The Functionalizer: Lossless Functional Decomposition for Subword Tokenization
The Functionalizer:面向子词分词的无损函数分解
arXiv:2609.15991 RAG 检索增强 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Functionalizer,一种无损预分词框架,将正字与结构变体分解为分词前的可组合操作码/操作数前缀流:以 Unicode 私有使用区编码的参数化变换操作符(操作码)为前缀,连接规范基础 token(操作数)。The Functionalizer is presented, a lossless pre-tokenizer framework that factors orthographic and structural variations into a compositional opcode/operand prefix stream before tokenization: a canonical base token (operand) prefixed by parametric transformation operators (opcodes) encoded in the Unicode Private Use Area.

A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal
大语言模型的测谎测试:读出模型不愿透露的知识
arXiv:2609.21996 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

内部识别探针(Probe of Internal Recognition, PIR)可区分"不愿回答"与"无法回答"的模型,支持装傻审计与遗忘验证,并从选择题扩展至自由生成。Probe of Internal Recognition (PIR) separates a model that will not answer from one that cannot, which supports sandbagging audits and unlearning verification, and extends from multiple-choice questions to free-form generation.

TAPe+ML: A Compact Structured Representation for Multi-Task Computer Vision
TAPe+ML:面向多任务计算机视觉的紧凑结构化表示
arXiv:2609.20869 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,将部分建模负担从网络参数转移至结构化输入表示,可在更低的数据、内存与算力需求下支持紧凑的多任务视觉系统。The results suggest that shifting part of the modeling burden from network parameters to a structured input representation can support compact multi-task vision systems with reduced data, memory, and compute requirements.

SkillSpec: Intent-Masked Specification Reasoning for Agent Skill Correctness
SkillSpec:面向 Agent Skill 正确性的意图掩码规约推理
arXiv:2609.06052 Agent 智能体 方法 OA · 绿色 被引 1 · S2

提出 SkillSpec,一种 Hoare 风格框架,将技能正确性建模为规约推理问题,并将异构技能仓库转化为统一的图表示,对齐描述、指令与代码构件。This work proposes SkillSpec, a Hoare-style framework that formulates skill correctness as a specification reasoning problem, and transforms a heterogeneous skill repository into a unified graph representation that aligns descriptions, instructions and code artifacts.

StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training
StableVQ:稳定的 Vector-Quantized Tokenizer 训练实用指南
arXiv:2609.26774 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 StableVQ,重新审视每个模块的合适学习目标,解决各模块独立训练以承担各自角色时产生的问题;其基于共享投影 codebook 构建,轻量且不引入可学习参数。StableVQ is proposed, which revisits the proper learning objective of each module and resolves the problems that arise when each is trained to fulfill its own role independently, and Built on top of shared-projection codebooks, is lightweight and introduces no learnable parameters.

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks
Tasteful Agent:长程任务中品味(taste)的度量与改进
arXiv:2609.25804 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

构建 Taste-Bench,一个由 Agent 在工程与研究任务中产生的轨迹自动构建的品味问题基准,并证明品味是可训练的。Taste-Bench is built, a benchmark of taste questions constructed automatically from trajectories that agents produced in engineering and research tasks, and it is shown that taste can be trained.

All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
脚本感知的 Mixture-of-Experts 统一多语言场景文本识别
arXiv:2609.24058 多模态 应用落地 OA · 绿色 被引 1 · S2

本文构建了 TextMuSS-10M,一个涵盖 10 种文字、229 种语言的大规模合成场景文本数据集,并提出了 ScriptMoE,一种具备文字感知能力的 Mixture-of-Experts (MoE) 架构。该架构在精度上达到最高,且比 per-language experts 更简单、比 VLM 更轻量,同时精度优于两者。This work constructs TextMuSS-10M, a large-scale synthetic scene text dataset spanning 10 scripts and 229 languages and proposes ScriptMoE, a script-aware Mixture-of-Experts (MoE) architecture that achieves the highest accuracy and is simpler than per-language experts, lighter than VLMs, and more accurate than both.

11. AgenticRAGTracer(arXiv 2602.19127)
11. AgenticRAGTracer(arXiv 2602.19127)
arXiv:2602.19127 Agent 智能体 评测集 OA · 绿色 被引 5 · S2

本文提出 AgenticRAGTracer,这是首个主要由大语言模型自动构建、专为支持逐步验证而设计的 Agentic RAG 基准。AgenticRAGTracer is introduced, the first Agentic RAG benchmark that is primarily constructed automatically by large language models and designed to support step-by-step validation, and is primarily constructed automatically by large language models and designed to support step-by-step validation.

Lean Pool: An AI-Maintained Archive of Formalized Mathematics
Lean Pool:AI 维护的形式化数学档案库
arXiv:2609.25199 Agent 智能体 方法 OA · 绿色 被引 1 · S2

Lean Pool 是一个形式化数学仓库,由 AI agent 进行生长、维护和优化。Lean Pool is a repository of formalized mathematics. It is grown, maintained and optimized by AI agents.

Emergent Collusion in Long-Horizon LLM Agent Interaction
长程 LLM Agent 交互中的涌现合谋
arXiv:2609.24967 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

总体而言,研究结果表明,长期交互会以产生安全风险的方式重塑 Agent 的协作模式;限制 Agent 可获取的交互历史数量与范围能够减少串通行为。Overall, the findings show that long-horizon interaction can reshape how agents coordinate in ways that create safety risks, and restricting the amount and scope of interaction history available to agents reduces collusion.

Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes
三维场景中交互理解的几何与语义耦合
arXiv:2609.25247 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出了 SEGMENT-SNAP,通过 part-handle coupling 融合几何与语义证据,在 Articulate3D Challenge 中获得第一名。This work presents SEGMENT-SNAP, which combines geometric and semantic evidence through part-handle coupling and achieved first place in the Articulate3D Challenge.

ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning
ALPINE:面向参数与样本高效少样本学习的自适应定位
arXiv:2609.22323 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种面向 few-shot 图像分类的超轻量级空间-关系架构,结合固定的 Gabor 边缘能量引导与窗口化的、内容自适应的 patch locator。实验表明,该架构中虽然存在 pairwise relational computation,但它并非性能的主要驱动因素;真正起决定作用的是 content-adaptive patch locator。An ultra-lightweight spatial-relational architecture for few-shot image classification that combines fixed Gabor edge-energy guidance with a windowed, content-adaptive patch locator is presented, showing that the architecture's pairwise relational computation, while present, is not the primary driver of its performance; the content-adaptive patch locator is.

Blaming Across the Aisle: Political Contrasting and Blame Attribution in the Danish Parliament
跨越党派指责:丹麦议会中的政治对比与归责
arXiv:2609.26346 LLM 基础设施 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

政治话语日益呈现更强的敌对感是普遍观感,但稳健证据仍稀缺。本文研究 1997 至 2026 年丹麦议会的归责行为,结合专门构建的分类器 BlameBERT(F1: 0.80)与多层统计建模。该分类器采用面向低至中等资源语言的标注高效流程。结果显示出一条香蕉形轨迹:归责水平约在 2016 年前下降,随后在近年(2019–2026)进入显著且持续的上升阶段。执政地位显著Political discourse is widely perceived to be growing more hostile, yet robust evidence remains scarce. This study examines blame attribution in the Danish Parliament from 1997 to 2026, combining a purpose-built classifier, BlameBERT (F1: 0.80), with multilevel statistical modeling. The classifier is constructed using an annotation-efficient pipeline for blame attribution in low-to-mid resource languages. The results reveal a banana-shaped trajectory, with blame declining until around 2016 before entering a significant and sustained increase in recent years (2019-2026). Government status consi

Agensh: Scaling Organizational Intelligence to 1,024 Agents
Agensh:将组织智能扩展到 1,024 个智能体
arXiv:2609.26781 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文揭示 Agent 数量是多 Agent 组织扩展通用智能边界的新 scaling 维度,为硬延迟约束或时间预算下的复杂任务提供了实用方案。The number of agents is revealed as a new scaling dimension for multi-agent organizations to expand the frontier of general intelligence, offering a practical solution for complex tasks under hard latency constraints or time budgets.

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
JEV-as-a-Judge:自信则接受,不确定则升级
arXiv:2609.26550 评测基准 方法 OA · 绿色 被引 14 · S2

LLM-as-a-judge 支持跨任务评估,但在大规模场景下推理成本与置信度可靠性成为关键问题。本文研究仅做判断的 judge 能否提供经济高效的首轮判别,并识别何时需要更强的评估。在与十六种生成式与奖励模型 judge 的对比中(采用盲法人类裁定作为参照),我们发现 jev-as-a-judge 在普通偏好与有证据支撑的事实性任务上,与作为最强对照的 SOTA LLM judge 仅相差 3 个百分点,成本仅为后者的 0.36%。在需要核查LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generative and reward-model judges, with blinded human adjudication, we find it within three percentage points of a state-of-the-art LLM judge, our strongest comparator, on ordinary preference and evidence-grounded factuality at 0.36% of the comparator's fee. Larger gaps arise when judgments require checking

ImIR: Image-Instruction Tuning for All-in-One Image Restoration
ImIR:面向一体化图像修复的图像-指令微调
arXiv:2609.25267 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

一种近期且有效的方法:通过一个小型 low-rank adapter,将大型预训练图像编辑模型适配到图像恢复任务,并以源自退化图像本身的 instruction 替代 text prompt;在同等条件下该方法优于文本条件。A recent and effective recipe adapts a large pretrained image-editing model to restoration using a small low-rank adapter with a text prompt to an instruction derived from the degraded image itself that outperforms text conditioning under a matched comparison.

LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay
LatentPort:超越 KV Cache——混合语言模型中循环记忆的跨模型迁移:无需目标前缀重放的 4B 到 9B 混合状态交接
arXiv:2609.25053 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文展示了在一对架构匹配的 Qwen3.5 4B-to-9B 兄弟模型之间实现有用的持久 hybrid-state transfer,是首次在大小不同的混合语言模型之间完成持久循环推理状态的跨模型交接,且无需 target prefix replay。This work demonstrates useful persistent hybrid-state transfer across one architecture-matched Qwen3.5 4B-to-9B sibling pair, the first demonstrated cross-model handoff of persistent recurrent inference state between differently sized hybrid language models without target prefix replay.

10. TritonForge: Automated Triton Kernel Optimization (arXiv 2512.09196)
10. TritonForge:自动化 Triton Kernel 优化(arXiv 2512.09196)
arXiv:2512.09196 LLM 基础设施 方法 OA · 绿色 被引 20 · S2

TritonForge 是一个面向自动化 Triton kernel 优化的 profiling 引导框架,融合 kernel 分析、运行时 profiling 与迭代式代码转换以简化优化流程,并为自动化 GPU 性能优化领域的未来研究奠定基础。TritonForge, a profiling-guided framework for automated Triton kernel optimization that integrates kernel analysis, runtime profiling, and iterative code transformation to streamline the optimization process and provides a foundation for future research in automated GPU performance optimization.

Embedding Physics Priors in Robot Learning: A Survey
将物理先验嵌入机器人学习:一篇综述
arXiv:2609.22319 RAG 检索增强 综述 OA · 绿色 被引 0 · S2 + OpenAlex

论文论证了物理先验提供了一种与机器人领域高度相关的归纳偏置,其作用是与数据驱动学习互补而非取代,从而为更通用、数据高效且可信的机器人系统铺平道路。It is argued that physics priors provide a particularly relevant robotics-specific inductive bias, complementing rather than replacing data-driven learning, and paving the way toward more generalizable, data-efficient, and trustworthy robotic systems.

Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts
Tri-PvP:通过感知-命题证据冲突揭示 Omni-Modal LLM 的模态偏差
arXiv:2609.06011 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 Tri-PvP,一个包含 8,000 样本、跨越视觉、音频与文本的三模态冲突基准,揭示了证据形式偏倚的系统性不对称:模型在视觉上更偏向感知信号,而在音频上更偏向命题性信号。Tri-PvP, an 8,000-sample tri-modal conflict benchmark crossing vision, audio, and text, is introduced, revealing a systematic asymmetry in evidence-form bias: models exhibit a stronger bias toward perceptual signal in vision but propositional in audio.

Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
Just-in-Time Memory:面向 LLM Agent 的任务自适应记忆策展学习
arXiv:2609.27334 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Just-in-Time Memory (JitMem) 一致优于无 memory 的 Agent 以及启发式与学习式写入 memory 方法,相较最强基线在成功率上分别提升了 16.2、16.3 和 3.9 个绝对百分点。Just-in-Time Memory (JitMem) consistently outperforms no-memory agents as well as heuristic and learned write-time memory methods, improving over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively.

EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics
EmbodiedSWE:面向长程灵巧机器人任务的 Coding Agent
arXiv:2609.27308 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

EMBODIEDSWE-GEN 将 coding agent 的单一解决方案扩展为大规模多样化轨迹用于训练 VLA,并表明仅在 coding agent 生成的仿真演示上微调的 VLA,即可在真实机器人上完成长时任务。EMBODIEDSWE-GEN, which expands a single solution from coding agent into large diverse trajectories for training a VLA, and shows that a VLA finetuned solely on coding-agent-generated simulation demonstrations completes a long-horizon task on real robot.

StudentBench: AI and human tutoring yield equivalent GRE learning gains
StudentBench:AI 辅导与人类辅导在 GRE 学习成效上相当
arXiv:2609.28470 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

论文证实,AI 辅导在 GRE 学习收益上与专家人类辅导具有统计等效性(p = .015),且在 GRE 的七个领域中,有五个领域的最佳 AI 导师平均超越了人类导师。It is established that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average.

HappyWorld-Bench
HappyWorld-Bench
arXiv:2609.24308 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 HappyWorld-Bench,一个用于评估生成世界在 Agent 交互过程中是否保持可靠的综合基准,并强调对世界模型的评估不仅应看视觉质量,还应考察状态一致性以及其对动作和干预响应的正确性。HappyWorld-Bench is introduced, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them, and highlights the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.

PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing
PackLab:面向机器人装箱场景中 MLLM 开发、训练与评估的综合框架
arXiv:2609.23784 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

PackLab 是一个用于开发、训练与评估闭环机器人装箱 MLLM 的综合框架,在不同物体集合与容器配置下均优于传统装箱启发式方法、经典强化学习方法以及通用 MLLM,展现了 MLLM 在长时任务机器人装箱中的潜力。PackLab is a comprehensive framework for developing, training, and evaluating MLLMs for closed-loop robotic bin packing that outperforms conventional packing heuristics, traditional reinforcement learning methods, and general-purpose MLLMs across object sets and container configurations, highlighting the potential of MLLMs for long-horizon robotic packing.

Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World
Spatial-Interactor:通过与可观测物理世界的交互学习空间推理
arXiv:2609.23038 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 Spatial-Interactor,一个通过交互训练 VLM 建模物理世界状态转换的框架,并将学习过程组织为三级课程:L1 被动世界状态转换、L2 主动自我状态转换、L3 长时交互轨迹。Spatial-Interactor is introduced, a framework that trains VLMs to model physical-world state transitions through interaction, organizing this learning process into a three-level curriculum covering L1 passive world-state transitions, L2 active self-state transitions, and L3 long-horizon interaction trajectories.

10. Cloud Native System for LLM Inference Serving(arXiv 2507.18007)
10. 面向 LLM 推理服务的 Cloud Native 系统(arXiv 2507.18007)
arXiv:2507.18007 LLM 基础设施 方法 OA · 绿色 被引 7 · S2

本文探讨容器化、微服务、动态调度等 Cloud Native 技术如何从根本上提升 LLM 推理服务,并展示 Cloud Native 系统在高需求场景下实现更高效资源分配、降低延迟与提升吞吐的能力。This article explores how Cloud Native technologies, such as containerization, microservices, and dynamic scheduling, can fundamentally improve LLM inference serving and demonstrates how a Cloud Native system enables more efficient resource allocation, reduces latency, and enhances throughput in high-demand scenarios.