研究库 论文知识库
Papers · organized/paper_cards

论文

1640 张论文卡片 · OA 绿色

开放获取 全部 绿色 · 1640
FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
FoMo:生成轨迹中的分叉时刻作为感知距离
arXiv:2609.25716 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一条全自动的数据生成流程,可在无需任何人工标注的情况下生成图像对之间的逐点感知距离标签,并证明 diffusion trajectory 与人类视觉系统高度一致。This paper proposes a fully automated data generation pipeline that generates pointwise perceptual distance labels between image pairs without any human annotation, and demonstrates that the diffusion trajectory aligns well with the human visual system.

CARD: Cluster-level Adaptation with Reward-guided Decoding for Personalized Text Generation
[标题中文] CARD:面向个性化文本生成的基于聚类级自适应与奖励引导解码
arXiv:2601.06352 工程化 应用落地 OA · 绿色 被引 2 · S2

本文提出 CARD,一种通过渐进式细化实现有效个性化的层级框架:先按共享风格模式对用户聚类,再为各组学习专用的 LoRA adapter,从而在低资源场景下也能实现稳健的泛化与强劲的性能。This work presents CARD, a hierarchical framework that achieves effective personalization through progressive refinement that first clusters users according to shared stylistic patterns and learns group-specific LoRA adapters, enabling robust generalization and strong low-resource performance.

Do Implicit Personalization and Explicit Styles Conflict? PsPLUG: A Lightweight Plug-in for Balancing Personalization and Style in Customized LLMs
[标题中文] 隐式个性化与显式风格是否冲突?PsPLUG:用于在定制化 LLM 中平衡个性化与风格的轻量级插件
arXiv:2601.06362 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

本文提出 PsPLUG,一个轻量级插件式模块,在刻画目标风格后学习用户专属的残差,能更好地保留用户偏好,并对个性化与风格遵循之间的平衡提供更精确的控制。PsPLUG is proposed, a lightweight plug-in that learns a user-specific residual after accounting for the requested style, and better preserves user preferences while providing precise control over the balance between personalization and style adherence.

Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching
[标题中文] 通过连续深度批处理实现循环语言模型的深度自适应推理
arXiv:2608.09444 LLM 基础设施 方法 OA · 绿色 被引 5 · S2

本文提出首个通过 continuous depth batching (CDB) 实现 depth-adaptive looped LM 的高效方法:在 loop 步骤之间重组 batch,动态调度架构中的 looped 与非 looped 部分,管理 looped KV-caching,并提前预测将退出 loop 的 token,以便异步准备 batch。This work introduces the first efficient method for depth-adaptive looped LMs via continuous depth batching (CDB), which forms new batches between loop steps, and dynamically schedules looped and non-looped parts of the architecture, manages looped KV-caching, and predicts which tokens will exit the loop in advance so it can prepare batches asynchronously.

Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations
[标题中文] 气候政策因果评估中识别假设的循证审计
arXiv:2609.30867 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

[TLDR中文] 双重差分(DID)研究被广泛用于评估气候政策,但评估支持其识别假设的证据仍具挑战。我们提出 ARGUS,一个结构化的语言模型流水线,针对十一维的假设—含义—证据评估标准对所报告的证据进行审计,并在无法检索到相关证据时选择弃答。我们通过注入缺陷、经济学论文以及一个使用经协调标签的小规模试点评估 ARGUS。在 11 类缺陷基准上,ARGUS 检测出 73% 的植入缺陷,而基于关键词的流水线仅能检测出 18%。Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-flaw benchmark, ARGUS detects 73% of planted flaws, compared with 18% for a keyword-based pipeline. Acr

MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation
[标题中文] MOPD-Router:重新思考多教师在策略蒸馏中的教师路由
arXiv:2609.30837 工程化 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ExpertAlign,一种无需域标签或独立 routing 模型的框架,可在每个 token 上对全体 teacher 池进行监督路由;研究表明 token 级路由能够利用跨域互补监督,并减少对 prompt 级域指派的单一依赖。This work proposes ExpertAlign, a framework that routes supervision over the full teacher pool at each token, without domain labels or training a separate routing model, and demonstrates token-level routing can exploit cross-domain complementary supervision, and reduce exclusive reliance on prompt-level domain assignment.

Systems 补充候选
arXiv:2511.02230 Agent 智能体 方法 Open MIND OA · 绿色 被引 68 · S2

Continnum,一种通过为 KV cache 保留引入 TTL 机制来优化多轮 Agent 工作负载任务完成时间的服务系统,与程序级 FCFS 结合时可保持多轮连续性,并降低 Agent 工作流的延迟。Continnum, a serving system to optimize job completion time for multi-turn agent workloads by introducing time-to-live mechanism for KV cache retention, and when combined with program-level first-come-first-serve preserves multi-turn continuity, and reduces delay for agentic workflows.

IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking
[标题中文] IndicBankBench:评估语言模型助手在印度零售银行中的安全性与可靠性
arXiv:2609.29167 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 IndicBankBench,一个涵盖五个业务领域、一个能力/拒答领域以及二十个主轴、共 799 个案例的零售银行 benchmark,并提供 case 级诊断分析,以区分那些提出不必要追问的系统与那些采取行动却未能调和客户上下文或完整解决诉求的系统。IndicBankBench is introduced, a 799-case benchmark for Indian retail banking spanning five operational domains, a capability/refusal domain, and twenty primary axes, and a case-level diagnostics that distinguish systems that ask unnecessary questions from those that act but fail to reconcile customer context or fully resolve the request.

Paragraph Boundaries Are Not White Space:Compression Depth as the Signature of Hierarchical Structure
[标题中文] 段落边界并非空白:压缩深度作为层级结构的特征
arXiv:2609.23551 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

使用一种分层旋转位置编码(hRoPE),将段落、句子和 token 索引表示为独立通道,保持 token 序列固定,对段落坐标 $p_1$ 进行干预,并使用 token 距离精确估计器测量跨段落注意力。A hierarchical rotary positional encoding (hRoPE) that represents paragraph, sentence, and token indices as separate channels, hold the token sequence fixed, intervene on the paragraph coordinate $p_1$, and measure cross-paragraph attention with a token-distance-exact estimator is used.

BoundInk: Boundary-Aware Online Handwriting Generation
[标题中文] BoundInk:面向边界感知的在线手写生成
arXiv:2604.02103 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

BoundInk 是一种以书写者为条件的框架,将字符间的边界视为显式生成单元,既能保留书写者特有的字形外观,又能在完整文本行内改善连接性与字距。BoundInk is introduced, a writer-conditioned framework that treats inter-character boundaries as explicit generation units and preserves writer-specific glyph appearance while improving connectivity and spacing across complete text lines.

VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
[标题中文] VLA-Precision:面向视觉-语言-动作模型高效真实世界在线强化学习的非对称协同自举
arXiv:2609.04355 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 VLA-Precision,一个面向真实世界的高效在线 RL 框架,包含 Asymmetric Co-Bootstrapping (ACoB) 算法与 ACoB-Stream 架构;其中 ACoB-Stream 以不变状态解耦与按需流式传输为设计原则,构建了经验-策略的闭环架构,可实现最高 10.9% 的吞吐与计算效率提升。VLA-Precision, an efficient real-world online RL framework featuring the Asymmetric Co-Bootstrapping (ACoB) algorithm and the ACoB-Stream architecture, and ACoB-Stream, a closed-loop experience--policy architecture that establishes invariant-state decoupling and on-demand streaming as design principles, deliver up to 10.9% improvements in throughput and computational efficiency.

QReason: Query-Focused Decoupled Chain-of-Thought for Efficient Passage Reranking
QReason:面向高效段落重排序的查询聚焦解耦思维链。
arXiv:2609.30904 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

QReason 是一种解耦框架,将面向 query 的推理与针对窗口的 passage 相关性评估分离,显著减少冗余推理,在取得与 reasoning-based reranker 相当乃至更优的排序性能的同时,超越了现有的 query rewriting 模型。QReason is a decoupled framework that separates query-focused reasoning from window-specific passage relevance assessment, and significantly reduces redundant reasoning, achieves ranking performance comparable to or better than strong reasoning-based rerankers, and outperforms existing query rewriting models.

EngramRAG: Dynamic Usage-Weighted Topology and Synaptic Consolidation for Multi-Hop Agentic Memory
EngramRAG:面向多跳 Agentic 记忆的动态使用加权拓扑与突触巩固。
arXiv:2609.32049 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

EngramRAG 是一种自适应记忆架构,将低延迟的 Waking State 反射与异步后台 Dreaming State 整合周期耦合,并引入融合密集向量、BM25 与 U-PPR 的三源混合检索,通过动态 Reciprocal Rank Fusion (RRF) 实现。The proposed EngramRAG is an adaptive memory architecture coupling a low-latency Waking State reflex with an asynchronous background Dreaming State consolidation cycle, and introduces triple-source hybrid retrieval fusing dense vectors, BM25, and U-PPR via dynamic Reciprocal Rank Fusion (RRF).

Evaluating the accuracy of KV cache reuse techniques
评估 KV cache 复用技术的准确性。
arXiv:2609.31415 LLM 基础设施 观点 OA · 绿色 被引 0 · S2 + OpenAlex

Boxoffice 是一个程序化生成评测数据集的工具,用以覆盖具有挑战性的 KV cache 复用模式;研究表明现有数据集并未展现出充分评估此类技术所需的复用动态特性。Boxoffice, a tool that programmatically generates evaluation datasets that exercise challenging KV cache reuse patterns, is introduced and it is shown that existing datasets do not exhibit the reuse dynamics needed to thoroughly evaluate such techniques.

Intent2Tc: Automated Intent-to-Traffic Control Translation with Language Models
Intent2Tc:基于语言模型实现意图到流量控制的自动化翻译。
arXiv:2609.31397 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

Intent2Tc 是一个由语言模型驱动的闭环框架,将业务级流量整形意图转换为声明式子意图,进而生成经过验证的可执行 Linux traffic control (tc) 配置,并展示了该框架的实际适用性。Intent2Tc is presented, a closed-loop language-model-driven framework that translates business-level traffic-shaping intents into declarative sub-intents and subsequently into validated, executable Linux traffic control (tc) configurations and demonstrates the practical applicability of the proposed framework.

Softmax Reparameterization for Output-Head Quantization
面向输出头量化的 Softmax 重参数化。
arXiv:2609.31291 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 softmax reparameterization,一种训练后方法,在量化前搜索功能等价的输出头,并展示了在总体 logit 误差增大的情况下保真度仍可提升。This work introduces softmax reparameterization, a post-training method that searches over functionally equivalent output heads before quantization and shows how fidelity can improve despite greater total logit error.

LightMIS: Ultra-Lightweight Medical Image Segmentation Without a Stage-Wise Decoder
LightMIS:无逐级解码器的超轻量医学图像分割。
arXiv:2609.28327 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

LightMIS 通过 Scale-Aligned Projection 块将五级编码器的输出对齐至统一分辨率,仅聚合一次,并利用结合 Adaptive Kernel Fusion 与所提 Progressive Receptive Fusion 模块的 Adaptive Fusion Cascade 精炼融合表示。LightMIS aligns the outputs of a five-level encoder to a common resolution using Scale-Aligned Projection blocks, aggregates them once, and refines the fused representation with an Adaptive Fusion Cascade, which combines Adaptive Kernel Fusion with the proposed Progressive Receptive Fusion module.

Systems 补充候选
arXiv:2606.01751 RAG 检索增强 方法 OA · 绿色 被引 3 · S2

SarseX 模型无关、无需训练,并与 Prefix Cache 兼容,可为多轮对话、检索增强生成 (RAG) 和 Agent 工作流等常见在线服务场景提供统一支持。SarseX is model-agnostic, training-free, and compatible with Prefix Cache, and it provides unified support for common online serving scenarios including multi-round chat, retrieval-augmented generation (RAG), and agent workflows.

BoundaryMORPH: Budgeted Reranking via Active Set Selection for Diffuse Retrieval
BoundaryMORPH:通过主动集合选择实现预算化重排序的扩散检索
arXiv:2609.27213 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

BoundaryMORPH 是一种新算法,专门为 LLM 的上下文容量 k 分配 CE 预算,在多个模型与数据集上针对开放性查询取得了 SOTA 集合检索质量。BoundaryMORPH is introduced, a novel algorithm that allocates CE budget specifically for the LLM's context capacity $k$ and achieves state-of-the-art set retrieval quality across multiple models and datasets with open-ended queries.

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance
SAGE:通过拓扑引导缓解长程推理偏置
arXiv:2609.30192 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 SAGE(Structural Admissibility-Guided Exploration),一个通过注入结构引导来缓解长程推理中探索偏差与累积偏差的统一框架,在 Andrews-Curtis 问题上取得了最高 8 倍的提升。This work proposes SAGE (Structural Admissibility-Guided Exploration), a unified framework that injects structural guidance to alleviate exploration bias and compounding bias in long-horizon reasoning and achieves up to an 8-fold improvement on the Andrews-Curtis problem.

LastOPD: Taming Collapse in Latent On-Policy Distillation
LastOPD:驯服潜在 On-Policy 蒸馏中的坍缩问题
arXiv:2609.28845 多模态 方法 OA · 绿色 被引 2 · S2

LastOPD 仅在最后一层状态(LM head 共同读取的接口)施加潜在信号,且仅在进入 token 级 OPD 前的 10 步交叉淡入阶段施加,从而保留潜在信号中的有用部分,并在坍缩发生前将 student 交由 token 级监督。LastOPD is proposed, which applies the latent signal only at the last-layer state, the common interface both LM heads read, and only during a 10-step crossfade into token-level OPD, which keeps the useful part of the latent signal and hands the student to token-level supervision before the collapse sets in.

TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent
TimeEvo:时序 Agent 的失败驱动自演化
arXiv:2609.27277 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

TimeEvo 将 Agent 诊断出的失败聚类为能力缺口,为每个缺口规划测量,合成只产出证据的工具来填补缺口,并仅通过配对的准入门控接纳候选工具库。TimeEvo is proposed, which clusters an agent's diagnosed failures into capability gaps, plans a measurement for each, synthesizes evidence-only tools that fill them, and admits the candidate library only through a paired admission gate.

Disaggregated Quantization: Specializing LLM Prefill and Decode
解耦量化:针对 LLM Prefill 与 Decode 的专用量化
arXiv:2609.26333 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

评估 vLLM 在分离式 serving 下的精度,并通过训练后量化在最高 2.8T 参数的模型上进一步验证共享权重格式的分离式 serving。Assessment of accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters are evaluated.

D-JEPA: A Decision-Aligned Latent World Model
D-JEPA:一种决策对齐的潜在世界模型
arXiv:2609.24749 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

D-JEPA 是一种决策对齐的潜在世界模型,从已执行结果中学习候选未来之间的决策相关关系,将决策相关关系结构确立为预测世界建模与有效控制之间的直接桥梁。D-JEPA is introduced, a decision-aligned latent world model that learns decision-relevant relations among candidate futures from executed outcomes, establishing decision-relevant relational structure as a direct bridge between predictive world modeling and effective control.

Just-In-Time Agent Memory with Runtime Agentic Research
基于运行时 Agentic 研究的 Just-In-Time Agent 记忆。
arXiv:2609.34385 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Just-In-Time Agent Memory (JAM) 是一个可训练框架,在运行时进行查询条件下的上下文构建,其任务性能优于 AOT 式记忆系统,同时比先前的可训练 Agent 记忆方法显著更高效。Just-In-Time Agent Memory (JAM), a trainable framework for query-conditioned context construction at runtime, is proposed, where it achieves stronger task performance than AOT-style memory systems while remaining substantially more efficient than prior trained agentic memory approaches.

Stashbird: Efficient Speaker-Indexed Memory for Conversational Agents
Stashbird:面向对话 Agent 的高效说话人索引记忆。
arXiv:2609.34242 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Stashbird 是一种 Agent 记忆系统,通过显式来源溯源将源事件链接到派生记忆状态,在 LongMemEval-S 和 GroupMemBench 上准确率高于 Hindsight,在 EverMemBench 上与之相当。Stashbird is presented, an agent memory system that links source episodes to derived memory state through explicit provenance and achieves higher accuracy than Hindsight on LongMemEval-S and GroupMemBench and comparable accuracy on EverMemBench.

InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video
InfiniHand:基于第一人称视频的流式世界空间手部运动估计。
arXiv:2609.35743 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

InfiniHand 是一个端到端流式前馈框架,联合估计 MANO 参数、相机轨迹与手部位置,直接从未标定的自我中心视频出发,相比 ViDiHand 在 ARCTIC PA-p 上降低 21%,并显著缓解世界空间漂移。InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, achieving a 21% reduction in ARCTIC PA-p compared to ViDiHand while substantially mitigating world-space drift.

GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space
GeoVerse: 几何潜空间中的世界一致新视角合成
arXiv:2609.35734 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

GeoVerse 是一个在世界一致条件下合成新视角的框架,在预训练 3D 基础模型的几何潜在空间内进行生成,并通过 ControlNet 风格 adapter 注入视频生成模型的外观先验。GeoVerse is a framework that synthesizes world-consistent novel views by performing generation within the geometric latent space of a pretrained 3D foundation model and injecting appearance priors from a video generative model via a ControlNet-style adapter.

Systems 补充候选
arXiv:2606.03910 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

NetKV,一种使用该 oracle 信息的 O(|D|) 每请求贪心策略,其层级排序被证明对过时遥测数据具有鲁棒性;并证明随着上下文长度增长,忽略网络项会使仅缓存感知的调度任意次优。NetKV, the O(|D|) per-request greedy that consumes this oracle, has tier rankings that are provably robust to stale telemetry, and it is proved that ignoring the network term renders cache-aware-only scheduling arbitrarily suboptimal as context length grows.

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
通过交错强化学习在统一模型中学习原生反思
arXiv:2609.35767 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

UMM-Reflection 在统一模型内利用强化学习(RL)完成完整反思轨迹:兄弟轨迹共享同一初始图像,因此组相对优势可比较不同反思策略,而单条轨迹级优势同时更新反思 token 与基于 flow 的修订,避免了逐轮信用分配的组合爆炸。UMM-Reflection is introduced, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment.

When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety
模型内部何时有用?探究表征工程在 LLM 安全中的作用
arXiv:2609.34771 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

总体而言,representation engineering 并不能普遍替代行为安全护栏,但在特定条件下具有实际优势,并可带来互补的安全收益。Overall, representation engineering does not generally replace behavioral safeguards, but offers practical advantages under specific conditions and can provide complementary safety benefits.

WideSWE: Can Coding Agents Coordinate Changes Across Repositories?
WideSWE: 编码 Agent 能否跨代码仓库协调变更?
arXiv:2609.33382 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 WideSWE,用于在跨仓库任务上评估编码 Agent,并在相同 prompt 下将其与联合执行进行对比,以考察逐仓库工作是否能缓解相关困难。This work introduces WideSWE to evaluate coding agents on cross-repository tasks, and compares it with joint execution under identical prompts to examine whether working on one repository at a time can alleviate difficulties.

Residual Transferability in Neural Image Watermarking
神经图像水印中的残差可迁移性
arXiv:2609.32241 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文识别出两种强化水印证据对载体图像依赖性、抑制残余可迁移性的机制,并提出 CoverLock,一种即插即用策略,可在不重新设计架构的前提下强化现有水印系统的图像依赖性。This work identifies two mechanisms that strengthen the dependence of watermark evidence on the cover image, thereby suppressing the residual transferability, and introduces CoverLock, a plug-and-play strategy for existing watermarking systems that strengthens such image dependence without architectural redesign.

FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching
FlowTool: 基于流匹配控制图像修图中的工具参数
arXiv:2609.35673 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 FlowTool,一个通过 conditional rectified flow 直接建模以输入图像和用户指令为条件的高质量工具参数分布的框架,并显著提升了推理效率。This work introduces FlowTool, a framework that directly models the distribution of high-quality tool parameters conditioned on the input image and user instruction using conditional rectified flow, and significantly improves inference efficiency.

Nereus: Adaptive Parallelism for LLM Post-Training
Nereus: 面向 LLM 后训练的自适应并行
arXiv:2609.34645 工程化 方法 OA · 绿色 被引 1 · S2

Nereus 是一种成本感知的运行时,将 RL 后训练任务适配为高效执行计划,并基于内存可行的全局计划执行状态转移,使用与运行任务校准的成本模型来接纳转移。Nereus is a cost-aware runtime that adapts RL post-training jobs into efficient execution plans and executes a transition using a memory-feasible global plan and admits the transition using a cost model calibrated against the running job.

Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision
难视觉,易视觉:GPT-6 Astra 揭示的计算机视觉全貌
arXiv:2609.35718 多模态 评测集 OA · 绿色 被引 1 · S2

本文勾勒了一幅计算机视觉版图:在其中,越来越复杂的视觉任务可通过通用接口访问,而精确且对保真度敏感的感知仍是重要前沿。A changing landscape of computer vision is mapped in which increasingly sophisticated visual tasks are accessible through a general-purpose interface, while precise and fidelity-sensitive perception remains an important frontier.