研究库 论文知识库
Papers · organized/paper_cards

论文

1130 张论文卡片 · 方法 · OA 绿色

开放获取 全部 绿色 · 1686
EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents
EMBL AI Librarian:面向 AI Agent 的生命科学知识层
arXiv:2607.28229 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出EMBL AI Librarian,一个升级Europe PMC接口的知识层,面向AI agent,提升多项任务表现:文献综合、claim验证、开放域问答,以及下游生物学任务如protocol问题与序列操作。EMBL AI Librarian is introduced, a knowledge layer that upgrades the Europe PMC interface for AI agents that improves performance across a range of tasks: literature synthesis, claim verification, open-domain question answering, and downstream biology tasks such as protocol questions and sequence manipulation.

Constitutional Midtraining: Content Presence Drives Alignment Gains
宪法式中训练:内容存在驱动对齐收益
arXiv:2607.26654 安全与风险 方法 OA · 绿色 被引 2 · S2

训练后对齐往往较浅,会在微调中被侵蚀。而中训练干预能否在干净隔离于训练后的情况下产生持久对齐,此前未经检验。我们通过宪法式中训练来测试:在 120B 规模上,插入基于原则与价值观的内容,与仅做回放的对照组进行对比。我们基于 Anthropic 的 Constitution 构建了 394M token 的宪法语料,并采用 2×2 析因设计(课程顺序 × 审慎推理),形成四种宪法式中训练条件与一组对照,随后在自生成与既有...Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce four constitutionally midtrained conditions plus a control, evaluated on self-generated and establis

UEmbed: Unified Sparse and Dense Multimodal Embeddings
UEmbed:统一的稀疏与密集多模态 Embedding
arXiv:2608.02583 RAG 检索增强 方法 OA · 绿色 被引 2 · S2

UEmbed (Unified Embedding)是一种decoder-only多模态嵌入模型,在单次因果前向中同时产出稀疏词项与稠密表示,提供新范式:在单一模型中统一稠密与稀疏嵌入,并将稀疏检索扩展以统一文本与多模态输入。UEmbed (Unified Embedding), a decoder-only multimodal embedding model that produces both sparse lexical and dense representations in one causal forward pass, offers a new paradigm: it unifies dense and sparse embeddings in one model, while further extending sparse retrieval to unify text and multimodal inputs.

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
LongHorizon-Harness:面向真实任务的长 Horizon Agent 推进
arXiv:2608.01964 Agent 智能体 方法 OA · 绿色 被引 9 · S2

将长视野执行重新表述为任务状态管理问题,提出LongHorizon-Harness,在执行外部显式维护任务状态,并仅用从环境中独立验证的事实更新它。This work reformulate long-horizon execution as a task-state management problem and proposes LongHorizon-Harness, which maintains the task state explicitly outside execution and updates it only with facts independently verified from the environment.

Progressive Agent Skill Generation via Reinforcement Learning
基于强化学习的渐进式 Agent Skill 生成
arXiv:2608.01678 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本工作提出 Skill-α,一种学习统一策略以进行渐进式 Skill 生成的强化学习方法,并引入一种新颖的回滚奖励,通过在锚定查询上比较原始技能与编辑后技能下的下游执行情况来评估每次编辑。This work proposesSkill-$\alpha, a reinforcement learning method that learns a unified policy for progressive skill generation and introduces a novel rollback reward that evaluates each edit by comparing downstream execution under the original and edited skills on an anchored query.

Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations
超越形态的运动:从抽象运动表示引导跨类别运动迁移
arXiv:2608.01628 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Motion Beyond Morphology(超越形态的运动)这一视角,旨在跨固定结构对应迁移运动,通过两阶段框架保留在不同目标形态间仍具有意义的动力学。This work introduces Motion Beyond Morphology, a perspective that seeks to transfer motion beyond fixed structural correspondence, by preserving dynamics that remain meaningful across different target morphologies by proposing a two-stage framework.

DAPD: Dual-Anchored Policy Distillation
DAPD:双锚点策略蒸馏
arXiv:2608.01735 LLM 基础设施 方法 OA · 绿色 被引 2 · S2

本文提出 DAPD,一种具有两级锚定的统一框架,可显著缓解特权错觉,在 Qwen3-4B 上以平均 +2.00 分优于 OPSD。DAPD is proposed, a unified framework with two levels of anchoring that significantly alleviates privilege illusion, outperforming OPSD on Qwen3-4B by +2.00 points on average across tasks.

3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering
3DZip:面向 3D 问答的空间感知特征多样性引导 token 压缩
arXiv:2608.01185 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 3DZip,一种三阶段 token 压缩框架:首先采用粗粒度体素化去除点级冗余,再通过 Determinantal Point Process 基于特征空间多样性选取锚点 token,最后在空间约束下融合剩余 token 以保持几何一致性。3DZip is proposed, a three-stage token compression framework that first applies coarse voxelization to remove point-level redundancy, then selects anchor tokens based on feature-space diversity via a Determinantal Point Process, and finally merges remaining tokens under spatial constraints to preserve geometric coherence.

CADENA: Stepwise CAD Reverse Engineering
CADENA:逐步式 CAD 逆向工程
arXiv:2608.00799 工程化 方法 OA · 绿色 被引 1 · S2

本文提出 CADENA(西班牙语意为"链"),一种将 3D 网格重建为参数化 CAD 程序的模型,按顺序逐个生成操作序列,并在每一步将目标与当前预测几何进行对比。This work introduces CADENA (Spanish for"chain"), a model that reconstructs a 3D mesh as a parametric CAD program, growing its sequence of operations one at a time and comparing the target with the currently predicted geometry at every step.

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation
LeapTalk:打破 Talking Head 生成中的延迟-质量权衡
arXiv:2608.00079 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 LeapTalk,一种通过单次前向实现稳定且实时说话头生成、可扩展至任意长视频的新颖框架,并引入音频驱动的无分类器引导机制,在极端步数缩减下保持细粒度唇形同步。LeapTalk is proposed, a novel framework that achieves stable and real-time talking-head generation with a single forward step, scaling to arbitrarily long videos, and an audio-driven classifier-free guidance mechanism to maintain fine-grained lip synchronization under extreme step reduction.

GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning
GradCuit:基于信用分配的梯度流实现鲁棒且可解释的测试时潜在推理
arXiv:2608.02585 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

GradCuit(梯度穿越电路)在所选 Transformer 层、提示隐藏表示与生成续写之间插入可优化的潜变量,开启了鲁棒且可解释的测试时缩放新维度,使 LLM 调整其推理方式,而不仅仅是重新生成、采样或重排输出。GradCuit (gradient through circuit), which inserts optimizable latent states at a selected Transformer layer between the hidden representations of the prompt and the generated continuation, opens a new axis of robust and interpretable test-time scaling, where LLMs adapt how they reason rather than merely regenerate, sample, or rerank outputs.

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
DeepVoyager-VL:激励视觉在环的搜索以应对长程多模态 Agent
arXiv:2608.01827 多模态 方法 OA · 绿色 被引 1 · S2

本文提出 DeepVoyager-VL,一种面向视觉在环搜索的长程多模态深度搜索框架,通过构建多模态事件图驱动数据合成,从而产出具有中间视觉依赖与长推理链的问题。DeepVoyager-VL is proposed, a long-horizon multimodal deep-search framework for vision-in-the-loop search that constructs a multimodal event graph to drive data synthesis, yielding problems with intermediate visual dependencies and long reasoning chains.

DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents
DreamTraj:通过读取未渲染的视频扩散潜在生成 6-DoF 物体轨迹
arXiv:2608.00486 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 DreamTraj,从单张 RGB 图像和任务指令预测物体 6-DoF 轨迹,推理时无需视频、深度或 CAD 模型,是首个直接从中间视频扩散表征(而非生成像素)解码物体 6-DoF 轨迹的方法This work proposes DreamTraj, which predicts a 6-DoF object trajectory from a single RGB image and a task instruction, requiring no video, depth, or CAD model at inference, and is the first approach to directly decode object 6-DoF trajectories from intermediate video diffusion representations rather than generated pixels.

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts
内部松弛、跨模态均衡:面向视觉-语言 Mixture-of-Experts 的几何引导负载均衡
arXiv:2608.00574 多模态 方法 OA · 绿色 被引 1 · S2

在四种拆分 backbone 上,ReBA 降低所有报告 benchmark 输入的负载,同时保持与 Std-Aux 相当的平均任务准确率,并在分辨率与分块变化下降低测试范围内的平均负载与最差物理负载Across four split backbones, ReBA lowers load on every reported benchmark input while keeping mean task accuracy comparable to Std-Aux, and lowers average load over the tested range and worst physical load under resolution and tiling shifts.

RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems
RecHarness:用于自演化推荐系统的 Bandit 路由 Agentic Harness
arXiv:2607.29241 评测基准 方法 OA · 绿色 被引 3 · S2

RecHarness 将优化过程分为两步:bandit 路由器根据历史验证反馈选择下一步修改方向,LLM 在选定方向内生成具体优化假设与可执行代码编辑RecHarness separates the optimization process into two steps: a bandit router selects the next modification direction according to historical validation feedback, while the LLM generates a concrete optimization hypothesis and executable code edit within the selected direction.

A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples
冻结的像素空间扩散模型可借助自身采样进行自我引导
arXiv:2607.29122 多模态 方法 OA · 绿色 被引 3 · S2

本文在中间层附加轻量预测头,保持 backbone 冻结,并利用中间预测与最终预测的差异作为采样时的自引导方向,训练一个能够自引导的冻结预训练像素扩散模型This work attaches a lightweight prediction head to an intermediate layer, keep the backbone frozen, and use the discrepancy between the intermediate and final predictions as a self-guidance direction during sampling to train a frozen, pretrained pixel diffusion model that can guide itself.

GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding
GPTQ-2D:立方时间的双侧自适应舍入
arXiv:2607.27042 LLM 基础设施 方法 OA · 绿色 被引 3 · S2

本文提出 GPTQ-2D,以三次时间复杂度生成相同的取整矩阵,并研究该任务的双侧版本——固定非奇异基矩阵同时作用于残差的左右两侧This work presents GPTQ-2D, which produces the identical rounded matrix in cubic time, and studies the two-sided version of this task, in which fixed nonsingular basis matrices act on both the left and the right of the residual.

Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge
标题 -> 标题中文:Wnuan:面向专有企业知识问答的分阶段后训练
arXiv:2608.01862 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Wnuan 三阶段流程:从文档构建任务导向监督,结合通用数据回放进行监督微调,对残余误差应用强化学习,阐述分阶段企业适配的收益与通用能力代价Wnuan, a three-stage pipeline that constructs task-oriented supervision from documents, performs supervised fine-tuning with general-data replay, and applies reinforcement learning to residual errors is presented, describing both the gains and the general-capability cost of staged enterprise adaptation.

Zero-Mem: Zero-Token Memory Operations for LLM Agents
标题 -> 标题中文:Zero-Mem:面向 LLM Agent 的零 token 记忆操作
arXiv:2607.29377 Agent 智能体 方法 OA · 绿色 被引 3 · S2

结果表明结构化 agent 记忆无需生成过去的中间表示,Zero-Mem 在消除记忆操作中 LLM 调用与 LLM-token 消耗的同时取得具有竞争力的性能The results show that structured agent memory need not generate an intermediate representation of the past, and Zero-Mem achieves competitive performance while eliminating LLM calls and LLM-token consumption from memory operations.

To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing
标题 -> 标题中文:增是机器,删是人工:LLM 代码编辑中删除回避的度量与缓解
arXiv:2607.28887 评测基准 方法 OA · 绿色 被引 1 · S2

在后训练中教授删除操作可减少删除回避行为并提升更广泛的代码编辑性能,表明该行为是训练不足而非不可达成Teaching deletion during post-training reduces deletion avoidance and improves broader code-editing performance, suggesting the behavior is undertrained rather than beyond reach.

MemSFT: Mitigating Alignment Tax with an External Parametric Memory
标题 -> 标题中文:MemSFT:借助外部参数化记忆缓解对齐税
arXiv:2607.25614 安全与风险 方法 OA · 绿色 被引 2 · S2

将 LLM 适配到专用领域常会带来对齐税:针对领域特定任务进行微调会导致灾难性遗忘,并显著降低在通用任务上的表现。我们提出 MemSFT,通过将领域专业化与主干参数更新解耦,以即插即用的参数化记忆来缓解对齐税。该记忆被训练为模仿在领域数据上运作的非参数化检索器,从而记住原本需通过检索获取的知识与模式。一旦在Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-specific tasks can cause catastrophic forgetting and substantially degrade performance on general tasks. We propose MemSFT, which mitigates the alignment tax by decoupling domain specialization from backbone parameter updates through a plug-and-play parametric memory. The memory is trained to imitate the behavior of a non-parametric retriever operating over domain data, thereby memorizing knowledge and patterns that would otherwise be accessed through retrieval. Once trained on

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
标题 -> 标题中文:看见还是知道?多模态大语言模型中的视觉上下文敏感性
arXiv:2607.26326 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

在作者研究的粗粒度属性上,MLLM 编码了视觉证据但无法可靠控制对其的依赖For the coarse attributes the authors study, MLLMs encode the visual evidence but cannot reliably control their reliance on it, indicating that for the coarse attributes they study, MLLMs cannot reliably control their reliance on it.

Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV
标题 -> 标题中文:全局计算,本地落盘:稀疏 Event-KV 的记忆契约
arXiv:2607.23693 Agent 智能体 方法 OA · 绿色 被引 1 · S2

结果是面向稀疏 event-KV 服务的记忆契约:写入什么、落在何处、源消失后什么得以保留The result is a memory contract for sparse event-KV serving: what to write, where it lands, and what survives once the source is gone.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
Video-DeepResearch:迈向下一代多模态深度研究 Agent
arXiv:2608.03979 多模态 方法 OA · 绿色 被引 2 · S2

提出 Video-DR,采用解耦的感知-探索流水线与分阶段工具解锁,强制在 web 检索前进行充分的跨帧视觉定位,实现突破模仿学习上限的自主探索。Video-DR is introduced, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval, enabling autonomous exploration that breaks the imitation-learning ceiling.

Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation
知识-几何解耦:面向流式推荐的可刷新预训练迁移
arXiv:2608.02738 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Knowledge-Geometry Decoupling (KGD) 并引入 Behavioral Multi-Token Prediction (BMTP),仅将协作或语义相关的未来项作为监督,从而得到更干净、更可迁移的行为知识。Knowledge-Geometry Decoupling (KGD) is proposed and Behavioral Multi-Token Prediction (BMTP) is introduced to retain only collaboratively or semantically related future items as supervision, yielding cleaner and more transferable behavioral knowledge.

Quo Vadis, World Modeling?
Quo Vadis, World Modeling?
arXiv:2608.02713 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

将 Agent-Centric Interactive World Proxies 概念化,将基础范式从物理状态转移转向 agent 可用的信息转移,如执行结果、检索到的经验或技能、以及验证信号,扩展了世界建模的范围,为持续改进的 agent 提供多样化反馈。This work conceptualize Agent-Centric Interactive World Proxies, shifting the fundamental paradigm from physical state transitions to agent-usable information transitions, such as execution outcomes, retrieved experiences or skills, and verification signals, broadening the scope of world modeling to provide versatile feedback for continually improving agents.

Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories
Push-Wiper:通过分段推送轨迹实现跨多种污渍与表面的通用机器人清洁
arXiv:2608.00730 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

高粘度污渍以其高粘度与复杂流变特性,仍是机器人表面清洁的主要挑战。传统擦拭往往扩散污渍,而擦洗摩擦力更强却存在损伤表面的风险。本文提出 Push-Wiper,一种将高粘度污渍清洁重构为聚集问题的框架。Push-Wiper 使用海绵通过分段推送轨迹渐进式聚集污渍,随后通过后处理阶段剥离已聚集物质并实现海绵自清洁。我们采用逐步Viscous stains, characterized by high viscosity and complex rheological properties, remain a major challenge for robotic surface cleaning. Conventional wiping often spreads the stain, while scrubbing provides stronger friction but risks damaging the surface. In this paper, we propose Push-Wiper, a framework that reformulates viscous stain cleaning as an aggregation problem. Push-Wiper employs a sponge to progressively gather stains through segmented pushing trajectories, followed by a post-processing phase that detaches the aggregated material and enables sponge self-cleaning. We adopt a stepw

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
ST-WAM:面向视觉分布偏移下鲁棒操作的语义-时序世界动作模型
arXiv:2607.28993 工程化 方法 OA · 绿色 被引 9 · S2

提出 Semantic-Temporal WAM (ST-WAM),使用 DINOv3 作为未来预测与历史检索的共享语义表示,同时保留细粒度 VAE 动力学,以提升动作鲁棒性;证明语义-时间建模能有效补充像素生成动力学,实现稳健的操作。Semantic-Temporal WAM (ST-WAM) is proposed to improve action robustness by using DINOv3 as a shared semantic representation for future prediction and history retrieval while retaining fine-grained VAE dynamics, demonstrating that semantic-temporal modeling effectively complements pixel-generative dynamics for robust manipulation.

ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
ContinualSkillBench:LLM Agent 能否真正进化其能力?
arXiv:2608.03874 Agent 智能体 方法 OA · 绿色 被引 4 · S2

提出 ContinualSkillBench,一个面向 in-context 持续 skill 学习的动态评估框架,表明当前 in-context skill 进化机制能够支持持续适应,但仍难以稳定地将经验整合为鲁棒且可迁移的 skill。ContinualSkillBench is introduced, a dynamic evaluation framework for in-context continual skill learning that shows that current in-context skill evolution mechanisms can support continual adaptation, but still struggle to consistently consolidate experience into robust and transferable skills.

PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs
PosterMELD:面向可控设计多样化的多 Agent 论文转海报生成,输出可编辑的可印刷成品
arXiv:2608.02218 Agent 智能体 方法 OA · 绿色 被引 3 · S2

PosterMELD 是一个模板条件的多 agent 流水线:capacity-aware slot 在渲染前引导写作,确定性 gate 与 VLM 审核将失败路由到有界修复,在生成的多种方法中获得最高的条件 CHE 并产出多个可印刷输出。PosterMELD is a template-conditioned multi-agent pipeline: capacity-aware slots guide writing before rendering, and deterministic gates plus vision-language model (VLM) review route failures to bounded repair result in the highest conditional CHE among generated methods with multiple print-ready outputs.

Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation
更优、更强、更快、更广:面向基于 MLLM 分割的结构化全 Mask 预测
arXiv:2608.02791 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

STAMPlus 解决了超出单目标预测的三难问题,解耦自回归对话与非自回归 mask 预测,取得 SOTA 分割性能,同时保持通用多模态指令遵循能力,并降低 12 类别延迟。STAMPlus resolves the trilemma beyond single-target prediction, decoupling autoregressive dialogue from non-autoregressive mask prediction and achieves state-of-the-art segmentation performance, preserves general multimodal instruction following, and reduces 12-category latency.

MiniWorld: Democratizing the Training of Video World Models from Scratch
MiniWorld:降低视频世界模型从零训练的门槛
arXiv:2608.01127 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 MiniWorld,一个从零训练流式视频世界模型的可复现框架,采用 chunk-wise 非递减噪声调度与两阶段继续训练,提升时间建模与稳定性,将促进未来视频世界模型的研究。MiniWorld is presented, a reproducible framework for training streaming video world models from scratch that adopts a chunk-wise non-decreasing noise schedule and two-stage continued training to improve temporal modeling and stability and will facilitate future research on video world modeling.

Decoding Children's Gait Behavior
解码儿童步态行为
arXiv:2608.00371 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

为人体动作识别引入新问题域:基于标准 RGB 视频对儿童步态行为进行细粒度分析,并描述一个统一的端到端框架用于解码儿科步态的基本组成。This work introduces a new problem domain for human action recognition: the fine-grained analysis of children's gait behaviors from standard RGB video, and describes a unified end-to-end framework for decoding fundamental components of pediatric gait.

BERTopic: Neural topic modeling with a class-based TF-IDF procedure
BERTopic:基于类内 TF-IDF 流程的神经主题建模
arXiv:2203.05794 工程化 方法 OA · 绿色 被引 3220 · S2

提出 BERTopic,一种通过开发类内 TF-IDF 变体来提取一致性主题表示,从而扩展主题建模流程的主题模型BERTopic is presented, a topic model that extends the process of topic modeling by extracting coherent topic representation through the development of a class-based variation of TF-IDF.

Stereo Matching by Training a Convolutional Neural Network to Compare Image Patches
通过训练卷积神经网络比较图像块进行立体匹配
arXiv:1510.05970 工程化 方法 OA · 绿色 被引 1478 · S2

提出一种从校正后的图像对中提取深度信息的方法,使用卷积神经网络在小图像块上学习相似性度量,并针对该任务考察了两种网络架构:一种面向速度优化,另一种面向精度优化This work presents a method for extracting depth information from a rectified image pair by learning a similarity measure on small image patches using a convolutional neural network and examines two network architectures for this task: one tuned for speed, the other for accuracy.

Random Erasing Data Augmentation
Random Erasing 数据增强
arXiv:1708.04896 工程化 方法 OA · 绿色 被引 4339 · S2

在训练过程中,Random Erasing 在图像中随机选择一个矩形区域并以随机值擦除其像素,在图像分类、目标检测与行人重识别任务中相较于强基线均带来稳定提升In training, Random Erasing randomly selects a rectangle region in an image and erases its pixels with random values and yields consistent improvement over strong baselines in image classification, object detection and person re-identification.