One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
- 类型:arxiv
- 标识:2610.12448
- 链接:https://arxiv.org/abs/2610.12448
- 主分类:multimodal
- 形态:method
- TLDR:In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank. A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space. We evaluate this design in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher. Across b
- 待LLM分类:否
- 标题中文:One Block, Multiple Depths:具有深度编程专家的循环视觉 Transformer
- TLDR中文:在本工作中,我们证明单个 Transformer block 循环应用即可在相近推理 FLOPs 下匹配全深度视觉编码器的精度,且无需中间特征蒸馏。reViT 通过将每个循环深度的 FFN 表示为小型共享专家库的凸组合来恢复深度特定的变换。一个连续归一化的深度坐标对该混合进行编程,在 FFN 参数空间中定义一条可重采样的轨迹。我们在两种场景下评估该设计:有监督 ImageNet-1k 训练以及从 DINOv2 教师模型蒸馏。在各规模下
- 来源文件:
- /inbox/tom/_candidates/2026-10-09-agent-rag-longcontext-candidates.json