研究库 论文知识库
Papers · organized/paper_cards

论文

1094 张论文卡片 · 方法 · OA 绿色

开放获取 全部 绿色 · 1640
Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation
我们能信任教师吗?面向自蒸馏的解耦信用方向-幅度
arXiv:2609.34848 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出解耦信用自蒸馏,理论上将信用方向与幅度解耦为两个可靠信号,并据此校准特权教师监督,从而完成策略优化的 step-to-token 信用分配。Decoupled Credit Self-Distillation is introduced, which theoretically decouples credit direction and magnitude into two reliable signals and uses them to calibrate privileged teacher supervision, enabling step-to-token credit assignment for policy optimization.

Selecting Diverse SFT Traces Improves Post-RL Generalization
选择多样化的 SFT 轨迹可提升 RL 后的泛化能力
arXiv:2609.33780 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,推理路径多样性可作为筛选 SFT 数据的实用准则,能更好地为 RL 准备模型,并据此提出一种轻量级、基于规则的指纹方法用于筛选。These results identify reasoning-route diversity as a practical criterion for selecting SFT data that better prepares models for RL, and propose a lightweight, rule-based fingerprint to select for it.

NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech
NVAlign:面向连续自回归流匹配文本到语音系统中非言语控制的直接梯度优化
arXiv:2609.31892 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

NVV-SuperBench 与人工听力评测的结果显示,NVAlign 在标签跟随准确率上优于 SFT 与 Flow-GRPO 基线,证明直接对奖励梯度进行优化可提升连续自回归流匹配 TTS 中的非语言控制能力。Results from NVV-SuperBench and human listening evaluations show that NVAlign improves tag-following accuracy over SFT and Flow-GRPO baselines, demonstrating that direct reward-gradient optimization can improve non-verbal control in continuous autoregressive flow-matching TTS.

Towards Semi-Automatically Comparing Keyword-Based and Semantic Search Accuracy
Towards Semi-Automatically Comparing Keyword-Based and Semantic Search Accuracy
arXiv:2609.37749 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个新颖的初步框架,可定量评估输出格式不同(如列表与消息)的检索系统的 IR 准确度,为客观评估基于关键词与基于语义的对话式检索方法奠定坚实基础。This work introduces a novel, preliminary framework to quantitatively assess IR accuracy of search systems that produce different output formats, such as lists and messages, and provides a strong foundation for objectively assessing keyword-based and semantic chat-based search methods.

Context Language Models
Context Language Models
arXiv:2609.37725 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Context Language Models,能原生管理自身上下文,并可通过经标准 skill 优化循环演化的自然语言指令进行引导,在上下文管理任务上将未见数据的准确率最高提升 35.9 分,同时降低计算开销。Context Language Models are introduced, language models that natively manage their own context and can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute.

Multimodal 补充候选
arXiv:2606.13578 多模态 方法 OA · 绿色 被引 4 · S2

构建了 RoboGenesis,一个基于仿真的工作流与数据引擎,可从原子技能组合配置好的实验工作流,对 rollout 进行验证与过滤,并跨支持的机器人配置导出结构化演示数据。RoboGenesis is built, a simulation-based workflow and data engine that composes configured laboratory workflows from atomic skills, validates and filters rollouts, and exports structured demonstrations across supported robot profiles.

Follow the Entities: A Corpus Map for Agentic Search
[标题中文] 跟随实体:面向 Agentic 搜索的语料库地图
arXiv:2609.37226 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 CorpusMap,一个以语料中重复出现的实体为核心的导航层;这些实体可从文档自身识别,并能在不同来源间把单个文档与众多其他文档相连,表明实体可作为大型文档集合导航的有效锚点。CorpusMap is introduced, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources, suggesting that entities serve as effective anchors for navigating large document collections.

Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
[标题中文] 相同字节,不同权威:Chat 模板 Prompt 注入中的保留 token 表示
arXiv:2609.35932 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

在所测试的每一对基座与指令微调模型中,指令微调都强化了模型对保留标记的偏好,且该差距在该通道上持续存在。In every base and instruction-tuned pair the authors test, instruction tuning strengthens the model's preference for reserved markers, and the gap persists on that channel.

Pretraining Transformers with Quantized Softmax in Attention
[标题中文] 使用量化 Softmax 的注意力机制预训练 Transformer
arXiv:2609.33591 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文推导了对应的反向传播规则(包括校准项的导数),并在模型、数据、优化器均一致的预训练实验中对不同选择进行了比较。This work derives the corresponding backward rules, including calibration derivatives, and compares these choices in pretraining experiments matched on model, data, and optimizer, and compares these choices in pretraining experiments matched on model, data, and optimizer.

Omni-IO Skills: Harnessing Your Agent Omni-Native
[标题中文] Omni-IO Skills:让你的 Agent 原生支持全模态
arXiv:2609.31847 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,harness 层面的能力组合是构建广泛、可演进 Omni 系统的实用路径,且无需改变宿主 agent 的推理核心。The results establish harness-level capability composition as a practical route to broad, evolvable Omni systems without changing the host agent's reasoning core.

Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents
[标题中文] Omni-Decision:面向全模态 Agent 的证据账本式规划
arXiv:2607.11433 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Omni-Decision,一个基于证据账本规划的 omni-modal agent:用显式的证据账本替代不断膨胀的对话历史,记录仍缺失的证据、已确认的内容以及记录间的冲突。Omni-Decision is presented, an omni-modal agent built on evidence-ledger planning: it replaces the growing dialogue history with an explicit evidence ledger that records what evidence is still missing, what has been confirmed, and where records conflict.

KUPAS MASTER: Distilling the Tacit Expertise of Master Practitioners into Agent-Ready Experience Corpora
[标题中文] KUPAS MASTER:将资深从业者的隐性专业知识蒸馏为 Agent 可用的经验语料
arXiv:2609.37673 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

KUPAS MASTER 是一个围绕九层认知语料构建的经验工程平台,将异构的工作记录与从业者访谈转化为 agent 可用的可追溯、可复用经验语料,提供了从个体隐性经验到组织知识与 agent 能力的可行路径。KUPAS MASTER, an experience engineering platform built around nine-layer cognitive corpus construction, turns heterogeneous work records and practitioner interviews into traceable, reusable experience corpora for agents, and provides a practical path from individual tacit experience to organizational knowledge and agent capabilities.

Improved Distributional Diffusion Models
[标题中文] 改进的分布式扩散模型
arXiv:2609.37147 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

一种随机少步生成器,其 FID 在采样预算从 4 增加到 50 NFE 时不会下降,且相同配方可迁移到 text-to-image 生成。A stochastic few-step generator whose FID does not degrade as the sampling budget grows from 4 to 50 NFE, and the same recipe transfers to text-to-image generation.

LLMs are General Asynchronous Agents
[标题中文] LLM 是通用异步 Agent
arXiv:2609.35427 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种异步 LLM 框架,允许用户(或 Agent 自身)定义具有重叠 memory 状态的推理协程,并展示 Qwen 3.x 模型无需任务专属训练即可在流式视频理解、电子游戏和监控中实现异步运行。This work develops an asynchronous LLM framework that lets users (or the agents themselves) define inference coroutines with overlapping memory states and showcases that Qwen 3.x models are capable of asynchronous operation for streaming video understanding, videogames, and monitoring, without task-specific training.

AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation
AutoRef:面向 Agentic 多参考图生成的 Harness 优化
arXiv:2609.35530 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 AutoRef,在保持两侧模型冻结的前提下自动优化 harness:一个 coding Agent 迭代重写 harness 代码,并发现 AutoRef-Harness,可改进开源权重模型 FLUX。This work proposes AutoRef, which optimizes the harness automatically while keeping both models frozen: a coding agent iteratively rewrites the harness code, and discovers AutoRef-Harness, which improves the open-weight FLUX.

Reasoning with Image Generation
结合图像生成的推理
arXiv:2609.16409 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

在包括多视角空间推理和碰撞预测在内的六项多样化视觉推理任务中,ReImaGin 始终优于纯文本推理和专用 vision-tool 基线,提升幅度高达 25%,证明了灵活、可生成的视觉推理的优势。Across six diverse visual reasoning tasks including multi-view spatial reasoning and collision prediction, ReImaGin consistently outperforms both text-only reasoning and specialist vision-tool baselines, with gains of up to 25\%, demonstrating the advantage of flexible, generative visual reasoning.

What Makes Recurrence Effective in Looped Language Models?
循环语言模型中的循环机制为何有效?
arXiv:2609.36636 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

通道级 history-state 注入结合 timestep 条件化构成了一种低成本且更有效的设计,能够在更长 unrolling 下更好地保留知识,同时提升在不同推理预算下的鲁棒性。It is shown that channel-wise history-state injection combined with timestep conditioning offers a low-cost and more effective design, better preserving knowledge under extended unrolling while improving robustness across inference budgets.

Fractional State Space Transition for Long Sequence Modeling
用于长序列建模的分数阶状态空间转移
arXiv:2609.36314 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 FRAC,一种源自分数阶动力学的选择性 SSM 架构,用幂律长记忆替代指数衰减,在长上下文性能上持续优于 SOTA SSM 基线,同时在短上下文上保持竞争力。FRAC is introduced, a selective SSM architecture derived from fractional dynamics that replaces this exponential decay with power-law long memory and consistently improves long-context performance over state-of-the-art SSM baselines while staying competitive on short-context.

Scaling Properties of Same-Family On-Policy Distillation
同家族同策略蒸馏的缩放特性
arXiv:2609.32722 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

研究发现,早期 OPD 训练动态均呈现一种规律的 *useful-transfer* 区间,其中留出准确率(即 *gold score*, $G$)随 $d=\sqrt{\mathrm{KL}(\pi_\theta \Vert \pi_{\mathrm{ref}})}$(学生初始化在 token 级反向 KL 散度的平方根)近似线性上升。It is found that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, $G$) rises approximately linearly in $d=\sqrt{\mathrm{KL}(\pi_\theta \Vert \pi_{\mathrm{ref}})}$, the square root of token-level reverse KL divergence from the student initialization.

4.2 Reliability 不等于成功率:12 指标拆出 consistency / robustness / predictability / safety
4.2 Reliability 不等于成功率:12 指标拆出 consistency / robustness / predictability / safety(⭐⭐⭐⭐⭐)
arXiv:2602.16666 安全与风险 方法 Open MIND OA · 绿色 被引 78 · S2

本工作提出 12 个具体指标,从一致性、鲁棒性、可预测性和安全性四个关键维度分解 Agent 可靠性,可与传统评估互补,并提供用于分析 Agent 表现、退化与失败方式的工具。This work proposes twelve concrete metrics that decompose agent reliability along four key dimensions: consistency, robustness, predictability, and safety, which complement traditional evaluations while offering tools for reasoning about how agents perform, degrade, and fail.

UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement
UniEvo-VL:一种用于多模态模型自我改进的 On-policy 自蒸馏训练方案
arXiv:2609.38721 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 UniEvo-VL,一种面向多模态模型的自演化框架,可在 test-time compute 阶段从这种建设性的自纠错反馈中学习,在无外部监督或指导的情况下提升用户使用多模态模型的体验。UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback during test-time compute, is introduced to enhance the user experience when using multimodal models without external supervision or guidance.

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
Mid-Harness:在 Model 与 Harness 之间扩展终端 Agent 的动作规模
arXiv:2609.39982 评测基准 方法 OA · 绿色 被引 1 · S2

本文提出 Mid-Harness,在执行前对候选动作进行采样和验证,同时保持 generator 和 harness 不变,并将 action scaling 识别为终端 Agent 中 test-time compute scaling 的一个有前景的目标。Mid-Harness is introduced, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged, and identifies action scaling as a promising target for test-time compute scaling in terminal agents.

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
False Frontiers:诊断与缓解自演化搜索 Agent 中的共作弊问题
arXiv:2609.39102 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出多样本验证(MSV),对同一模型在有源信息和无源信息下各查询三次,以决定任务接纳并替换不可靠的伪标签,这部分减少了虚假一致性,但仍残留大量 co-cheating。This work introduces multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels, which partially reduces false agreement but leaves substantial residual co-cheating.

Scaling Laws for Looped Mixture of Experts
循环化 Mixture of Experts 的扩展定律
arXiv:2609.40316 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Loop Scaling Laws,是首个将 recurrence 和 sparsity 与模型规模、数据联合建模的 scaling law,为在算力和显存约束下设计 looped MoE 模型提供了原则性基础。Loop Scaling Laws are introduced, the first scaling law to jointly model recurrence and sparsity alongside model size and data, and provide a principled foundation for designing looped MoE models under compute and memory constraints.

Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation
打破巴别塔:面向长字幕翻译的自进化多 Agent 系统
arXiv:2609.38660 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

SMART,面向长篇字幕翻译的自演化多 Agent 系统,构建持久化的剧集级 memory,并通过动态 router 和 Mixture-of-Agents 层翻译部分句子,配合术语验证、字幕约束校验和上下文检索工具。SMART, a Self-evolving Multi-Agent system for long-foRm subtitle Translation, builds persistent series-level memory and translates a subset of sentences through a dynamic router and Mixture-of-Agents layer with tools for terminology verification, subtitle constraint validation, and contextual retrieval.

RoPE at the End of Its Rope? Theory, Diagnosis, and Mitigation of Long-Context Failures
RoPE 走到尽头了吗?长上下文失效的理论、诊断与缓解
arXiv:2609.39929 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

基于 RoPE 的语言模型出现长上下文失效,根源在于 RoPE 在维持稳定 token 偏好与区分相近位置之间存在固有权衡。要判断应处理哪种弱点以及如何处理,需要更精确地刻画 RoPE 在不同上下文长度下训练后模型中的行为。我们通过允许 RoPE 各频率下 query-key 尺度不一致,弥补了先前理论的一个关键局限,使之与实际经验观测高度吻合。我们的理论使得上述脆弱性对单个注意力头与输入可测量,并量化了高频Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and distinguishing nearby positions. Determining which weakness to address, and how, requires a more precise characterization of RoPE's behavior in trained models across context lengths. We address a key limitation of prior theory by allowing unequal query-key scales across RoPE frequencies, which aligns well with practical empirical observations. Our theory makes both vulnerabilities measurable for individual heads and inputs, and quantifies how high-freque

4.1 LogicalRAG:把 Agentic RAG 的重点从“更重 backend”转向“更强 retrieval control”
arXiv:2605.27123 RAG 检索增强 方法 OA · 绿色 被引 3 · S2

本文提出一个 Agentic RAG 框架,使 LLM 能够使用逻辑表达式构建检索意图,同时将检索后端简化为基于倒排索引的系统,并表明将检索过程锚定在逻辑查询上可显著降低生成响应中的幻觉。This paper proposes an agentic RAG framework that enables LLMs to formulate retrieval intents using logical expressions while simplifying the retrieval backend to an inverted-index-based system, and shows that anchoring the retrieval process in logical queries substantially reduces hallucinations in generated responses.

Safety of Latent Communication in Multi-Agent Systems
多 Agent 系统中潜在通信的安全
arXiv:2609.39788 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

即使是良性的链接训练也会在底层安全对齐 agent 不变的情况下,相对于基于文本的通信增加有害合规性;安全对齐需要将多智能体系统作为整体来考虑。This work shows that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged, and shows that safety alignment requires considering the multi-agent system as a whole.

ATLAS: Aligned Transport of Latent Structure for Reliable World Model Planning
ATLAS:面向可靠世界模型规划的对齐潜在结构传输
arXiv:2609.36333 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ATLAS(Aligned Transport of Latent Structure),一种在显式保持关系几何结构的同时校准全局潜空间分布的训练目标,通过一维 Wasserstein-2 传输进行 Wasserstein 嵌入匹配来校准其边缘分布。This work introduces Aligned Transport of Latent Structure (ATLAS), a training objective that explicitly preserves relational geometry while calibrating the global latent distribution and uses Wasserstein embedding matching to calibrate its marginal through one-dimensional Wasserstein-2 transport.

It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them
关键不在图像本身:无关上下文会扰动 VLM 评判且不提供有效信息
arXiv:2609.37863 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MIST(Misleading-Image Stress Test):200 个英文句子,每句围绕一个可作比喻或字面理解的短语,配以对齐图像(描绘其读法)、误导图像(描绘相反读法)或无图像三种条件。MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all is introduced.

Generative Moment Matching Networks
生成矩匹配网络
arXiv:1502.02761 多模态 方法 OA · 绿色 被引 954 · S2

本文提出一种方法,通过多层感知机的一次前馈传播生成独立样本(与近期提出的 GAN 类似),并使用 MMD 学习生成可被解码为样本的 codes。This work forms a method that generates an independent sample via a single feedforward pass through a multilayer perceptron, as in the recently proposed generative adversarial networks, using MMD to learn to generate codes that can then be decoded to produce samples.

DAGent: Evaluate-then-Grow Planning for Deep Research Agents
DAGent:面向深度研究 Agent 的 Evaluate-then-Grow 规划
arXiv:2609.39154 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

DAGent 是一个基于 DAG 的多智能体框架,采用先评估再生长的增量规划:Orchestrator 逐批扩展任务图,每一步扩展都以已完成节点的置信度与不确定性信号为条件;证据条件化规划在更低的每任务 token、工具调用和步骤开销下达到了比 Plan-then-Patch 更高的准确率。DAGent, a DAG-based multi-agent framework with Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph one batch at a time, conditioning each expansion on confidence and uncertainty signals from completed nodes, which shows that evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart.

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
MILO:通过编排式多 Agent 进化实现自动化 Harness 发现
arXiv:2609.38349 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MILO(Meta-evolutionary Island Orchestration),一个共同演化 agent harness 及其发现策略的框架,使用前沿模型(Opus 4.8)与开源权重模型(gpt-oss-120b)超越了八个 SOTA harness 与六种搜索方法。This work introduces MILO (Meta-evolutionary Island Orchestration), a framework that co-evolves agent harnesses and the strategy used to discover them and outperforms eight state-of-the-art harnesses and six search methods using frontier (Opus 4.8) and open-weight (gpt-oss-120b) models.

2.3 本轮补充公开检索
arXiv:2606.14589 Agent 智能体 方法 OA · 绿色 被引 3 · S2

本文对一个自 2026 年 3 月起持续运行的个人助理 Agent 运行时中的静默失败进行纵向研究,该系统包含约 40 个定时任务、8 个 LLM 提供商、一个工具治理代理以及一个知识库记忆层,由 4,286 个单元测试和 827 项治理检查守护。A longitudinal study of silent failures in a personal-assistant agent runtime in continuous production since March 2026, with roughly 40 scheduled jobs, 8 LLM providers, a tool-governance proxy, and a knowledge-base memory plane, defended by 4,286 unit tests and 827 governance checks is presented.

Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Tacit-TTS:从自回归解码到掩码预测的高效免转录语音克隆
arXiv:2609.38658 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Tacit-TTS 是一个从 IndexTTS2 蒸馏而来的高效无需转录的零样本语音克隆系统,用掩码非自回归生成替换自回归的文本到语义解码,引入无需训练的声学长度估计,并通过 ReFlow 蒸馏加速流匹配渲染器。Tacit-TTS is presented, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2 that replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation.

PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation
PixelUMM:无编码器的统一图像与视频理解与生成
arXiv:2609.38597 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

统一多模态模型(UMM)通常依赖独立的视觉表示分别完成理解与生成,这增加了视觉上下文长度,并使其难以与既有视觉-语言预训练流程集成。近期 pixel-space modeling 的进展提供了一种无编码器的替代方案,但将该范式从图像扩展到视频并非易事:视频理解与生成采用不同的时间表示,统一视觉接口的设计仍是开放问题。本文提出 PixelUMM,一种用于统一图像与视频理解的无编码器模型……Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending this paradigm from images to videos is non-trivial: video understanding and generation adopt different temporal representations, leaving the design of a unified visual interface an open question. We present PixelUMM, an encoder-free model for unified image and video understanding