Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

  • 类型:arxiv
  • 标识:2609.13053
  • 链接:https://arxiv.org/abs/2609.13053
  • 主分类:multimodal
  • 形态:method
  • TLDR:Visual goal and dynamics prediction can provide language-conditioned robot policies with both a target outcome and a representation of action-dependent scene changes. We bring these predictions into action generation and selection through a shared trajectory model. Dynin-Robotics implements this formulation on Dynin-Omni, an omnimodal masked-diffusion backbone, representing language, visual observations, goals, and actions as discrete tokens. By varying conditioning and target spans, the same model learns action prediction, action-conditioned next-observation prediction, terminal goal-state pr
  • 待LLM分类:否
  • 标题中文:Dynin-Robotics:全模态统一扩散视觉-语言-动作模型
  • TLDR中文:视觉目标与动力学预测可为语言条件机器人策略同时提供目标结果和动作相关的场景变化表征。我们通过共享轨迹模型将这些预测引入动作生成与选择。Dynin-Robotics 在 Dynin-Omni(全模态掩码扩散骨干)上实现该公式化方法,将语言、视觉观测、目标和动作表示为离散 token。通过改变条件跨度与目标跨度,同一模型学习动作预测、动作条件下一观测预测、终端目标状态 pr
  • 来源文件
  • /inbox/tom/_candidates/2026-09-15-agent-rag-longcontext-candidates.json