Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue

  • 类型:arxiv
  • 标识:2609.04250
  • 链接:https://arxiv.org/abs/2609.04250
  • 主分类:multimodal
  • 形态:method
  • TLDR:An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities live in separate model families: spoken dialogue models produce speech without motion, and co-speech motion models produce motion only from audio handed to them. The standard remedy is a cascade that first generates the spoken response and then runs a motion model over the finished audio, which requires a second full inference pass and precludes any joint optimisation between the two. We present Motion-Omni, an end-to-end framework in which a spoken dialogue model natively outputs exp
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-07-agent-rag-longcontext-candidates.json