SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

  • 类型:arxiv
  • 标识:2608.05137
  • 链接:https://arxiv.org/abs/2608.05137
  • 主分类:multimodal
  • 形态:method
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:SmartMage is proposed, a unified MLLM that dynamically orchestrates heterogeneous modalities for semantic-aware 3D scene understanding and achieves state-of-the-art performance across five 3D scene understanding benchmarks, and attains competitive results on RGB-only video understanding benchmarks.
  • 待LLM分类:否
  • 标题中文:SmartMage:面向 3D 场景理解的动态模态编排
  • TLDR中文:SmartMage 是一个统一的 MLLM,动态调度异构模态以实现语义感知的 3D 场景理解,在五个 3D 场景理解基准上达到 SOTA,并在仅 RGB 视频理解基准上取得具有竞争力的结果。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-07-agent-rag-longcontext-candidates.json
  • [S2 enrich]