SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding
- 类型:arxiv
- 标识:2608.05137
- 链接:https://arxiv.org/abs/2608.05137
- 主分类:multimodal
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:SmartMage is proposed, a unified MLLM that dynamically orchestrates heterogeneous modalities for semantic-aware 3D scene understanding and achieves state-of-the-art performance across five 3D scene understanding benchmarks, and attains competitive results on RGB-only video understanding benchmarks.
- 待LLM分类:否
- 标题中文:SmartMage:面向 3D 场景理解的动态模态编排
- TLDR中文:SmartMage 是一个统一的 MLLM,动态调度异构模态以实现语义感知的 3D 场景理解,在五个 3D 场景理解基准上达到 SOTA,并在仅 RGB 视频理解基准上取得具有竞争力的结果。
- 来源文件:
- /inbox/tom/_candidates/2026-08-07-agent-rag-longcontext-candidates.json
- [S2 enrich]