Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
- 类型:arxiv
- 标识:2609.00111
- 链接:https://arxiv.org/abs/2609.00111
- 主分类:multimodal
- 形态:method
- TLDR:We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-eye-view (BEV) perception head jointly performs 3D object detection, semantic occupancy prediction, and BEV map segmentation. It serves as a probe of the 3D information accessible from the shared representations and provides an explicit, inspectable interface to 3D scene structur
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-02-agent-rag-longcontext-candidates.json