Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

  • 类型:arxiv
  • 标识:2609.00111
  • 链接:https://arxiv.org/abs/2609.00111
  • 主分类:multimodal
  • 形态:method
  • TLDR:We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-eye-view (BEV) perception head jointly performs 3D object detection, semantic occupancy prediction, and BEV map segmentation. It serves as a probe of the 3D information accessible from the shared representations and provides an explicit, inspectable interface to 3D scene structur
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-02-agent-rag-longcontext-candidates.json