← All posts / Models

One Brain for the Whole Car: Alibaba Open-Sources Qwen-Drive-1.0, a 4B VLM That Sees, Explains, and Drives

Alibaba's Qwen team has open-sourced Qwen-Drive-1.0-4B, a vision-language foundation model that unifies 3D perception, driving Q&A, and trajectory planning in one Apache 2.0 package — while keeping the base Qwen3.5-4B entirely untouched.

One Brain for the Whole Car: Alibaba Open-Sources Qwen-Drive-1.0, a 4B VLM That Sees, Explains, and Drives

Autonomous driving stacks have historically been split-brained: a perception network that sees the world in tensors, a planner that outputs trajectories nobody can read, and — increasingly — a separate chatbot in the cockpit that can talk about the drive but has no idea what the car is about to do. On September 7, Alibaba’s Qwen team, working with Huazhong University of Science and Technology, released Qwen-Drive-1.0-4B under Apache 2.0, an attempt to collapse that stack into a single vision-language foundation model that performs 3D perception, answers free-form questions about the scene, and generates the actual driving trajectory — all while leaving its pretrained VLM architecture entirely untouched.

What exactly shipped

Qwen-Drive-1.0 starts from the natively multimodal Qwen3.5-4B as a shared backbone and bolts on two external modules:

  • A BEV perception head (0.5 GB) that jointly performs 3D object detection, semantic occupancy prediction, and bird’s-eye-view map segmentation. The team deliberately kept this head simple so it acts as an explicit, inspectable probe of what 3D information is actually latent in the VLM’s representations, rather than a benchmark-chasing specialist.
  • A Planning Expert that conditions on the shared VLM representations and generates future ego trajectories through flow matching. Two variants ship: planner-sft (2.1 GB, imitation-trained, supports both direct and reasoning-conditioned planning) and planner-rl (2.1 GB, further reward-optimized on NAVSIM PDMS, Waymo Open Dataset E2E RFS, and a displacement term).

The whole repository is one directory: the 9.1 GB VLM at the root, with each task head in a subfolder beside it. Load the root alone and you have a driving-scene chatbot; attach a planner and run(InferenceMode.REASONING_PLANNING, ...) returns sampled trajectories shaped (6, 50, 3) — six candidate paths, five seconds at 10 Hz, each with x, y, and heading. A demo script runs four bundled planning scenes and six perception frames end to end with no extra data. The technical report is on arXiv (2609.00111, submitted August 31 by Xin Zhou and fifteen co-authors including Xiang Bai’s group), and the code lives at QwenLM/Qwen-Drive-1.0 on GitHub.

The numbers

The headline claim — that driving competence can be added without lobotomizing the general model — holds up in the published tables.

Motion planning. On the open-loop WOD-E2E benchmark, the RL planner posts an RFS of 8.45 val / 7.91 test, beating MindVLA-U1’s 8.20/7.87, and a 5-second ADE of 1.27 on val — nearly halving the next-best 2.28. On NAVSIM (pseudo-closed-loop), it reaches 90.7 PDMS against 90.3 for SpanVLA and SimWAM, and 91.4 in best-of-6 sampling. The one benchmark it doesn’t win is AlpaSim closed-loop, where Alpamayo-1.5’s 0.45 at-fault score beats Qwen-Drive’s 0.37 — though Qwen-Drive’s RL variant improves its own SFT score of 0.27 by 37% there. The team’s framing is honest: after reinforcement learning, the model “trades only a marginal open-loop displacement for comprehensive gains in human-preference alignment and closed-loop safety.”

Driving VQA. The results are more lopsided. On LingoQA it scores 77.8 against 72.0 for MiMo-Embodied-7B; on WaymoQA safety, 70.7 versus 66.5; on DRAMA causal-reasoning benchmark CoC, it posts 41.3 while every competitor — including dedicated driving VLMs from NVIDIA’s Cosmos line — sits at 4.0 or below. Ego3D depth RMSE drops to 7.78 against 9.85 for the next best.

General capability. This is the part that usually breaks. MMStar actually rises slightly to 75.9 (base: 75.3), MMMU dips only from 73.4 to 72.7, and MMBench holds at 85.5 versus 87.1. A staged training recipe — unifying trajectory annotations across public datasets, rewriting responses, filtering for consistency, and interleaving general vision-language data throughout — appears to have averted catastrophic forgetting, which the team calls out as a design goal rather than a happy accident.

The catch, and why it matters

The Decoder’s coverage adds a caveat worth taking seriously: in the paper’s own analysis, the model’s explanation of a maneuver doesn’t always match the maneuver it actually executes. The finding cuts both ways. It confirms the researchers’ core thesis — that text-image models do not automatically understand three-dimensional space, and spatial awareness has to be trained on purpose (hence the BEV probe). But it also means the natural-language rationale a passenger might hear is not yet a faithful account of the physics actually commanding the wheel. For an industry increasingly interested in explainable autonomy — regulators included — that gap is the difference between a chat feature and an audit log.

Why this release is significant

First, the economics: a 4B-parameter model with heads totaling under 12 GB is deployable on automotive-grade hardware, not just an H100 cluster. Second, the provenance: this is trained purely on public datasets (nuScenes, Waymo, NAVSIM and kin) with unified trajectory formats, meaning any lab can reproduce, fine-tune, or extend it — a genuine foundation-model play rather than a demo. Third, the strategic signal: Alibaba is positioning Qwen as the default open-weight stack for embodied AI the way it did for text, and Chinese teams now openly target the cockpit-and-drivetrain unification that Western incumbents have only demoed behind closed doors.

It is telling that the team titles this “an initial step.” Qwen-Drive-1.0 doesn’t beat every specialist on every benchmark, and its explanations still occasionally diverge from its actions. But as a single, open, inspectable artifact that unifies seeing, explaining, and planning — it’s the clearest blueprint yet for where driving models go next.