Qwen-Drive-1.0-4B: Generate Driving Trajectories with a 4B VLM
Qwen's driving model adds a BEV perception head and a flow-matching planner to Qwen3.5-4B. Install it and sample ego trajectories from demo scenes.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
Qwen-Drive-1.0-4B is a vision-language model for autonomous driving from the Qwen team. It is built on Qwen3.5-4B. What sets it apart is that the team left the base VLM’s architecture completely unchanged and attached two external modules to it. The first is a BEV perception head that does 3D detection, semantic occupancy and BEV map segmentation. The second is a Planning Expert that generates future ego trajectories with flow matching. The VLM by itself still answers free-form questions about driving scenes. One download gives you VQA, 3D perception and motion planning, all sharing the same backbone. All of it is released under Apache-2.0.
Key specs
| Base model | Qwen/Qwen3.5-4B (architecture unchanged) |
| VLM weights (repo root) | 9.1 GB, serves VQA by itself |
planner-sft/ |
2.1 GB, imitation-trained, supports direct and reasoning planning |
planner-rl/ |
2.1 GB, reward-optimized on NAVSIM PDMS, WOD-E2E RFS and a displacement term |
perception/ |
0.5 GB BEV head |
| Trajectory output | (num_samples, 50, 3): x, y, heading, 5 s at 10 Hz |
Selected numbers from the card:
- NAVSIM PDMS: 90.7 for the RL model, against 90.3 for SpanVLA and SimWAM and 89.6 for AutoVLA. The card also reports a best-of-6 score of 91.4, which keeps the best of six samples. None of the compared methods report a best-of-6 number, so don’t compare 91.4 with their scores.
- WOD-E2E RFS (val/test): 8.45 / 7.91 for the RL model, against 8.20 / 7.87 for MindVLA-U1.
- AlpaSim closed-loop at-fault score: 0.37 for the RL model. Alpamayo-1.5 leads at 0.45.
- PAI-AV ADE 3s: 0.37 for SFT and 0.42 for RL. Alpamayo-1.5 is best at 0.35.
- Driving VQA (SFT): LingoQA 77.8 (base Qwen3.5-4B: 70.4), WaymoQA all 74.5 (base: 67.1), CoC all 41.3 (base: 2.6).
- General VL (SFT vs. base): The card describes the SFT model as “on par” with the base model and says it “largely preserves” general ability. SFT is a little behind the base on most of these benchmarks: MMBench 85.5 vs. 87.1, MMMU 72.7 vs. 73.4, MMMU-Pro std 62.7 vs. 64.9, CountQA 31.7 vs. 35.9. It is ahead on a few, such as MMStar 75.9 vs. 75.3 and RealWorldQA 79.0 vs. 76.3.
Install
The card’s planning code uses the inference package from the GitHub repo:
git clone https://github.com/QwenLM/Qwen-Drive-1.0 qwen-drive && cd qwen-drivepip install -e . --no-build-isolationDownload everything: the VLM at the root plus all task heads in their subfolders.
hf download Qwen/Qwen-Drive-1.0-4B --local-dir Qwen-Drive-1.0-4BThe repo includes demo data under data/demo/. The examples below use it. The card also mentions scripts/demo.py in the GitHub repo, which runs four bundled planning scenes and six perception frames end to end.
Run it
This snippet comes from the card. It loads the VLM with the RL planner attached, reads one demo scene and samples six trajectories in reasoning mode:
import torchfrom qwen_drive import InferenceMode, QwenDriveForPlanningfrom qwen_drive.benchmarks import read_scene_filefrom qwen_drive.images import ImageArchive
model = QwenDriveForPlanning.from_pretrained( "Qwen-Drive-1.0-4B", planner="Qwen-Drive-1.0-4B/planner-rl", dtype=torch.bfloat16, attn_implementation="flash_attention_2",).to("cuda").eval()
scene = next( read_scene_file( "data/demo/planning_scenes.jsonl", image_archive=ImageArchive.open("data/demo/frames.parquet"), )).scene
result = model.run(InferenceMode.REASONING_PLANNING, scene=scene, num_samples=6)print(result.reasoning)print(result.trajectories.shape) # (6, 50, 3) -> (x, y, heading), 5 s at 10 Hzresult.reasoning holds the model’s text rationale. result.trajectories holds the sampled paths.
Batch-plan a scene file and save trajectories
The card calls next() on read_scene_file, so it returns an iterator. That means you should be able to loop over every scene in a file and write the results out for offline scoring or visualization. The card only shows reading the first scene, so looping over the whole file is our assumption. The card also doesn’t give the exact types of the result fields. Wrapping trajectories in torch.as_tensor and reasoning in str() keeps the output serializable either way.
import json
import torchfrom qwen_drive import InferenceMode, QwenDriveForPlanningfrom qwen_drive.benchmarks import read_scene_filefrom qwen_drive.images import ImageArchive
model = QwenDriveForPlanning.from_pretrained( "Qwen-Drive-1.0-4B", planner="Qwen-Drive-1.0-4B/planner-rl", dtype=torch.bfloat16, attn_implementation="flash_attention_2",).to("cuda").eval()
archive = ImageArchive.open("data/demo/frames.parquet")scenes = read_scene_file("data/demo/planning_scenes.jsonl", image_archive=archive)
with open("trajectories.jsonl", "w") as out: for i, item in enumerate(scenes): result = model.run( InferenceMode.REASONING_PLANNING, scene=item.scene, num_samples=6 ) traj = torch.as_tensor(result.trajectories).float().cpu() out.write(json.dumps({ "scene_index": i, "reasoning": str(result.reasoning), "trajectories": traj.tolist(), # [6][50][x, y, heading] }) + "\n") print(f"scene {i}: {tuple(traj.shape)}")Look at the spread between samples
With num_samples=6 you get six trajectories for each scene. Measuring how far apart their endpoints land can be useful when you inspect results. The card doesn’t say how much the samples differ, and it doesn’t say whether the spread tracks scene difficulty. If you want to use spread to flag scenes, validate that on your own data first. The card also gives no units or coordinate frame for x and y. The threshold below uses the same units as the model output and is only a placeholder.
import torchfrom qwen_drive import InferenceMode, QwenDriveForPlanningfrom qwen_drive.benchmarks import read_scene_filefrom qwen_drive.images import ImageArchive
# Placeholder, in the model's output units (the card doesn't state them).SPREAD_THRESHOLD = 2.0
model = QwenDriveForPlanning.from_pretrained( "Qwen-Drive-1.0-4B", planner="Qwen-Drive-1.0-4B/planner-rl", dtype=torch.bfloat16, attn_implementation="flash_attention_2",).to("cuda").eval()
archive = ImageArchive.open("data/demo/frames.parquet")for i, item in enumerate( read_scene_file("data/demo/planning_scenes.jsonl", image_archive=archive)): result = model.run( InferenceMode.REASONING_PLANNING, scene=item.scene, num_samples=6 ) traj = torch.as_tensor(result.trajectories).float().cpu() endpoints = traj[:, -1, :2] # (6, 2) x, y at the last step (t = 5 s) centroid = endpoints.mean(dim=0) spread = (endpoints - centroid).norm(dim=1).max().item() flag = "WIDE" if spread > SPREAD_THRESHOLD else "ok" print(f"scene {i}: max endpoint deviation {spread:.2f} [{flag}]")Compare planner-sft and planner-rl
The two Planning Experts trade off differently. In the card’s results, RL wins on NAVSIM, WOD-E2E RFS and AlpaSim, while SFT has lower PAI-AV displacement. Both support reasoning mode, so you can run them on the same scene. This script loads them one at a time so that only one planner is in memory.
import torchfrom qwen_drive import InferenceMode, QwenDriveForPlanningfrom qwen_drive.benchmarks import read_scene_filefrom qwen_drive.images import ImageArchive
scene = next( read_scene_file( "data/demo/planning_scenes.jsonl", image_archive=ImageArchive.open("data/demo/frames.parquet"), )).scene
for name in ["planner-sft", "planner-rl"]: model = QwenDriveForPlanning.from_pretrained( "Qwen-Drive-1.0-4B", planner=f"Qwen-Drive-1.0-4B/{name}", dtype=torch.bfloat16, attn_implementation="flash_attention_2", ).to("cuda").eval()
result = model.run(InferenceMode.REASONING_PLANNING, scene=scene, num_samples=6) traj = torch.as_tensor(result.trajectories).float().cpu() print(f"=== {name} ===") print(result.reasoning) print("mean endpoint (x, y):", traj[:, -1, :2].mean(dim=0).tolist())
del model torch.cuda.empty_cache()Gotchas
- Run
planner-rlin reasoning mode only. It was reward-optimized only on reasoning-conditioned rollouts. If you want direct planning (no reasoning text), useplanner-sft. The card doesn’t show the enum name for direct mode, so check the GitHub repo. - The card only shows code for planning. It says VQA works by loading the root directory without a planner, and that you attach the perception head with
QwenDrivePerception.from_pretrained("Qwen-Drive-1.0-4B/perception"). It doesn’t show the calls or output formats for either. Use the GitHub repo andscripts/demo.pyfor those. - Attention backend. The card’s example loads the model on CUDA in
bfloat16withattn_implementation="flash_attention_2". The card doesn’t say whether FlashAttention 2 is required or whether other attention backends work, and it doesn’t cover installing flash-attn. - Memory. The VLM alone is 9.1 GB of weights, and each planner adds 2.1 GB. Activations for
num_samples=6come on top of that. The card gives no VRAM numbers. - Benchmark decoding isn’t default decoding. The VQA scores use near-greedy settings (top-p=0.001, top-k=1, temperature=0.01). The LingoQA score of 77.8 was judged by Qwen-Plus. Under the official LingoJudge, the score is 79.4.
- The perception head is a probe. The authors kept it simple on purpose so it works as an inspectable view into the VLM’s 3D representations. It isn’t meant to compete on perception benchmarks. The card’s text gives no perception metrics. It only includes a results image.
- RL trades some open-loop accuracy for safety. RL is behind SFT on PAI-AV ADE (0.42 vs. 0.37 at 3 s), and Alpamayo-1.5 still leads AlpaSim (0.45 vs. 0.37).
When to pick it
Pick Qwen-Drive-1.0-4B if you do driving research and want one Apache-2.0 model that plans trajectories, answers questions about scenes and gives you inspectable 3D outputs from a shared backbone. Driving QA is where its lead is clearest. LingoQA, WaymoQA, CoC and Ego3D all improve substantially over both the base Qwen3.5-4B and the driving specialists in the card. On general VL benchmarks it stays roughly on par with the base model, a little behind on most of them. For planning, the RL model has the top score among the compared methods on NAVSIM PDMS and WOD-E2E RFS.
Look elsewhere in these cases:
- You need the best closed-loop safety score in the card’s comparison. Alpamayo-1.5 scores higher on AlpaSim.
- You need state-of-the-art 3D detection or occupancy. The perception head isn’t designed for that.
- You want planning or perception without an extra package. The card’s code for both uses the
qwen_drivepackage and its scene format. The card doesn’t show how the VLM is loaded for VQA.
All the evaluations in the card are open-loop, pseudo-closed-loop or simulation. Treat it as a research foundation model, not a component ready to control a vehicle.
Related
DeepSeek-V4.1-Flash: a 1M-context multimodal MoE built for agent work
DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.
deepseek-ai/DeepSeek-V4.1-Flash
Run Gemma 4 31B for Image, Video and Reasoning Tasks
Google DeepMind's largest dense Gemma 4 model handles images, video frames and 256K-token context. Here's how to run it with Transformers.
google/gemma-4-31B-it
GLM-5.3-Flash: What You Can Run With Z.ai's 18B-Active Multimodal Model
Z.ai's first natively multimodal GLM-5 model: 320B total / 18B active parameters, hybrid sparse-linear attention, MIT license. What the card covers.
zai-org/GLM-5.3-Flash