Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 6 min read

Qwen-Drive-1.0-4B: Generate Driving Trajectories with a 4B VLM

Qwen's driving model adds a BEV perception head and a flow-matching planner to Qwen3.5-4B. Install it and sample ego trajectories from demo scenes.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

Qwen-Drive-1.0-4B is a vision-language model for autonomous driving from the Qwen team. It is built on Qwen3.5-4B. What sets it apart is that the team left the base VLM’s architecture completely unchanged and attached two external modules to it. The first is a BEV perception head that does 3D detection, semantic occupancy and BEV map segmentation. The second is a Planning Expert that generates future ego trajectories with flow matching. The VLM by itself still answers free-form questions about driving scenes. One download gives you VQA, 3D perception and motion planning, all sharing the same backbone. All of it is released under Apache-2.0.

Key specs

Base model Qwen/Qwen3.5-4B (architecture unchanged)
VLM weights (repo root) 9.1 GB, serves VQA by itself
planner-sft/ 2.1 GB, imitation-trained, supports direct and reasoning planning
planner-rl/ 2.1 GB, reward-optimized on NAVSIM PDMS, WOD-E2E RFS and a displacement term
perception/ 0.5 GB BEV head
Trajectory output (num_samples, 50, 3): x, y, heading, 5 s at 10 Hz

Selected numbers from the card:

  • NAVSIM PDMS: 90.7 for the RL model, against 90.3 for SpanVLA and SimWAM and 89.6 for AutoVLA. The card also reports a best-of-6 score of 91.4, which keeps the best of six samples. None of the compared methods report a best-of-6 number, so don’t compare 91.4 with their scores.
  • WOD-E2E RFS (val/test): 8.45 / 7.91 for the RL model, against 8.20 / 7.87 for MindVLA-U1.
  • AlpaSim closed-loop at-fault score: 0.37 for the RL model. Alpamayo-1.5 leads at 0.45.
  • PAI-AV ADE 3s: 0.37 for SFT and 0.42 for RL. Alpamayo-1.5 is best at 0.35.
  • Driving VQA (SFT): LingoQA 77.8 (base Qwen3.5-4B: 70.4), WaymoQA all 74.5 (base: 67.1), CoC all 41.3 (base: 2.6).
  • General VL (SFT vs. base): The card describes the SFT model as “on par” with the base model and says it “largely preserves” general ability. SFT is a little behind the base on most of these benchmarks: MMBench 85.5 vs. 87.1, MMMU 72.7 vs. 73.4, MMMU-Pro std 62.7 vs. 64.9, CountQA 31.7 vs. 35.9. It is ahead on a few, such as MMStar 75.9 vs. 75.3 and RealWorldQA 79.0 vs. 76.3.

Install

The card’s planning code uses the inference package from the GitHub repo:

Terminal window
git clone https://github.com/QwenLM/Qwen-Drive-1.0 qwen-drive && cd qwen-drive
pip install -e . --no-build-isolation

Download everything: the VLM at the root plus all task heads in their subfolders.

Terminal window
hf download Qwen/Qwen-Drive-1.0-4B --local-dir Qwen-Drive-1.0-4B

The repo includes demo data under data/demo/. The examples below use it. The card also mentions scripts/demo.py in the GitHub repo, which runs four bundled planning scenes and six perception frames end to end.

Run it

This snippet comes from the card. It loads the VLM with the RL planner attached, reads one demo scene and samples six trajectories in reasoning mode:

quickstart.py
import torch
from qwen_drive import InferenceMode, QwenDriveForPlanning
from qwen_drive.benchmarks import read_scene_file
from qwen_drive.images import ImageArchive
model = QwenDriveForPlanning.from_pretrained(
"Qwen-Drive-1.0-4B",
planner="Qwen-Drive-1.0-4B/planner-rl",
dtype=torch.bfloat16,
attn_implementation="flash_attention_2",
).to("cuda").eval()
scene = next(
read_scene_file(
"data/demo/planning_scenes.jsonl",
image_archive=ImageArchive.open("data/demo/frames.parquet"),
)
).scene
result = model.run(InferenceMode.REASONING_PLANNING, scene=scene, num_samples=6)
print(result.reasoning)
print(result.trajectories.shape) # (6, 50, 3) -> (x, y, heading), 5 s at 10 Hz

result.reasoning holds the model’s text rationale. result.trajectories holds the sampled paths.

Batch-plan a scene file and save trajectories

The card calls next() on read_scene_file, so it returns an iterator. That means you should be able to loop over every scene in a file and write the results out for offline scoring or visualization. The card only shows reading the first scene, so looping over the whole file is our assumption. The card also doesn’t give the exact types of the result fields. Wrapping trajectories in torch.as_tensor and reasoning in str() keeps the output serializable either way.

batch_plan.py
import json
import torch
from qwen_drive import InferenceMode, QwenDriveForPlanning
from qwen_drive.benchmarks import read_scene_file
from qwen_drive.images import ImageArchive
model = QwenDriveForPlanning.from_pretrained(
"Qwen-Drive-1.0-4B",
planner="Qwen-Drive-1.0-4B/planner-rl",
dtype=torch.bfloat16,
attn_implementation="flash_attention_2",
).to("cuda").eval()
archive = ImageArchive.open("data/demo/frames.parquet")
scenes = read_scene_file("data/demo/planning_scenes.jsonl", image_archive=archive)
with open("trajectories.jsonl", "w") as out:
for i, item in enumerate(scenes):
result = model.run(
InferenceMode.REASONING_PLANNING, scene=item.scene, num_samples=6
)
traj = torch.as_tensor(result.trajectories).float().cpu()
out.write(json.dumps({
"scene_index": i,
"reasoning": str(result.reasoning),
"trajectories": traj.tolist(), # [6][50][x, y, heading]
}) + "\n")
print(f"scene {i}: {tuple(traj.shape)}")

Look at the spread between samples

With num_samples=6 you get six trajectories for each scene. Measuring how far apart their endpoints land can be useful when you inspect results. The card doesn’t say how much the samples differ, and it doesn’t say whether the spread tracks scene difficulty. If you want to use spread to flag scenes, validate that on your own data first. The card also gives no units or coordinate frame for x and y. The threshold below uses the same units as the model output and is only a placeholder.

spread.py
import torch
from qwen_drive import InferenceMode, QwenDriveForPlanning
from qwen_drive.benchmarks import read_scene_file
from qwen_drive.images import ImageArchive
# Placeholder, in the model's output units (the card doesn't state them).
SPREAD_THRESHOLD = 2.0
model = QwenDriveForPlanning.from_pretrained(
"Qwen-Drive-1.0-4B",
planner="Qwen-Drive-1.0-4B/planner-rl",
dtype=torch.bfloat16,
attn_implementation="flash_attention_2",
).to("cuda").eval()
archive = ImageArchive.open("data/demo/frames.parquet")
for i, item in enumerate(
read_scene_file("data/demo/planning_scenes.jsonl", image_archive=archive)
):
result = model.run(
InferenceMode.REASONING_PLANNING, scene=item.scene, num_samples=6
)
traj = torch.as_tensor(result.trajectories).float().cpu()
endpoints = traj[:, -1, :2] # (6, 2) x, y at the last step (t = 5 s)
centroid = endpoints.mean(dim=0)
spread = (endpoints - centroid).norm(dim=1).max().item()
flag = "WIDE" if spread > SPREAD_THRESHOLD else "ok"
print(f"scene {i}: max endpoint deviation {spread:.2f} [{flag}]")

Compare planner-sft and planner-rl

The two Planning Experts trade off differently. In the card’s results, RL wins on NAVSIM, WOD-E2E RFS and AlpaSim, while SFT has lower PAI-AV displacement. Both support reasoning mode, so you can run them on the same scene. This script loads them one at a time so that only one planner is in memory.

sft_vs_rl.py
import torch
from qwen_drive import InferenceMode, QwenDriveForPlanning
from qwen_drive.benchmarks import read_scene_file
from qwen_drive.images import ImageArchive
scene = next(
read_scene_file(
"data/demo/planning_scenes.jsonl",
image_archive=ImageArchive.open("data/demo/frames.parquet"),
)
).scene
for name in ["planner-sft", "planner-rl"]:
model = QwenDriveForPlanning.from_pretrained(
"Qwen-Drive-1.0-4B",
planner=f"Qwen-Drive-1.0-4B/{name}",
dtype=torch.bfloat16,
attn_implementation="flash_attention_2",
).to("cuda").eval()
result = model.run(InferenceMode.REASONING_PLANNING, scene=scene, num_samples=6)
traj = torch.as_tensor(result.trajectories).float().cpu()
print(f"=== {name} ===")
print(result.reasoning)
print("mean endpoint (x, y):", traj[:, -1, :2].mean(dim=0).tolist())
del model
torch.cuda.empty_cache()

Gotchas

  • Run planner-rl in reasoning mode only. It was reward-optimized only on reasoning-conditioned rollouts. If you want direct planning (no reasoning text), use planner-sft. The card doesn’t show the enum name for direct mode, so check the GitHub repo.
  • The card only shows code for planning. It says VQA works by loading the root directory without a planner, and that you attach the perception head with QwenDrivePerception.from_pretrained("Qwen-Drive-1.0-4B/perception"). It doesn’t show the calls or output formats for either. Use the GitHub repo and scripts/demo.py for those.
  • Attention backend. The card’s example loads the model on CUDA in bfloat16 with attn_implementation="flash_attention_2". The card doesn’t say whether FlashAttention 2 is required or whether other attention backends work, and it doesn’t cover installing flash-attn.
  • Memory. The VLM alone is 9.1 GB of weights, and each planner adds 2.1 GB. Activations for num_samples=6 come on top of that. The card gives no VRAM numbers.
  • Benchmark decoding isn’t default decoding. The VQA scores use near-greedy settings (top-p=0.001, top-k=1, temperature=0.01). The LingoQA score of 77.8 was judged by Qwen-Plus. Under the official LingoJudge, the score is 79.4.
  • The perception head is a probe. The authors kept it simple on purpose so it works as an inspectable view into the VLM’s 3D representations. It isn’t meant to compete on perception benchmarks. The card’s text gives no perception metrics. It only includes a results image.
  • RL trades some open-loop accuracy for safety. RL is behind SFT on PAI-AV ADE (0.42 vs. 0.37 at 3 s), and Alpamayo-1.5 still leads AlpaSim (0.45 vs. 0.37).

When to pick it

Pick Qwen-Drive-1.0-4B if you do driving research and want one Apache-2.0 model that plans trajectories, answers questions about scenes and gives you inspectable 3D outputs from a shared backbone. Driving QA is where its lead is clearest. LingoQA, WaymoQA, CoC and Ego3D all improve substantially over both the base Qwen3.5-4B and the driving specialists in the card. On general VL benchmarks it stays roughly on par with the base model, a little behind on most of them. For planning, the RL model has the top score among the compared methods on NAVSIM PDMS and WOD-E2E RFS.

Look elsewhere in these cases:

  • You need the best closed-loop safety score in the card’s comparison. Alpamayo-1.5 scores higher on AlpaSim.
  • You need state-of-the-art 3D detection or occupancy. The perception head isn’t designed for that.
  • You want planning or perception without an extra package. The card’s code for both uses the qwen_drive package and its scene format. The card doesn’t show how the VLM is loaded for VQA.

All the evaluations in the card are open-loop, pseudo-closed-loop or simulation. Treat it as a research foundation model, not a component ready to control a vehicle.

Related