Realtime AI News
Qwen debuts Qwen-Drive-1.0-4B, a vision-language model for autonomous driving
Qwen has listed Qwen-Drive-1.0-4B on Hugging Face, a 4B-parameter vision-language model built for autonomous driving scenarios. The image-text-to-text model covers motion planning, 3D perception and visual question answering, extending Qwen's push from general chat models toward physical-world tasks.

Qwen has listed Qwen/Qwen-Drive-1.0-4B on the official Hugging Face model hub as a model update, marking the team's latest step into multimodal models for autonomous driving.
The model card describes an image-text-to-text model built on the transformers framework with safetensors weights and a 4B parameter scale.
The page's tags make the positioning clear: qwen_drive, autonomous-driving, motion-planning, 3d-perception and visual-question-answering sit side by side, pointing to a multimodal model aimed squarely at driving research.
In other words, Qwen-Drive-1.0-4B is designed to combine visual perception with language reasoning, letting the model read road scenes while responding to driving-related planning and question-answering tasks.
The compact 4B size is another key signal: compared with general-purpose models that run to hundreds of billions of parameters, a model at this scale is far easier to experiment with, fine-tune, and eventually deploy on vehicle-side or edge hardware.
The release extends Qwen's push to bring large-model capabilities down to physical-world scenarios, and it reinforces how vision-language models are becoming a more common ingredient in the autonomous driving stack.
What to watch next: real-road and closed-loop simulation results for the model, community experiments in motion planning and 3D perception, and whether Qwen follows up with larger Drive-series versions.
Why it matters
A compact 4B driving-focused vision-language model is a meaningful signal that multimodal AI is moving into physical-world tasks. Watch for community benchmarks and whether a larger Drive model follows.
Nearby Updates
All09/02, 14:20
Ant Group's OmniTable wins VLDB 2026 industrial best paper, handling 35PB of LLM corpus
QbitAI reports that Ant Group's unified wide-table system, OmniTable, has won the best-paper award in the industrial track at VLDB 2026. Built for large-model data preparation, the system handles a 35PB corpus on a single wide-table architecture and reportedly lifts processing efficiency by 5.6x.
09/02, 14:10
Alibaba updates flagship Qwen3.8-Max, front-end coding tops global leaderboard
Alibaba has refreshed its flagship Qwen3.8-Max with post-training tuned for coding and professional office work, delivering a significant performance gain over the previous version. The new model scored 1691 on CodeArena, a leading front-end coding leaderboard, up 22 points, to place first overall ahead of models such as Claude Opus 5 and Kimi K3.
09/02, 14:10
Alibaba updates flagship Qwen3.8-Max, tops CodeArena front-end coding leaderboard
On September 2, Alibaba updated its flagship Qwen3.8-Max model with post-training focused on coding and professional office work, lifting it to first place on the CodeArena front-end WebDev leaderboard with a score of 1691. Alibaba says the new version averages about $5 per million tokens, and it is now live on the Qwen AI platform's API with Qwen Office, Qoder, and the Qwen app already integrated.
09/02, 14:06
Former ByteDance RL expert Sun Peng joins Physical AI firm 星尘智能 to round out its stack
QbitAI reports that Sun Peng, a former reinforcement-learning expert at ByteDance and former head of the agent center at Tencent Robotics X, formally joined 星尘智能 on September 2. The hire is aimed at completing the company's full-stack technology layout in Physical AI.