Qwen Blog ยท 2026-05-29

Qwen-VLA: From Understanding the World to Acting in It

Over the past few years, multimodal large language models have become increasingly capable of understanding images, videos, and real-world scenes. They can recognize objects, reason about spatial relationships, answer visual questions, and solve complex multimodal reasoning ta...

Open original