
Qwen2.5 VL 3B is a multimodal LLM from the Qwen Team with the following key enhancements:
SoTA understanding of images of various resolution & ratio: Qwen2.5-VL achieves state-of-the-art performance on visual understanding benchmarks, including MathVista, DocVQA, RealWorldQA, MTVQA, etc.
Agent that can operate your mobiles, robots, etc.: with the abilities of complex reasoning and decision making, Qwen2.5-VL can be integrated with devices like mobile phones, robots, etc., for automatic operation based on visual environment and text instructions.
Multilingual Support: to serve global users, besides English and Chinese, Qwen2.5-VL now supports the understanding of texts in different languages inside images, including most European languages, Japanese, Korean, Arabic, Vietnamese, etc.
For more details, see this blog post(opens in new tab) and GitHub repo(opens in new tab).
Usage of this model is subject to Tongyi Qianwen LICENSE AGREEMENT(opens in new tab).
Modalities
Context
64K
Released
Mar 26, 2025
Knowledge Cutoff
Jun 2024
Qwen2.5 VL 3B is a multimodal LLM from the Qwen Team with the following key enhancements: - SoTA understanding of images of various resolution & ratio: Qwen2.5-VL achieves state-of-the-art performance on visual understanding benchmarks, including MathVista, DocVQA, RealWorldQA, MTVQA, etc.
Qwen2.5 VL 3B Instruct has a 64,000 token context window.
Qwen2.5 VL 3B Instruct accepts text and images as input and returns text.
Qwen3.8 27B, Qwen3.8 2.4T A95B, Qwen3.8 Max and 47 more are other text models from Qwen.
Qwen2.5 VL 3B Instruct was released on March 26, 2025. Its knowledge cutoff is June 30, 2024.
Token volume and request traffic to this model over time.