Skip to content
Not available in this workspace
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Favicon for qwen

Qwen: Qwen2.5-VL 7B Instruct

qwen/qwen-2.5-vl-7b-instruct

Model weights

Qwen2.5 VL 7B is a multimodal LLM from the Qwen Team with the following key enhancements:

  • SoTA understanding of images of various resolution & ratio: Qwen2.5-VL achieves state-of-the-art performance on visual understanding benchmarks, including MathVista, DocVQA, RealWorldQA, MTVQA, etc.

  • Understanding videos of 20min+: Qwen2.5-VL can understand videos over 20 minutes for high-quality video-based question answering, dialog, content creation, etc.

  • Agent that can operate your mobiles, robots, etc.: with the abilities of complex reasoning and decision making, Qwen2.5-VL can be integrated with devices like mobile phones, robots, etc., for automatic operation based on visual environment and text instructions.

  • Multilingual Support: to serve global users, besides English and Chinese, Qwen2.5-VL now supports the understanding of texts in different languages inside images, including most European languages, Japanese, Korean, Arabic, Vietnamese, etc.

For more details, see this blog post(opens in new tab) and GitHub repo(opens in new tab).

Usage of this model is subject to Tongyi Qianwen LICENSE AGREEMENT(opens in new tab).

Modalities

Context

33K

Released

Aug 28, 2024

Knowledge Cutoff

Jun 2024

ActivityFAQ

Activity

Token volume and request traffic to this model over time.

About Qwen: Qwen2.5-VL 7B Instruct

OpenRouter makes Qwen: Qwen2.5-VL 7B Instruct available through a unified, OpenAI-compatible API using the model ID qwen/qwen-2.5-vl-7b-instruct.

Qwen: Qwen2.5-VL 7B Instruct accepts text and images and returns text. It has a 32,768-token context window.

It was released on August 28, 2024; its knowledge cutoff is June 30, 2024.

More models from Qwen

  • Qwen3.8 27B
  • Qwen3 Reranker 8B
  • Qwen3 ASR 1.7B

Frequently asked questions

Qwen2.5 VL 7B is a multimodal LLM from the Qwen Team with the following key enhancements: - SoTA understanding of images of various resolution & ratio: Qwen2.5-VL achieves state-of-the-art performance on visual understanding benchmarks, including MathVista, DocVQA, RealWorldQA, MTVQA, etc.

Qwen2.5-VL 7B Instruct has a 32,768 token context window.

Qwen2.5-VL 7B Instruct accepts text and images as input and returns text.

Qwen3.8 27B, Qwen3.8 2.4T A95B, Qwen3.8 Max and 47 more are other text models from Qwen.

Qwen2.5-VL 7B Instruct was released on August 28, 2024. Its knowledge cutoff is June 30, 2024.