Alibaba: qwen2.5-vl-72b-instruct API
Qwen2.5-VL-72B-Instruct is the flagship vision-language model of the Qwen2.5 series, featuring 72 billion parameters. It serves as a powerful multi-modal foundation capable of state-of-the-art performance in complex visual reasoning, document parsing, and agentic tasks. Compared to its predecessor (Qwen2-VL), the 2.5 series introduces significant architectural refinements, including window attention in the vision encoder for higher efficiency and native support for dynamic resolution. This model is engineered to be a "Visual Agent," capable of precisely interacting with computer and mobile UIs, understanding videos longer than an hour, and providing stable structured outputs (JSON) for complex data extraction from invoices, forms, and technical charts.
- Input: text, image, video
- Output: text
- Reasoning: Supported
- Tool calling: Supported
- File input: Supported
- Released: 2025-01-27
- Knowledge cutoff: 2024-11-30
Frequently Asked Questions
Does qwen2.5-vl-72b-instruct support function calling?
Yes. qwen2.5-vl-72b-instruct supports tool / function calling.
Does qwen2.5-vl-72b-instruct support reasoning?
Yes. qwen2.5-vl-72b-instruct is a reasoning-capable model.
What is the knowledge cutoff of qwen2.5-vl-72b-instruct?
The knowledge cutoff of qwen2.5-vl-72b-instruct is 2024-11-30.