Qwen API
Access Qwen instantly with Puter.js, and add AI to any app in a few lines of code without backend or API keys.
// npm install @heyputer/puter.js
import { puter } from '@heyputer/puter.js';
puter.ai.chat("Explain AI like I'm five!", {
model: "qwen/qwen3.8-flash"
}).then(response => {
console.log(response);
});
<html>
<body>
<script src="https://js.puter.com/v2/"></script>
<script>
puter.ai.chat("Explain AI like I'm five!", {
model: "qwen/qwen3.8-flash"
}).then(response => {
console.log(response);
});
</script>
</body>
</html>
List of Qwen Models
Qwen3.6 Flash
qwen/qwen3.6-flash
Qwen3.6 Flash is the speed-optimized tier of Alibaba's Qwen3.6 model family, designed for high-throughput, low-latency inference pipelines. It sits alongside Qwen3.6 Max Preview, Plus, and 35B-A3B in the product lineup, targeting use cases where fast response times matter more than peak benchmark scores. Like other Qwen3.6 models, it builds on a hybrid architecture combining linear attention with sparse mixture-of-experts routing. It is best suited for high-volume production workloads such as classification, extraction, summarization, and lightweight agent tasks where latency and cost efficiency are the primary constraints.
ChatQwen3.5 Plus 2026-04-20
qwen/qwen3.5-plus-20260420
Qwen3.5 Plus is a proprietary hosted model from Alibaba, built on the Qwen3.5-397B-A17B Mixture-of-Experts architecture with 397 billion total parameters and 17 billion active per token. Its headline feature is a 1-million-token native context window — among the largest available via API — making it well suited for processing entire codebases, long documents, or extended multi-turn conversations in a single request. It supports both a deep-thinking mode and an "Auto" mode that adaptively invokes tools like web search and code interpreters. This April 20, 2026 snapshot reflects ongoing improvements to the model since its original February 2026 launch. The Qwen3.5 series demonstrated strong multimodal performance across reasoning, coding, and vision tasks. A solid general-purpose option for developers needing large-context capabilities without migrating to the newer Qwen3.6 line.
ChatQwen3.6 27B
qwen/qwen3.6-27b
Qwen3.6 27B is a dense 27-billion-parameter multimodal model from Alibaba's Qwen team, purpose-built for agentic coding and repository-level reasoning. It scores 77.2% on SWE-bench Verified and 59.3% on Terminal-Bench 2.0, outperforming the previous-generation Qwen3.5-397B-A17B across all major coding benchmarks despite being far smaller. It natively supports text, image, and video inputs with a 262K-token context window, extendable to 1M tokens. A standout feature is Thinking Preservation, which retains reasoning traces across conversation turns — reducing redundant computation in multi-step agent loops. The model uses a hybrid attention architecture combining Gated DeltaNet with traditional self-attention. Ideal for developers building coding agents, multi-turn tool-use workflows, or frontend generation pipelines.
ChatQwen3.6 Max Preview
qwen/qwen3.6-max-preview
Qwen3.6 Max Preview is Alibaba's most capable language model to date — a proprietary flagship that claimed the top score on six major coding benchmarks at its April 20, 2026 release. It leads on SWE-bench Pro, Terminal-Bench 2.0, SkillsBench, QwenClawBench, QwenWebBench, and SciCode. The Artificial Analysis Intelligence Index rates it at 52, well above the median for reasoning models in its price tier. It supports a 256K-token context window and is text-only at launch. As a preview release, Alibaba is still actively iterating on the model. Best suited for teams building coding agents, scientific computing tools, or frontend generation systems that need peak benchmark performance.
ChatQwen3.6 35B A3B
qwen/qwen3.6-35b-a3b
Qwen3.6 35B A3B is a sparse Mixture-of-Experts model with 35 billion total parameters but only 3 billion active per token, making it highly efficient for inference. Developed by Alibaba's Qwen team, it scores 73.4% on SWE-bench Verified and 51.5% on Terminal-Bench 2.0 — significantly outperforming dense models like Gemma 4-31B (52.0% on SWE-bench Verified). It natively handles text, image, and video with a 262K-token context window, extendable to 1M tokens. The model supports Thinking Preservation for stable multi-turn reasoning and includes native tool-calling capabilities. Released under Apache 2.0, it was the first open-weight model in the Qwen3.6 family. A strong choice for developers who want frontier-adjacent coding performance at a fraction of the compute cost of larger models.
ChatQwen3.6 Plus
qwen/qwen3.6-plus
Qwen 3.6 Plus is Alibaba's flagship large language model, built on a hybrid architecture combining linear attention with sparse mixture-of-experts routing for high throughput and scalability. It's optimized for agentic coding and complex multi-step workflows. On Terminal-Bench 2.0, it scores 61.6, surpassing Claude 4.5 Opus (59.3), while its 78.8 on SWE-bench Verified places it close behind. It also leads on MCPMark (48.2%) for tool-calling reliability. A native multimodal model, it handles text, images, and documents within a 1M-token context window with up to 65K output tokens. Notable features include always-on chain-of-thought reasoning, native function calling, and a preserve_thinking parameter that retains reasoning across multi-turn agent loops. A strong fit for developers building AI coding agents, terminal automation, and tool-using pipelines.
ChatQwen3.6 Plus Preview
qwen/qwen3.6-plus-preview:free
Qwen 3.6 Plus Preview is a next-generation large language model from Alibaba's Qwen team, built on a hybrid architecture designed for improved efficiency and scalability. Released as an early preview in March 2026, it succeeds the Qwen 3.5 Plus series with stronger reasoning and more reliable agentic behavior. The model offers a 1-million-token context window and up to 65,536 output tokens, making it well suited for processing large codebases, lengthy documents, or multi-step workflows in a single request. It supports tool use and function calling natively, with built-in chain-of-thought reasoning that is always active. Qwen 3.6 Plus Preview is particularly strong in agentic coding, front-end component generation, and complex problem-solving. It's a good fit for developers building AI-driven code review tools, multi-step agents, or applications that benefit from deep reasoning over large inputs.
ChatQwen3.5-9B
qwen/qwen3.5-9b
Qwen 3.5 9B is a 9-billion parameter open-source multimodal model by Alibaba's Qwen Team, featuring a 262K native context window (extendable to ~1M tokens), support for text, image, and video input, and coverage of 201 languages. It uses a hybrid Gated DeltaNet architecture and outperforms much larger models like Qwen3-30B and OpenAI's gpt-oss-120B on key benchmarks including reasoning, vision, and document understanding.
ImageQwen Image 2.0
qwen/qwen-image-2.0
Qwen Image 2.0 is Alibaba's second-generation image foundation model, delivering a major upgrade over the original Qwen Image with a leaner 7B-parameter architecture that outperforms its 20B predecessor across the board. It generates natively at 2048×2048 resolution and unifies text-to-image generation and image editing into a single model — no separate pipelines needed. The model scores 88.32 on DPG-Bench, surpassing FLUX.1 (83.84) and GPT Image 1 (85.15), and ranks #1 on AI Arena's blind human evaluation for both generation and editing. Its headline feature is professional typography rendering: it handles prompts up to 1,000 tokens and can generate complete infographics, PPT slides, posters, and comics with accurate bilingual text layout. Ideal for developers building design-oriented workflows where text accuracy and prompt adherence are critical.
ImageQwen Image 2.0 Pro
qwen/qwen-image-2.0-pro
Qwen Image 2.0 Pro is the highest-fidelity configuration of Alibaba's Qwen Image 2.0, built on the same 7B-parameter architecture but tuned to maximize visual quality over speed. Compared to the standard tier, Pro delivers richer color accuracy, finer detail rendering — visible in textures like hair strands, fabric weaves, and metallic reflections — and stronger adherence to complex, multi-element prompts. Text rendering is also crisper, making it better suited for commercial assets like branded posters and packaging. The standard Qwen Image 2.0 is optimized for fast iteration and prototyping. Pro is where you go for final production renders where every pixel matters. Best for developers building pipelines that need polished, client-ready output from a single API call.
ChatQwen3.5-Flash
qwen/qwen3.5-flash-02-23
Qwen 3.5 Flash is the production-optimized API version of the 35B-A3B model. It features a default 1M token context window, built-in tool/function calling support, and is priced at ~$0.10/M input tokens for low-latency agentic workflows. The '02-23' suffix indicates the February 23, 2026 snapshot/version date.
ChatQwen3.5-122B-A10B
qwen/qwen3.5-122b-a10b
Qwen 3.5 122B (10B Active) is Alibaba's largest medium-sized MoE model, activating only 10B of its 122B total parameters per inference pass. It excels at agentic tasks like tool use and multi-step reasoning, leading the Qwen 3.5 lineup on benchmarks such as BFCL-V4 and BrowseComp. It supports 262K native context (extendable to 1M), native multimodal input, and 201 languages under Apache 2.0.
ChatQwen3.5-27B
qwen/qwen3.5-27b
Qwen 3.5 27B is the only dense (non-MoE) model in the Qwen 3.5 medium series, activating all 27B parameters on every forward pass for maximum per-token reasoning density. It ties GPT-5 mini on SWE-bench Verified at 72.4 and is competitive with Claude Sonnet 4.5 on visual reasoning benchmarks. It runs well on consumer hardware and is open-weight under Apache 2.0.
ChatQwen3.5-35B-A3B
qwen/qwen3.5-35b-a3b
Qwen 3.5 35B (3B Active) is a sparse MoE model that activates just 3B of its 35B total parameters, yet outperforms the previous-generation 235B flagship across language, vision, coding, and agent tasks. It uses a hybrid Gated DeltaNet + MoE architecture and can run on GPUs with as little as 8GB VRAM when quantized. It's the base model behind the hosted Qwen 3.5 Flash API.
ChatQwen3.5 Plus 02-15
qwen/qwen3.5-plus-02-15
Qwen3.5-Plus is the hosted flagship model in the Qwen3.5 series, available through Alibaba Cloud Model Studio. It offers a 1 million token context window by default and includes built-in tools with adaptive tool use, including web search and code interpreter capabilities. The model supports reasoning mode (chain-of-thought), search, and a fast response mode without extended thinking. It is accessible via an OpenAI-compatible API and can be integrated with third-party coding tools like Claude Code, Cline, and OpenClaw. Qwen3.5-Plus is designed for agentic workflows that combine multimodal reasoning with tool use.
ChatQwen3.5 Plus
qwen/qwen3.5-plus
Qwen3.5 Plus is Alibaba's hosted flagship model in the Qwen3.5 series, built on the Qwen3.5-397B-A17B Mixture-of-Experts architecture with 397 billion total parameters and 17 billion active per token. Its headline feature is a 1-million-token native context window — among the largest available via API — making it well suited for processing entire codebases, long documents, or extended multi-turn conversations in a single request. It supports both a deep-thinking mode and an "Auto" mode that adaptively invokes tools like web search and code interpreters. A solid general-purpose option for developers needing large-context capabilities and agentic workflows that combine multimodal reasoning with tool use.
ChatQwen3.5 397B A17B
qwen/qwen3.5-397b-a17b
Qwen3.5-397B-A17B is an open-weight native vision-language model from Alibaba's Qwen team, released in February 2026. It uses a hybrid architecture combining Gated Delta Networks (linear attention) with a sparse mixture-of-experts design, totaling 397 billion parameters but activating only 17 billion per forward pass for efficient inference. The model delivers strong performance across reasoning, coding, agent tasks, and multimodal understanding, competing with frontier models like GPT-5.2, Claude 4.5 Opus, and Gemini-3 Pro. It supports 201 languages and dialects and features a 250k-token vocabulary. Its decoding throughput is reported at 8.6x that of Qwen3-Max under a 32k context length.
ChatQwen3 Max Thinking
qwen/qwen3-max-thinking
Qwen3 Max Thinking is Alibaba Cloud's flagship proprietary reasoning model with a 256K context window, featuring test-time scaling and adaptive tool-use capabilities (web search, code interpreter, memory) that allow it to reason iteratively and autonomously. It scores competitively against GPT-5.2 and Gemini 3 Pro on benchmarks like Humanity's Last Exam and HMMT, excelling in math, complex reasoning, and instruction following.
ChatQwen3 Coder Next
qwen/qwen3-coder-next
Qwen3-Coder-Next is an open-weight coding model from Alibaba's Qwen team with 80B total parameters but only 3B active per token, designed specifically for coding agents and local development with a 256K context window. It uses a sparse Mixture-of-Experts (MoE) architecture with hybrid attention, trained on 800K executable coding tasks using reinforcement learning to excel at long-horizon reasoning, tool calling, and recovering from execution failures. It achieves performance comparable to models with 10-20x more active parameters on benchmarks like SWE-Bench while maintaining low inference costs.
ChatQwen3 VL 32B Instruct
qwen/qwen3-vl-32b-instruct
Qwen3 VL 32B Instruct is a dense vision-language model with strong text and visual capabilities, featuring visual coding, spatial understanding, and 256K context support.
ChatQwen3 VL 8B Instruct
qwen/qwen3-vl-8b-instruct
Qwen3 VL 8B Instruct is a compact vision-language model matching flagship text performance while supporting image/video understanding, visual coding, and 256K context length.
ChatQwen3 VL 8B Thinking
qwen/qwen3-vl-8b-thinking
Qwen3 VL 8B Thinking is the reasoning-enhanced compact vision model for complex visual analysis requiring step-by-step reasoning with efficient resource usage.
ChatQwen3 VL 30B A3B Instruct
qwen/qwen3-vl-30b-a3b-instruct
Qwen3 VL 30B A3B Instruct is an efficient vision-language MoE model offering strong image/video understanding with 3B active parameters and 256K context support.
ChatQwen3 VL 30B A3B Thinking
qwen/qwen3-vl-30b-a3b-thinking
Qwen3 VL 30B A3B Thinking is the reasoning-enhanced vision-language variant optimized for complex visual reasoning tasks with extended thinking capabilities.
ChatQwen3 Max
qwen/qwen3-max
Qwen3 Max is the most powerful Qwen3 API model with SOTA agent programming and tool usage capabilities. It features non-thinking mode optimized for complex agent scenarios.
ChatQwen3 VL 235B A22B Thinking
qwen/qwen3-vl-235b-a22b-thinking
Qwen3 VL 235B A22B Thinking is the reasoning-enhanced vision-language model excelling at visual math, detail analysis, and causal reasoning with extended chain-of-thought processing.
ChatQwen3-VL Plus
qwen/qwen3-vl-plus
Qwen3-VL Plus is Alibaba Cloud's hosted vision-language API model in the Qwen3-VL series, offering strong multimodal understanding without requiring self-hosted infrastructure. It handles a wide range of visual tasks including document parsing, chart analysis, OCR, image reasoning, and GUI interaction for PC and mobile interfaces. With a 262K token context window, it is well suited for processing lengthy documents, multi-page PDFs, and extended visual conversations in a single request. The model supports structured output and tool calling, making it a practical choice for developers building document intelligence pipelines, visual agents, and multimodal data extraction workflows via the OpenAI-compatible Alibaba Cloud Model Studio API.
ChatQwen Plus 0728
qwen/qwen-plus-2025-07-28
Qwen Plus (2025-07-28) is a snapshot version of Qwen Plus from July 2025, offering consistent behavior and performance for production deployments requiring version stability.
ChatQwen Plus 0728 (thinking)
qwen/qwen-plus-2025-07-28:thinking
Qwen Plus (2025-07-28) Thinking is the reasoning-enhanced version that uses chain-of-thought processing for complex problems, providing step-by-step reasoning before delivering answers.
ChatQwen3 Next 80B A3B Instruct
qwen/qwen3-next-80b-a3b-instruct
Qwen3 Next 80B A3B Instruct is an innovative MoE model with hybrid attention (Gated DeltaNet + Gated Attention), achieving 10x inference throughput for 32K+ contexts while matching Qwen3-235B performance.
ChatQwen3 Next 80B A3B Thinking
qwen/qwen3-next-80b-a3b-thinking
Qwen3 Next 80B A3B Thinking is the reasoning-enhanced variant outperforming Gemini-2.5-Flash-Thinking on complex reasoning tasks with hybrid attention and multi-token prediction.
ImageQwen Image
qwen/qwen-image
Qwen Image is a 20B-parameter image generation foundation model from Alibaba's Qwen series, built for text-to-image generation, image editing, and image understanding tasks. Its standout capability is high-fidelity text rendering — it accurately places readable text in both English and Chinese within generated images, making it especially strong for posters, slides, and design-heavy visuals. Beyond text, it supports a wide range of styles from photorealism to anime, and handles advanced editing operations like style transfer, object insertion/removal, and in-image text modification. The model also performs image understanding tasks including object detection, segmentation, depth estimation, and super-resolution. A versatile choice for developers who need generation, editing, and visual analysis in a single model. Licensed under Apache 2.0.
ChatQwen3 Coder Flash
qwen/qwen3-coder-flash
Qwen3 Coder Flash is a cost-effective coding model balancing performance and speed, suitable for scenarios requiring fast responses at lower cost while maintaining coding quality.
ChatQwen Flash
qwen/qwen-flash
Qwen Flash is Alibaba's latency-optimized general-purpose language model, designed as the successor to Qwen Turbo for cost-efficient, high-throughput workloads. It offers a 1 million token context window with native support for context caching, making repeated or large-context requests significantly cheaper. The model supports function calling and is accessible via an OpenAI-compatible API through Alibaba Cloud Model Studio. Qwen Flash is a strong choice for developers running high-volume production tasks — classification, extraction, summarization, and lightweight agentic pipelines — where low latency and predictable pricing matter more than peak reasoning capability. Its flexible tiered pricing and context cache support make it especially cost-effective at scale.
ChatQwen3 Coder Plus
qwen/qwen3-coder-plus
Qwen3 Coder Plus is the strongest Qwen coding API model, ideal for complex project generation and in-depth code reviews with up to 1M token context support.
ChatQwen3 235B A22B Instruct 2507
qwen/qwen3-235b-a22b-2507
Qwen3 235B A22B (2507) is the July 2025 updated version with significant improvements in instruction following, reasoning, coding, tool usage, and 256K long-context understanding.
ChatQwen3 235B A22B Thinking 2507
qwen/qwen3-235b-a22b-thinking-2507
Qwen3 235B A22B Thinking (2507) is the reasoning-enhanced variant using extended chain-of-thought processing for complex math, coding, and logical problems with enhanced performance.
ChatQwen3 30B A3B Instruct 2507
qwen/qwen3-30b-a3b-instruct-2507
Qwen3 30B A3B Instruct (2507) is the July 2025 updated instruction-tuned version with improved capabilities in reasoning, coding, and tool usage at high efficiency.
ChatQwen3 30B A3B Thinking 2507
qwen/qwen3-30b-a3b-thinking-2507
Qwen3 30B A3B Thinking (2507) is the reasoning-enhanced variant optimized for complex problem-solving with extended chain-of-thought processing at high parameter efficiency.
ChatQwen3 4B
qwen/qwen3-4b
Qwen3 4B is a small dense model in Alibaba's Qwen3 family, released in April 2025. It supports hybrid thinking modes, switching between step-by-step reasoning for math, coding, and logic and a faster non-thinking mode for general dialogue. The mode can be toggled per request, including with /think and /no_think tags in prompts. The Qwen team reports that it rivals the much larger Qwen2.5-72B-Instruct, though independent evaluations have been more mixed. It handles 119 languages and dialects and supports tool calling, with the Qwen-Agent framework recommended for agent workloads. Native context is 32K tokens, extendable to 131K with YaRN scaling. It fits developers who want reasoning and multilingual coverage at low cost, such as high-volume chat, translation, or lightweight agent tasks.
ChatQwen3 14B
qwen/qwen3-14b
Qwen3 14B is a dense language model with hybrid thinking/non-thinking modes, matching Qwen2.5-32B performance. It supports 119 languages and excels in math, coding, and reasoning tasks.
ChatQwen3 235B A22B
qwen/qwen3-235b-a22b
Qwen3 235B A22B is the flagship MoE model with 235B total and 22B active parameters, rivaling DeepSeek-R1 and o1. It features hybrid thinking modes and supports 119 languages with strong agentic capabilities.
ChatQwen3 30B A3B
qwen/qwen3-30b-a3b
Qwen3 30B A3B is an efficient MoE model with 30B total and 3B active parameters, outperforming QwQ-32B while using 10x fewer active parameters. It offers hybrid thinking modes and 119 language support.
ChatQwen3 32B
qwen/qwen3-32b
Qwen3 32B is a dense language model matching Qwen2.5-72B performance with hybrid thinking/non-thinking modes. It excels in STEM, coding, and reasoning while supporting 119 languages.
ChatQwen3 8B
qwen/qwen3-8b
Qwen3 8B is a dense model matching Qwen2.5-14B performance with hybrid thinking modes and 128K context. It offers strong reasoning, coding, and multilingual capabilities in a mid-sized package.
ChatQwen3 Coder 480B A35B
qwen/qwen3-coder-480b-a35b-instruct
Qwen3 Coder is the most agentic code model in the Qwen series, available in 30B and 480B MoE variants. It achieves SOTA on SWE-Bench with 256K native context, extendable to 1M tokens.
ChatQwen3 Coder 30B A3B Instruct
qwen/qwen3-coder-30b-a3b-instruct
Qwen3 Coder 30B A3B Instruct is an efficient MoE coding model with 30B total and 3.3B active parameters, offering strong agentic coding capabilities with 256K context support.
ChatQwen3 VL 235B A22B Instruct
qwen/qwen3-vl-235b-a22b
Qwen3 VL 235B A22B Instruct is the flagship vision-language MoE model with 256K context, offering superior visual coding, spatial understanding, and long video comprehension up to 20 minutes.
ChatQwen2.5 VL 32B Instruct
qwen/qwen2.5-vl-32b-instruct
Qwen 2.5 VL 32B Instruct is a mid-sized vision-language model offering enhanced image/video understanding with better alignment to human preferences. It bridges the gap between 7B and 72B variants.
ChatQwQ 32B
qwen/qwq-32b
QwQ 32B is a 32B parameter reasoning model rivaling DeepSeek-R1 (671B) through scaled reinforcement learning. It excels in math, coding, and complex reasoning with 131K context and agent capabilities.
ChatQwen2.5 VL 72B Instruct
qwen/qwen2.5-vl-72b-instruct
Qwen 2.5 VL 72B Instruct is the flagship open-source vision-language model excelling in document understanding, visual reasoning, and long video comprehension up to 1 hour with event pinpointing.
ChatQwen2.5 Coder 32B Instruct
qwen/qwen-2.5-coder-32b-instruct
Qwen 2.5 Coder 32B Instruct is a code-specialized model matching GPT-4o's coding capabilities, supporting 40+ programming languages. It excels in code generation, repair, and reasoning with 128K context support.
ChatQwen-Turbo
qwen/qwen-turbo
Qwen Turbo is a fast, cost-effective API model with up to 1M context length, ideal for simple tasks requiring quick responses. It supports multiple languages and offers flexible tiered pricing.
ChatQwen-VL OCR
qwen/qwen-vl-ocr
Qwen-VL OCR is Alibaba's specialized vision-language model purpose-built for text extraction and document parsing, derived from the Qwen-VL series. Unlike general-purpose VL models, it's optimized for OCR across scanned documents, tables, receipts, exam papers, forms, and handwritten content. It supports multilingual recognition including English, Chinese, French, German, Japanese, Korean, Russian, Italian, and Arabic. Capabilities include skewed image recognition, text localization with bounding box coordinates, table-to-HTML parsing, document-to-LaTeX conversion, and formula transcription. Built-in task modes return structured output as plain text, JSON, HTML, or LaTeX depending on the workflow. It's the right Qwen API choice for developers building document digitization, receipt parsing, or information extraction pipelines that need OCR-focused accuracy rather than general visual reasoning.
ChatQwen2.5 Coder 7B Instruct
qwen/qwen2.5-coder-7b-instruct
Qwen 2.5 Coder 7B Instruct is a compact code-specialized model with strong code generation, reasoning, and repair capabilities. It supports multiple programming languages while being deployable on consumer hardware.
ChatQwen2.5 72B Instruct
qwen/qwen2-5-72b-instruct
Qwen 2.5 72B Instruct is Alibaba's flagship open-source language model with 72 billion parameters, trained on 18 trillion tokens with 128K context support. It excels in coding, math, instruction following, and multilingual tasks across 29+ languages.
ChatQwen2.5 7B Instruct
qwen/qwen2-5-7b-instruct
Qwen 2.5 7B Instruct is a compact yet capable language model offering strong performance in coding, math, and general tasks. It supports 128K context length and 29+ languages while being efficient enough for smaller deployments.
ChatQwen2.5-VL 7B Instruct
qwen/qwen2-5-vl-7b-instruct
Qwen 2.5 VL 7B Instruct is a vision-language model capable of understanding images, documents, charts, and videos up to 1 hour. It supports OCR, visual reasoning, and can act as a visual agent for computer/phone use.
ChatQwen VL Max
qwen/qwen-vl-max
Qwen VL Max is Alibaba's most capable vision-language API model based on Qwen2.5-VL, offering superior image/video understanding, OCR, document analysis, and visual reasoning capabilities.
ChatQwen-Max
qwen/qwen-max
Qwen Max is Alibaba's most powerful proprietary API model, a large-scale MoE with hundreds of billions of parameters. It delivers top-tier performance in reasoning, coding, math, and multilingual tasks via Alibaba Cloud Model Studio.
ChatQwen-Plus
qwen/qwen-plus
Qwen Plus is a high-performance proprietary API model balancing capability and cost, suitable for complex tasks requiring strong reasoning and multilingual support. Available through Alibaba Cloud Model Studio.
ChatQwen VL Plus
qwen/qwen-vl-plus
Qwen VL Plus is a balanced vision-language API model offering good performance at lower cost, suitable for image understanding, OCR, and multimodal tasks without requiring maximum capability.
ChatQwen3.5-4B
qwen/qwen3.5-4b
Qwen3.5-4B is a compact multimodal model from Alibaba's Qwen3.5 Small series, released alongside 0.8B, 2B, and 9B siblings in early 2026 as part of the family succeeding Qwen3. Unlike models that bolt a vision tower onto a text backbone, it processes text and images in a unified latent space, using a hybrid architecture that interleaves linear-attention (Gated DeltaNet) blocks with standard attention. Alibaba reports 79.1% on MMLU-Pro, 76.2% on GPQA Diamond, and 74-77% on HMMT, scores that Qwen says close the gap with much larger models. Qwen positions the 4B size as a multimodal base for lightweight agents rather than a pure edge model like its 0.8B and 2B siblings. Via this API it offers a 256K context window and 32K max output tokens, a fit for developers who want multimodal reasoning and OCR at low per-token cost.
ChatQwen3.7 Flash
qwen/qwen3.7-flash
Qwen3.7 Flash is the low-cost, fast tier of Alibaba's Qwen3.7 family, released in July 2026. It is a vision-language model that accepts text, image, and video input across a 1 million-token context window, with reasoning enabled by default and a 262K-token thinking budget. Alibaba positions it as an upgrade over Qwen3.6 Flash in multimodal understanding and agent execution, with better object recognition and spatial intelligence. It supports function calling and structured outputs. Pricing is tiered by prompt length. Requests under 32K input tokens cost $0.03/$0.13 per million, rising to $0.20/$0.80 above 256K. Alibaba published no benchmarks at launch. An independent vision evaluation by Roboflow measured strong object identification (84.4%) but weak OCR and object detection, so it fits high-volume multimodal tasks (classification, visual agents, lightweight extraction) better than document-heavy pipelines.
ChatQwen3.7 Plus
qwen/qwen3.7-plus
Qwen3.7 Plus is Alibaba's multimodal agent model, released in June 2026, combining vision-language understanding with full agentic capabilities across a 1 million-token context window. Unlike the text-only Qwen3.7 Max, Plus ingests images and video alongside text, processed through early-fusion training so vision and language are jointly understood from the first layer. This enables GUI grounding — the model can interpret screenshots and issue precise on-screen actions — scoring 79.0 on ScreenSpot Pro, placing it alongside Claude Computer Use and OpenAI Operator in the GUI automation tier. Beyond vision, it adds deep reasoning, self-programming, tool invocation, and autonomous iteration: the model writes and tests code, calls external APIs, and loops until the task is done. On the Artificial Analysis Intelligence Index it scores 53. Choose it over Qwen3.7 Max when your workflow requires image or video inputs, browser/desktop automation, or end-to-end agentic pipelines that combine seeing, reasoning, and doing.
ChatQwen3.7 Max
qwen/qwen3.7-max
Qwen3.7 Max is Alibaba's flagship proprietary reasoning model, released in May 2026, built for long-horizon agentic workloads with a 1 million-token context window and a chain-of-thought reasoning architecture. It is purpose-built for complex, multi-step autonomous tasks. Alibaba demonstrated the model running for 35 hours without degradation, executing over 1,000 tool calls in a single session — making it a strong candidate for coding agents, automated pipelines, and deep document analysis. On benchmarks, it ranks 13th globally on LM Arena's text leaderboard and scores 56.6 on the Artificial Analysis Intelligence Index, making it the highest-ranked Chinese model on that index. It posted 90.2 on Arena-Hard v2 and 72.5 on SWE-Bench Verified. Qwen3.7 Max supports the Anthropic API protocol natively, so it integrates cleanly with tooling like Claude Code. It is well-suited for developers building coding assistants, research agents, or any API use case requiring extended reasoning over large contexts.
ChatQwen 3.5 Flash
qwen/qwen3.5-flash
Qwen 3.5 Flash is the fast, low-cost tier of the Qwen 3.5 family from Alibaba's Qwen team, served as the hosted API version of the Qwen3.5-35B-A3B model. It uses a hybrid architecture that pairs linear attention with sparse mixture-of-experts routing, which keeps compute close to linear across its 1M token context window. The model supports tool and function calling and accepts image input alongside text, with output up to 64K tokens per request. At $0.10 per million input tokens and $0.40 per million output tokens, it is priced for high-volume work such as classification, extraction, and lightweight agent loops where cost and context length matter more than peak reasoning. Qwen has not published benchmark scores for the Flash tier specifically. This rolling route tracks the latest snapshot; the dated 02-23 snapshot is listed separately.
ChatQwen3-VL Flash
qwen/qwen3-vl-flash
Qwen3-VL Flash is the fast, low-cost hosted vision-language model in Alibaba Cloud's Qwen3-VL series, served on Model Studio as the cheaper tier below Qwen3-VL Plus. It accepts text, image, and video input (video files up to one hour or 2 GB per request) and covers the Qwen3-VL task set: OCR in 32 languages, document and chart parsing, spatial reasoning, and GUI-based visual agent tasks. Alibaba reports it beats the smaller Qwen3-VL 30B A3B in both accuracy and response speed. Like Qwen3-VL Plus, it is a hybrid thinking model, answering directly by default with an optional thinking mode via the enable_thinking parameter. With a 256K context window, 32K max output, and function calling support, it fits high-volume visual workloads (screenshot parsing, receipt OCR, video tagging) where per-request cost matters more than peak accuracy.
ChatQwen3 Omni 30B A3B Instruct
qwen/qwen3-omni-30b-a3b-instruct
Qwen3 Omni 30B A3B Instruct is a natively multimodal Mixture-of-Experts model from Alibaba's Qwen team, with 30 billion total parameters and about 3 billion active per token. It accepts text, image, audio, and video input, and uses a Thinker-Talker design that can stream speech as well as text, with a reported end-to-end first-packet latency of 234 ms. The Qwen team reports state-of-the-art results on 22 of 36 audio and audio-visual benchmarks, and open-source state-of-the-art on 32, with speech recognition and voice conversation performance comparable to Gemini 2.5 Pro and ahead of GPT-4o-Transcribe. It covers 119 text languages, 19 speech input languages, and 10 speech output languages. This Instruct variant answers directly without chain-of-thought (a separate Thinking variant does extended reasoning). It supports function calling and fits voice assistants, transcription, audio analysis, and multimodal chat.
ChatQwen3 Omni 30B A3B Thinking
qwen/qwen3-omni-30b-a3b-thinking
Qwen3 Omni 30B A3B Thinking is the chain-of-thought reasoning variant of Qwen3-Omni, a natively multimodal Mixture-of-Experts model from Alibaba's Qwen team with 30 billion total parameters and about 3 billion active per token. It accepts text, image, audio, and video input and understands text in 119 languages and speech in 19. Unlike the Instruct variant, which can stream speech replies, the Thinking variant contains only the Thinker component and outputs text, producing explicit reasoning before the answer. Reported scores include 73.7 on AIME25, 73.1 on GPQA, and 88.8 on MMLU-Redux for text reasoning, 62.9 on MathVision and 69.7 on Video-MME for vision, and 75.8 on DailyOmni for audio-visual understanding. It fits harder cross-modal tasks, such as math problems presented in images or reasoning over audio and video content.
ChatTongyi DeepResearch 30B A3B
qwen/tongyi-deepresearch-30b-a3b
Tongyi DeepResearch 30B A3B is an agentic deep-research model from Alibaba's Tongyi lab, built on Qwen3-30B-A3B with a mixture-of-experts design that activates 3.3B of its 30.5B parameters per token. It is trained for long-horizon web research, multi-step information seeking, and report synthesis rather than general chat. It scores 32.9 on Humanity's Last Exam, 43.4 on BrowseComp, 46.7 on BrowseComp-ZH, 75.0 on xbench-DeepSearch, and 90.6 on FRAMES, and the Tongyi team reports performance on par with OpenAI's Deep Research across these benchmarks. Training combines agentic continual pre-training, supervised fine-tuning, and reinforcement learning on fully synthetic data. It runs in two inference modes, a native ReAct loop that needs no prompt engineering and a Heavy mode based on the IterResearch paradigm, which rebuilds a streamlined workspace each research round and can run agents in parallel. A fit for developers building autonomous research agents and deep search pipelines.
ChatQwen3-Omni Flash
qwen/qwen3-omni-flash
Qwen3-Omni Flash is a fast, cost-efficient omni-modal model from Alibaba's Qwen3 series, designed for real-time multimodal applications. As a member of the Qwen3-Omni family, it ingests text, images, audio, and video in a single end-to-end architecture — no separate pipelines or modality-switching required. It produces text responses and supports low-latency streaming, making it well suited for voice assistants, live audio/video analysis, and cost-sensitive production workloads. The Flash tier prioritizes speed and throughput over the maximum capability of the full Qwen3-Omni model, with a 65K context window and 16K output limit optimized for shorter media clips and high-volume inference. Developers building real-time assistants, transcription tools, or multimodal agents who need broad input coverage at a lower cost point will find it a practical choice.
ChatQwen3 235B A22B Instruct 2507 FP8 Throughput
qwen/qwen3-235b-a22b-instruct-2507-tput
Qwen3 235B A22B Instruct 2507 FP8 Throughput is a multilingual mixture-of-experts model from Alibaba's Qwen team, with 235B total parameters and 22B active per token. This non-thinking, instruction-tuned variant is tuned for fast, direct responses across reasoning, math, coding, and general knowledge. It posts strong published benchmarks, including 77.5% on GPQA, 70.3% on AIME25, 51.8% on LiveCodeBench v6, and 79.2% on Arena-Hard v2, with results that rival or exceed peers like GPT-4o and DeepSeek-V3 on math and coding tasks. With a 256K-token context window and solid tool-calling support, it's a strong, cost-effective pick for developers building agentic, long-context, and multilingual applications.
ChatQwen3 Coder 480B A35B Instruct Fp8
qwen/qwen3-coder-480b-a35b-instruct-fp8
Qwen3 Coder 480B A35B Instruct is a code-specialized Mixture-of-Experts model from Alibaba's Qwen team, with 480B total parameters and 35B active per token. It is purpose-built for agentic coding, where the model plans, calls tools, and works across multi-step developer workflows. It natively handles a 256K-token context, making it well suited to large files, sprawling repositories, and long-context code understanding. The model supports function calling and tool use out of the box. On SWE-bench Verified it scores 69.6%, state-of-the-art among open models and comparable to Claude Sonnet 4 on agentic coding, browser-use, and tool-use tasks. Choose it for autonomous coding agents, refactoring, debugging, and repository-scale tasks where strong agentic performance and a long context matter.
ChatQwen-MT Flash
qwen/qwen-mt-flash
Qwen-MT Flash is a machine translation model from Alibaba's Qwen team, built on the Qwen3 architecture and positioned as the balanced tier of the Qwen-MT line for quality, speed, and cost. It translates between 92 languages, including Chinese, English, Japanese, Korean, French, Spanish, German, Arabic, Thai, Indonesian, and Vietnamese. Like the other Qwen-MT models, it supports terminology intervention (custom glossaries for brand names and technical terms), domain prompting, and translation memory for consistent output across large documents. Unlike Qwen-MT Plus and Turbo, it also supports incremental streaming, where each response contains only newly generated text. Alibaba's documentation recommends it for general translation scenarios such as website content, product descriptions, and everyday communication, with Qwen-MT Plus reserved for professional domains where output quality matters most.
ChatQwen-MT Lite
qwen/qwen-mt-lite
Qwen-MT Lite is the lowest-cost tier of Alibaba's Qwen-MT machine translation family, built for simple, latency-sensitive translation such as real-time chat and live comment translation. It supports 31 languages, fewer than the 92 covered by Qwen-MT Plus and Qwen-MT Turbo, and Alibaba's model comparison rates it as basic quality at the fastest speed and lowest cost of the tiers. Unlike the higher tiers, it does not support domain prompting. It does support incremental streaming output, where each response contains only newly generated content. Best suited for developers translating short, high-volume text (chat messages, comments, UI strings) where response speed and price matter more than translation fidelity. For terminology control or higher output quality, Qwen-MT Plus or Turbo are the better fit.
ChatQwen3 1.7B
qwen/qwen3-1.7b
Qwen3 1.7B is a small dense chat model from Alibaba's Qwen team, part of the Qwen3 family released in April 2025. It has 1.7 billion parameters (1.4 billion non-embedding) and a 32,768-token context window. Like the rest of the family, it supports hybrid operation, a thinking mode for math, coding, and multi-step reasoning and a non-thinking mode for fast general dialogue, switchable per request via the enable_thinking flag. The Qwen team reports that its small Qwen3 models outperform larger Qwen2.5 instruct models on math, code generation, and commonsense logical reasoning. The model supports tool calling for agent workflows and covers over 100 languages and dialects. This route allows up to 8,192 output tokens. At these prices it fits high-volume tasks such as classification, routing, extraction, and simple chat.
ChatQwen3 0.6B
qwen/qwen3-0.6b
Qwen3 0.6B is the smallest dense model in Alibaba's Qwen3 family, released in April 2025 with 0.6 billion parameters (0.44 billion non-embedding). Like the larger Qwen3 models, it supports hybrid thinking modes, switching between step-by-step reasoning and a faster direct-answer mode per request, including with /think and /no_think tags in prompts. The Qwen3 technical report lists 55.6 on MMLU-Redux and 59.6 on GSM8K in thinking mode, ahead of Qwen2.5-0.5B-Instruct (41.6 on GSM8K). The family was pretrained on 36 trillion tokens covering 119 languages, and the model supports tool calling, with the Qwen-Agent framework recommended for agent workloads. At this size it fits high-volume, cost-sensitive tasks such as classification, routing, and drafting, and it is also used as a draft model for speculative decoding with larger Qwen3 models. Complex multi-step math stays well below what 1B-plus models manage.
ChatQwen3-VL 30B-A3B
qwen/qwen3-vl-30b-a3b
Qwen3-VL 30B-A3B is a compact mixture-of-experts vision-language model from Alibaba's Qwen team, with 30B total parameters and only 3B active per token for efficient inference. It supports image and text inputs with a 131K context window and delivers strong multimodal performance on benchmarks including MMMU and visual-math evaluations. Capabilities include document and chart understanding, OCR, visual coding (generating HTML/CSS/JS from images), 2D spatial grounding, and GUI agent tasks across desktop and mobile interfaces. The MoE architecture gives it the knowledge breadth of a much larger model while matching the latency and cost profile of a 3B dense model — making it a practical choice for developers who need reliable vision-language capabilities without the compute cost of the 235B flagship variant. Supports tool calling.
ChatQwQ Plus
qwen/qwq-plus
QwQ Plus is a proprietary reasoning model from Alibaba's Qwen team, serving as the hosted API counterpart to the open-weight QwQ-32B release. Like QwQ-32B, it uses reinforcement learning to develop extended chain-of-thought reasoning, excelling at math competition problems, scientific reasoning, and complex coding tasks. QwQ-32B achieved 79.5% on AIME 2024, 90.6% on MATH-500, and 63.4% on LiveCodeBench — rivaling much larger models. QwQ Plus exposes these capabilities through a managed API endpoint with a 131K token context window and tool call support. Best suited for developers building applications that require step-by-step mathematical reasoning, algorithmic problem-solving, or multi-step logical inference.
ChatQwen2.5 7B Instruct 1M
qwen/qwen2.5-7b-instruct-1m
Qwen2.5 7B Instruct 1M is a 7-billion-parameter instruction-tuned chat model from Alibaba's Qwen team, released in January 2025 as the long-context variant of Qwen2.5 7B Instruct in the Qwen2.5-1M series. It handles contexts up to 1 million tokens, enough to fit entire codebases, books, or large document collections in a single request. In Qwen's 1M-token passkey retrieval test it finds hidden information with near-perfect accuracy (minor errors reported for the 7B size), and on RULER, LV-Eval, and LongBench-Chat it outperforms its 128K predecessor, especially on sequences beyond 64K tokens. On short-context academic benchmarks it performs about the same as the 128K version, which Qwen reports as comparable to GPT-4o-mini, while supporting a context eight times longer. A low-cost option for long-document question answering, retrieval, and analysis over very large inputs.
ChatQwen-Omni Turbo
qwen/qwen-omni-turbo
Qwen-Omni Turbo is Alibaba's cost-optimized omnimodal API model, built to process text, image, audio, and video inputs and return text responses in a single unified interface. It is the lighter, faster tier in the Qwen-Omni family, designed for developers who need full multimodal coverage at lower latency and cost than the flagship Qwen-Omni model. Audio files up to 40 seconds and video files up to 150 MB are supported, spanning common formats such as MP3, WAV, MP4, and MOV. The model handles tool calling natively and is accessible via an OpenAI-compatible API. Best suited for developers building applications that need to reason across mixed media inputs — such as audio transcription pipelines, video understanding workflows, or multimodal chatbots — where throughput and cost efficiency matter.
ChatQwen-MT Plus
qwen/qwen-mt-plus
Qwen-MT Plus is a specialized machine translation model from Alibaba's Qwen team, purpose-built for high-quality text translation across 92 languages covering over 95% of the world's population. Unlike general-purpose language models, Qwen-MT Plus is fine-tuned specifically for translation tasks, offering term intervention, domain prompting, and translation memory features that give developers fine-grained control over output. It supports translation between major languages including Chinese, English, Japanese, Korean, French, Spanish, German, Arabic, Thai, Indonesian, and Vietnamese. Best suited for developers building multilingual applications, content localization pipelines, or customer-facing translation features where accuracy, terminology consistency, and domain fidelity matter more than general conversational ability.
ChatQwen-MT Turbo
qwen/qwen-mt-turbo
Qwen-MT Turbo is a fast, cost-effective machine translation model from Alibaba's Qwen team, designed for high-volume text translation across 92 languages. As the Turbo tier of the Qwen-MT family, it trades some of the output fidelity of Qwen-MT Plus for significantly lower cost and faster throughput — making it the practical choice for latency-sensitive or budget-constrained translation workflows. Like its sibling, it supports term intervention, domain prompting, and translation memory, giving developers control over terminology and style. Best suited for developers building high-volume localization pipelines, real-time translation features, or cost-sensitive multilingual applications where speed and price efficiency matter more than maximum output quality.
ChatQwen Plus Character
qwen/qwen-plus-character
Qwen Plus Character is the role-playing model in Alibaba's Qwen series, offered through Alibaba Cloud Model Studio with an OpenAI-compatible chat API. It is tuned for staying in character across a conversation. Alibaba's documentation highlights following predefined character instructions, advancing the dialogue, and showing active listening and empathy, with support for detailed persona definitions set through the system prompt. Compared with the general-purpose Qwen Plus, it trades broad task coverage for character consistency and natural conversational flow. It handles text input and output only and does not support function calling, so it is not a fit for tool-using agents. Alibaba lists virtual social apps, game NPCs, IP character replication, and smart hardware assistants as target use cases. A Japanese-tuned variant (qwen-plus-character-ja) is also available.
ChatQwen2.5-Omni 7B
qwen/qwen2-5-omni-7b
Qwen2.5-Omni 7B is Alibaba's end-to-end omni-modal model capable of perceiving text, images, audio, and video simultaneously while generating text and natural speech in real time. Built on a Thinker-Talker architecture with TMRoPE (Time-aligned Multimodal RoPE) for synchronizing audio and video streams, the 7B model achieves strong benchmark results across all modalities. It ranked first on the MMAU audio understanding leaderboard, scored 59.2 on MMMU image reasoning (near GPT-4o-mini's 60.0), and achieved 64.3 on Video-MME for video understanding without subtitles. On OmniBench, which tests cross-modal integration, it reached 56.13%. The model supports tool/function calling and targets developers building voice assistants, video analysis tools, and multimodal pipelines that require a single model to handle diverse input types.
ChatQwen2.5 Coder 32B Instruct
qwen/qwen2.5-coder-32b-instruct
Qwen2.5 Coder 32B Instruct is the largest model in Alibaba's Qwen2.5-Coder family, a code-specialized model released in November 2024 and trained for code generation, code repair, and code reasoning. On code generation it scores 92.7 on HumanEval and 86.3 on EvalPlus, and the Qwen team reports overall coding performance competitive with GPT-4o. On Aider's code repair benchmark it scores 73.7, comparable to GPT-4o, and it scores 75.2 on MdEval, a multi-language code repair benchmark, first among open models at release. It covers more than 40 programming languages, scoring 65.9 on McEval, with notably strong results in less common languages like Haskell and Racket. This route offers a 32K context window at low per-token cost. It fits developers who want code completion, bug fixing, and code explanation without paying frontier-model prices.
ChatQwen2.5 7B Instruct Turbo
qwen/qwen2.5-7b-instruct-turbo
Qwen2.5 7B Instruct Turbo is a 7-billion-parameter instruction-tuned chat model from Alibaba's Qwen team, served as a fast, low-cost Turbo endpoint via Together AI. It is known for strong coding and math performance relative to its size, scoring 84.8 on HumanEval and 75.5 on MATH, and it excels at instruction following, structured-data understanding, and reliable JSON output. The model supports tool/function calling and handles 29+ languages. Published comparisons show it outperforming similarly sized open models such as Llama 3.1 8B Instruct and Gemma 2 9B across most tasks. Choose it when you want a cheap, capable small model for coding assistants, structured extraction, multilingual chat, and high-throughput API workloads.
ChatQwen2.5 32B Instruct
qwen/qwen2.5-32b-instruct
Qwen2.5 32B Instruct is a 32-billion-parameter instruction-tuned chat model from Alibaba's Qwen team, released in September 2024 as part of the Qwen2.5 family, which was pretrained on 18 trillion tokens. On instruct-model benchmarks it scores 83.1 on MATH, 88.4 on HumanEval, 84.0 on MBPP, 69.0 on MMLU-Pro, and 74.5 on Arena-Hard. In Qwen's published comparison it leads GPT-4o mini and Gemma2 27B IT on most of these tasks, with math and coding as its strongest areas. The model supports function calling, structured JSON output, more than 29 languages, a 131,072-token context window, and up to 8,192 output tokens. It fits developers who want stronger reasoning, math, and coding than the 14B variant without paying for the 72B flagship.
ChatQwen2.5 14B Instruct
qwen/qwen2-5-14b-instruct
Qwen2.5 14B Instruct is a 14.7-billion-parameter open-weight model from Alibaba's Qwen team, trained on 18 trillion tokens and released under Apache 2.0. It hits a practical sweet spot in the Qwen2.5 lineup — outperforming both the 7B variant and models like Gemma 2 27B and GPT-4o mini on seven key benchmarks, while remaining far more efficient than the flagship 72B. Core strengths include strong instruction following, structured output (JSON) generation, math, and code. It reaches ~97% tool-call success across hardware, making it reliable for agentic workflows. Multilingual support spans 29+ languages with a 128K context window and up to 8K output tokens. A strong choice for developers who need GPT-4o-mini-class quality at a fraction of the cost of larger frontier models.
ChatQwen2.5 32B Instruct
qwen/qwen2-5-32b-instruct
Qwen2.5 32B Instruct is a general-purpose language model from Alibaba's Qwen team, sitting at the practical sweet spot between the 14B and 72B variants in the Qwen2.5 series — delivering stronger reasoning and language understanding than the 14B while remaining far more cost-efficient than the 72B. Trained on 18 trillion tokens, the model scores 57.7 on MATH and outperforms Qwen2-72B on comprehensive evaluations despite having fewer parameters. It excels at instruction following, multi-step reasoning, mathematics, coding assistance, and multilingual tasks across 29+ languages, with a 131K token context window and full tool-call support. A well-rounded choice for developers who need reliable general-purpose performance — complex enough for demanding workflows, light enough to keep inference costs manageable.
ChatQwen2.5-VL 72B Instruct
qwen/qwen2-5-vl-72b-instruct
Qwen2.5-VL 72B Instruct is Alibaba's flagship open-source vision-language model, matching state-of-the-art closed models like GPT-4o and Claude 3.5 Sonnet on multimodal tasks. The model excels at document understanding (96.4 on DocVQA), OCR (88.8 on OCRBench), and structured data extraction from invoices, forms, tables, and charts. On MMMU it scores 70.2, and across 21 benchmarks it outperforms Gemini 2.0 Flash, GPT-4o, and Claude 3.5 Sonnet on 13 of them. Video understanding extends to over one hour of footage with second-level event pinpointing, enabled by dynamic FPS sampling and absolute time encoding. The model also functions as a visual agent capable of computer and phone use. A strong choice for developers building document pipelines, OCR workflows, visual Q&A systems, or multimodal agents.
ChatQwen3.5 27B Anko
qwen/qwen3.5-27b-anko
Qwen3.5 27B Anko is a community fine-tune of Qwen3.5 27B, built by Allura (allura-org) on top of ArliAI's derestricted version of the base model, which removes built-in refusal behavior via a weight-abliteration technique. Anko adds a LoRA trained on reasoning traces and responses, using data generated by Doubao Seed 2.0 Pro and Mini, aimed at improving coherence and cutting down on repetitive output. It's distributed through the Infron aggregator and positioned for creative writing, roleplay, and general chat. The model card recommends a Claude-style system prompt and non-default sampling (temperature around 1.25 with min_p) rather than Qwen's stock settings. Context is 262K tokens. No independent benchmark scores have been published for this variant.
ChatQwen3.5 27B InfraCelestial
qwen/qwen3.5-27b-infracelestial
Qwen3.5 27B InfraCelestial is a community fine-tune of Qwen3.5 27B distributed through the Infron aggregator. It follows the naming and pricing pattern of other Qwen3.5 27B tunes on the same platform (a 262K context window and per-token pricing matching the "derestricted" family of roleplay- and creative-writing-oriented fine-tunes), which suggests a similar orientation toward uncensored creative and roleplay use. No model card, training details, base-model lineage beyond Qwen3.5 27B, or creator attribution could be found for this specific model. It does not appear on Hugging Face or other model hubs under this name at the time of writing. Context is 262K tokens, and no independent benchmark scores have been published for this model.
ChatQwen3.5 27B RPRMax v1
qwen/qwen3.5-27b-rprmax-v1
Qwen3.5 27B RPRMax v1 is a community fine-tune of Qwen3.5 27B distributed through the Infron aggregator. Its name and pricing follow the same pattern as other roleplay-oriented Qwen3.5 27B tunes on the platform (a 262K context window and per-token cost matching the "derestricted" fine-tune family), suggesting a similar orientation toward roleplay or creative-writing use, but no model card or listing description could be found to confirm training details or intended use. No creator attribution, base-lineage confirmation beyond Qwen3.5 27B, or third-party coverage could be located for this model at the time of writing. Context is 262K tokens, and no independent benchmark scores have been published for it.
ChatQwen3.8 Max Prime
qwen/qwen3.8-max-prime
Qwen3.8 Max Prime is a high-throughput serving tier of Alibaba's Qwen3.8 Max, the Qwen team's 2.4-trillion-parameter Mixture-of-Experts flagship released August 3, 2026. It runs the same underlying weights as standard Qwen3.8 Max, a 1M-token-context model with 95B active parameters and native text, image, and video input, but on faster inference infrastructure. Third-party trackers report roughly 1.5 to 2 times the output throughput of the standard endpoint, at a higher per-token price. Because the weights are unchanged, Prime inherits Qwen3.8 Max's benchmark results, including 67.7 on SWE-bench Pro and a top-three ranking on LMArena Code, putting it behind Claude Fable 5 but ahead of GPT-6 Sol on several agentic coding evaluations. It suits developers running latency-sensitive coding agents or high-concurrency workloads who want Qwen3.8 Max's capability without the standard endpoint's slower response time.
ChatQwen3.8 OmniFlash
qwen/qwen3.8-omni-flash
Qwen3.8 OmniFlash is Alibaba's first omni-modal model built around agentic capabilities, released September 18, 2026, succeeding Qwen3.5-Omni-Plus. It accepts text, images, audio, and video as input and returns text, pairing native audio-video understanding with reasoning and tool use. It can watch or listen to content, plan a task, call tools, and deliver a finished result, such as an edited video or a meeting summary. Alibaba reports more than a 26% average improvement across 30 evaluations versus Qwen3.5-Omni-Plus, with gains in audio-video agents, coding, long-context tasks, and real-time interaction. It says audio-visual performance approaches Gemini 3.8 Flash, with audio performance exceeding it. With a 1,000,000 token context window, 128,000 token output limit, and pricing of $0.08 per million input tokens and $0.24 per million output tokens, it suits developers building video and audio agents, such as meeting summarization or video-editing pipelines that need tool-calling built in.
ChatQwen3.8 Max (0902)
qwen/qwen3.8-max-0902
Qwen3.8 Max (0902) is a September 2026 snapshot of Alibaba's Qwen3.8 Max, a 2.4-trillion-parameter mixture-of-experts flagship with roughly 95 billion active parameters per token. It keeps the 1M-token context window, thinking mode, and tool ecosystem of the base model, with additional post-training on coding and office-work tasks. On Code Arena WebDev it ranks first overall at 1,691 points, ahead of Claude Opus 5 Max (1,687) and Kimi K3 Max (1,674). The largest gains are on TerminalBench 3.0 and ProgramBench Almost Solved, both more than doubling versus the prior snapshot. Reported core scores include GPQA Diamond 92.6 and PaperBench 93.0. Against Claude Opus 5, it leads on several coding benchmarks and on WorkArena, though Opus 5 still leads on most agentic coding rows and both office-work benchmarks. Pricing is unchanged from the base Qwen3.8 Max, making it a straightforward upgrade for coding and office-automation workloads already on the platform.
ChatQwen3.8 Flash
qwen/qwen3.8-flash
Qwen3.8 Flash is a multimodal model from Alibaba's Qwen team, released August 26, 2026, as the fast, lower-cost tier of the Qwen3.8 family alongside Qwen3.8 Max and Qwen3.8 27B. It uses a mixture-of-experts architecture with 125B total parameters and 6B active per token, an early preview of the architecture planned for Qwen4. It accepts text, image, and video input and returns text, with a 1,000,000 token context window and output capped at 128,000 tokens. The API supports tool calling, structured outputs via JSON schema, and prompt caching, with cached input billed at $0.016 per million tokens. Pricing is $0.08 per million input tokens and $0.24 per million output tokens, about one-twelfth the cost of Qwen3.8 Max. Alibaba says it was trained at roughly one-ninth the cost of Qwen3.7-Plus and reports higher scores on benchmarks including SWE-bench Pro and CoWorkBench, an agentic office-task benchmark.
ChatQwen3.8 27B
qwen/qwen3.8-27b
Qwen3.8 27B is a dense, open-weight multimodal model from Alibaba's Qwen team, released August 14, 2026 as a smaller member of the Qwen3.8 family alongside the flagship Qwen3.8 Max. It combines Gated DeltaNet linear attention with standard gated attention across 64 layers, giving a 27 billion parameter dense model a native 262K token context window, extendable to 1M tokens. It accepts text, image, and video input, including hour-scale video and STEM diagrams. Alibaba reports 61.7 on SWE-bench Pro and 73.0 on Terminal Bench 2.1, both improvements over the earlier Qwen3.6 27B, and 89.2 on GPQA Diamond. Released under Apache 2.0, it gives developers an open-weight alternative to Qwen3.8 Max for coding and agentic tasks, at a fraction of the parameter count.
ChatQwen3.8 2.4T A95B
qwen/qwen3.8-2.4t-a95b
Qwen3.8 2.4T A95B is Alibaba's open-weight release of its Qwen3.8 Max flagship, a sparse mixture-of-experts model with 2.4 trillion total parameters and 95 billion active per token, routed across 512 experts. It uses a hybrid attention design (Gated DeltaNet and Gated Attention layers) across 92 layers, with a native 262K context window and thinking mode enabled for every response. Alibaba reports 93.0 on PaperBench (ahead of GPT-5.6 Sol's 90.5), 92.6 on GPQA Diamond, 86.6 on Terminal-Bench 2.1, and 67.7 on SWE-bench Pro, positioning it for coding, research, and long-horizon agentic work. It gives developers access to Qwen-Max-class capability under open weights, useful for teams that want frontier-level coding and agentic performance without a closed API.
ChatQwen3.8 Max
qwen/qwen3.8-max
Qwen3.8 Max is Alibaba's flagship large language model, released August 3, 2026 as the most capable model in the Qwen family to date. It uses a mixture-of-experts architecture with 2.4 trillion total parameters and about 95 billion active per request, and accepts text, image, and video input with a context window of up to 1 million tokens. Alibaba positions it for coding and long-horizon agentic work: in testing the model ran autonomously for over 10 days building a self-evolving software harness. Reported benchmarks include 93.0 on PaperBench, 82.8 on IFBench, 86.6 on Terminal-Bench 2.1, and 86.1 on OSWorld-Verified, ahead of Claude Opus 4.8 on several coding and agent tasks and roughly matching Claude Fable 5 and GPT-5.6 Sol, though it trails both on some evaluations. On the Arena.AI leaderboard it ranks as the top Chinese model for text tasks. Alibaba plans to open-source the weights on Hugging Face and ModelScope.
ChatQwen3.7 Max (2026-06-08)
qwen/qwen3.7-max-2026-06-08
Qwen3.7 Max (2026-06-08) is a dated snapshot of Alibaba's flagship reasoning model, captured about three weeks after the initial 2026-05-20 release build. Qwen3.7 Max is built for long-horizon agentic workloads, pairing a 1 million-token context window with a chain-of-thought reasoning architecture. Alibaba demonstrated it running for 35 hours without degradation across more than 1,000 tool calls in a single session. It ranked 13th globally on LM Arena's text leaderboard and scored 56.6 on the Artificial Analysis Intelligence Index, the highest of any Chinese model on that index, alongside 90.2 on Arena-Hard v2 and 72.5 on SWE-Bench Verified. Use this snapshot when you want the same reasoning and agentic capabilities as the 2026-05-20 build but pinned to Alibaba's later update, so API calls keep returning consistent results as the rolling qwen/qwen3.7-max alias continues to move forward.
ChatQwen3.7 Max (2026-05-20)
qwen/qwen3.7-max-2026-05-20
Qwen3.7 Max (2026-05-20) is a dated snapshot of Alibaba's flagship reasoning model, pinned to its initial May 2026 release build. Qwen3.7 Max is built for long-horizon agentic workloads, pairing a 1 million-token context window with a chain-of-thought reasoning architecture. Alibaba demonstrated it running for 35 hours without degradation across more than 1,000 tool calls in a single session. It ranked 13th globally on LM Arena's text leaderboard and scored 56.6 on the Artificial Analysis Intelligence Index, the highest of any Chinese model on that index, alongside 90.2 on Arena-Hard v2 and 72.5 on SWE-Bench Verified. Pinning this snapshot instead of the rolling qwen/qwen3.7-max alias gives reproducible behavior for production use. It supports tool calling and the Anthropic API protocol, fitting coding agents, automated pipelines, and long-context document analysis where output has to stay consistent across model updates.
ChatQwen3.5 Plus (2026-04-20)
qwen/qwen3.5-plus-2026-04-20
Qwen3.5 Plus (2026-04-20) is Alibaba's hosted flagship chat model in the Qwen3.5 series, built on the Qwen3.5-397B-A17B architecture, a hybrid design that combines linear attention with a sparse mixture-of-experts layer, activating 17 billion of 397 billion total parameters per token. It takes text, image, and video input within a 1-million-token context window, and supports native tool calling for agentic workflows such as web search and code execution. This makes it a fit for developers processing long documents, large codebases, or multi-turn agent sessions in a single request. On Puter.js, this April 2026 snapshot is priced at $0.40 per million input tokens and $2.40 per million output tokens, higher than other Qwen3.5 Plus snapshots listed here, and it adds video as an input modality that those don't support. Accessible through an OpenAI-compatible API, it suits developers who want a recent, version-locked build of Qwen3.5 Plus for reproducible production behavior.
ChatQwen3.5-Omni Flash
qwen/qwen3.5-omni-flash
Qwen3.5-Omni Flash is the fast, cost-efficient tier of Alibaba's Qwen3.5-Omni family, a natively multimodal model that takes text, images, video, and audio as input in a single end-to-end architecture and returns text. Both Flash and Plus share a Thinker-Talker design built on a Hybrid-Attention Mixture-of-Experts architecture, an upgrade over the prior Qwen3-Omni generation, with an ARIA module aligning text and speech tokens for streaming synthesis; this API route returns text only. Alibaba's technical report shows Flash trading accuracy for speed against Plus: 79.9 vs 85.9 on MMLU-Pro, 81.9 vs 86.8 on MLVU video understanding, and a higher LibriSpeech word error rate (1.30-2.43 vs 1.11-2.23), while answering faster with a first-packet audio latency of 235ms versus 435ms. With a 48K token context, tool calling, and $0.40/$2.20 per million input/output tokens, it fits voice assistants, live video analysis, and high-volume multimodal apps where latency and cost matter more than peak accuracy.
ChatQwen3.5-Omni Plus
qwen/qwen3.5-omni-plus
Qwen3.5-Omni Plus is the flagship tier of Alibaba's Qwen3.5-Omni family, a natively multimodal model taking text, images, video, and audio as input in one end-to-end architecture and returning text. It shares Flash's Thinker-Talker design, now built on a Hybrid-Attention Mixture-of-Experts architecture, an upgrade over the prior Qwen3-Omni generation. Alibaba reports it reaches state-of-the-art results across 215 audio and audio-visual benchmarks, and says it surpasses Gemini 3.1 Pro on audio understanding while matching it on audio-visual comprehension. Compared with Flash, Plus scores higher on MMLU-Pro (85.9 vs 79.9) and MLVU video understanding (86.8 vs 81.9), and posts a lower LibriSpeech word error rate (1.11-2.23 vs 1.30-2.43), at the cost of higher latency (435ms vs 235ms first-packet audio). With a context window approaching one million tokens, tool calling, and speech recognition spanning 113 languages and dialects, it fits long-context multimodal pipelines, transcription, and translation work where accuracy matters more than raw speed.
ChatQwen3.5 Plus (2026-02-15)
qwen/qwen3.5-plus-2026-02-15
Qwen3.5 Plus (2026-02-15) is Alibaba's hosted flagship chat model in the Qwen3.5 series, built on the Qwen3.5-397B-A17B architecture, a hybrid design that combines linear attention with a sparse mixture-of-experts layer, activating 17 billion of 397 billion total parameters per token. It takes text, image, and video input within a 1-million-token context window, and supports native tool calling for agentic workflows such as web search and code execution. This makes it a fit for developers processing long documents, large codebases, or multi-turn agent sessions in a single request. On Puter.js, this dated snapshot is priced at $0.40 per million input tokens and $2.40 per million output tokens, higher than other Qwen3.5 Plus snapshots listed here, and it adds video as an input modality that those don't support. Accessible through an OpenAI-compatible API, it's a good choice when reproducible, version-pinned behavior matters more than using the latest rolling release.
ChatQwen3 Max (2026-01-23)
qwen/qwen3-max-2026-01-23
Qwen3 Max (2026-01-23) is Alibaba's flagship proprietary large language model, a later dated snapshot of the same trillion-parameter-class Mixture-of-Experts line launched in September 2025 and served only through the API. Compared with the September 2025 snapshot, this version adds hybrid thinking, letting a request run in fast non-thinking mode (the default) or switch to a slower reasoning mode via a /think suffix. In thinking mode it can call web search, web-page extraction, and a code interpreter mid-reasoning, which Alibaba says improves accuracy on complex, multi-step problems. That combination of a large context window, tool calling, and optional deliberate reasoning fits agentic coding, research assistants, and workflows that need to switch between quick answers and careful, tool-augmented problem solving in the same deployment.
ChatQwen3-VL Flash (2026-01-22)
qwen/qwen3-vl-flash-2026-01-22
Qwen3-VL Flash (2026-01-22) is Alibaba's fast, low-cost vision-language model, a dated snapshot of the Qwen3-VL Flash series pinned to its January 22, 2026 release. It covers the same Qwen3-VL task set as the Plus tier, including OCR, document and chart parsing, spatial reasoning, long video understanding, and GUI-based visual agent tasks, with hybrid thinking and non-thinking modes toggled per request. At $0.05 per million input tokens and $0.40 per million output tokens, it costs a quarter of Qwen3-VL Plus on both input and output, with the same 262,144 token context and 32,768 max output tokens. It fits high-volume visual workloads such as screenshot parsing, receipt OCR, and video tagging, where per-request cost matters more than peak accuracy, and where a fixed snapshot is preferable to a rolling alias.
ChatQwen3-VL Plus (2025-12-19)
qwen/qwen3-vl-plus-2025-12-19
Qwen3-VL Plus (2025-12-19) is Alibaba's hosted vision-language model, a snapshot of the Qwen3-VL Plus series pinned to its December 19, 2025 update, roughly three months after the original September 2025 release. Like other Qwen3-VL Plus versions, it covers document and chart parsing, OCR, spatial reasoning, and long video understanding, with hybrid thinking and non-thinking modes and visual-agent support for GUI tasks on desktop and mobile. It keeps the same 262,144 token context window, 32,768 max output tokens, and $0.20/$1.60 per million token pricing as the September snapshot. Choose this dated version over the rolling alias when you need a fixed model behind your API calls, and prefer it over the September snapshot if you want the most recent Plus-tier weights.
ChatQwen Plus (2025-12-01)
qwen/qwen-plus-2025-12-01
Qwen Plus (2025-12-01) is Alibaba's most recent dated snapshot of Qwen Plus among the pins covered here, following the September 2025 update in the same Qwen3-series line. Alibaba's release notes describe improved reasoning over the July 2025 snapshot, along with further enhanced agent capabilities and multi-turn tool invocation, plus a noted improvement in subjective, creative-writing tasks. With tool calling support and a 1 million token context window, it fits agentic workflows that call functions across many turns as well as tasks that mix long input documents with open-ended writing. As a version-pinned snapshot served through Alibaba Cloud Model Studio, it keeps behavior fixed for teams that need reproducibility rather than the rolling qwen-plus alias's automatic updates.
ChatQwen3-Omni Flash (2025-12-01)
qwen/qwen3-omni-flash-2025-12-01
Qwen3-Omni Flash (2025-12-01) is a dated snapshot of Alibaba's fast, cost-efficient omni-modal model, capturing what Alibaba calls a "massive upgrade" to the Flash tier released December 1, 2025. Like the rest of the Qwen3-Omni family, it ingests text, images, audio, and video in one end-to-end architecture and returns text, keeping the same 65K context window, 16K output limit, and pricing as earlier Flash snapshots. Alibaba says the update improves multi-turn audio and video conversation handling and adds support for customizing the model's personality through system prompts. For developers, the difference from the September 2025 snapshot is behavioral, not architectural. Context size, output limit, and cost stay the same; what changes is smoother multi-round voice and video dialogue and a more controllable persona, useful for voice assistants and conversational agents handling extended spoken or video exchanges.
ChatQwen3 Coder Plus (2025-09-23)
qwen/qwen3-coder-plus-2025-09-23
Qwen3 Coder Plus (2025-09-23) is Alibaba's API version of Qwen3-Coder, refreshed by the Qwen team in September 2025 to improve terminal-based agentic coding. It's built for multi-turn interaction with a development environment, planning steps, calling tools, reading back results, and adjusting rather than producing code in one shot. Qwen described this update as improving Terminal Bench performance with Qwen Code and Claude Code, and adding more secure code generation. A Qwen team member reported a SWE-Bench Verified score of 69.6 for this version. The API exposes 1,048,576 tokens of context, enough to keep a large repository or a long agent run in scope. This is the September 23, 2025 dated snapshot, pinned separately from the undated Qwen3 Coder Plus entry so its behavior stays fixed. Choose it for terminal-heavy agentic workflows with Claude Code or Qwen Code, or when you need a reproducible baseline newer than the July release.
ChatQwen3 Max (2025-09-23)
qwen/qwen3-max-2025-09-23
Qwen3 Max (2025-09-23) is Alibaba's flagship proprietary large language model, a trillion-parameter-class Mixture-of-Experts system trained on 36 trillion tokens and offered only through the API rather than as open weights. This snapshot marks Qwen's official, non-preview Qwen3 Max launch, following the earlier September 5 preview. Alibaba reported 72.5 on SWE-Bench Verified, 69.6 on Tau2-Bench Verified, 81.6 on AIME25, and 74.8 on LiveCodeBench v6, and the model reached the top three on the LMArena text leaderboard, ahead of GPT-5-Chat in that ranking. It runs non-thinking, answering directly without a separate reasoning pass, which keeps latency down for coding assistants and tool-calling agents. Pinning to this dated snapshot instead of the general qwen3-max alias keeps behavior fixed for production use, ahead of the hybrid-thinking update Alibaba shipped in the January 2026 snapshot.
ChatQwen3-VL Plus (2025-09-23)
qwen/qwen3-vl-plus-2025-09-23
Qwen3-VL Plus (2025-09-23) is Alibaba's hosted vision-language model, a dated snapshot of the Qwen3-VL series pinned to its September 23, 2025 launch. It handles document and chart parsing, OCR, spatial reasoning, and long video understanding, and supports hybrid thinking and non-thinking modes for balancing depth of reasoning against latency. It also functions as a visual agent for GUI-based tasks on desktop and mobile screens. With a 262,144 token context window and 32,768 max output tokens, it suits multi-page PDFs, long visual conversations, and tool-calling workflows through Alibaba Cloud Model Studio's OpenAI-compatible API. Pin this snapshot instead of the rolling qwen3-vl-plus alias when you need reproducible behavior across requests, such as in production pipelines where a later silent model swap could change outputs.
ChatQwen3 VL 235B A22B Instruct
qwen/qwen3-vl-235b-a22b-instruct
Qwen3 VL 235B A22B Instruct is Alibaba's flagship vision-language model, a mixture-of-experts architecture with 235 billion total parameters and 22 billion active per token, served through Alibaba's API. It's built for visual coding (generating Draw.io, HTML, and CSS from screenshots or mockups), spatial reasoning about object positions and occlusion, and GUI/PC navigation as a visual agent. OCR covers 32 languages, including rare and ancient scripts, and holds up on low-light, blurred, or tilted text. The model natively handles 256K tokens of interleaved text, image, and video, extendable to 1M tokens, and Alibaba reports near-full accuracy retention at that length, enough for hours-long video with second-level indexing. On OmniDocBench it scores 88.9 overall for document parsing. It fits developers building document-processing pipelines, coding-from-screenshot tools, or agents that need to watch and reason over long video.
ChatQwen3-Omni Flash (2025-09-15)
qwen/qwen3-omni-flash-2025-09-15
Qwen3-Omni Flash (2025-09-15) is a dated snapshot of Alibaba's fast, cost-efficient omni-modal model, capturing its state as first released alongside the broader Qwen3-Omni family on September 22, 2025. It ingests text, images, audio, and video in a single end-to-end architecture and returns text, with low-latency streaming support suited to voice assistants and live audio/video analysis. The Flash tier trades peak capability for speed and throughput, with a 65K context window and 16K output limit tuned for high-volume, cost-sensitive inference. Pinning to this snapshot locks in behavior from before Alibaba's December 2025 update, which added smoother multi-turn audio/video conversation handling and system-prompt personality customization. Choose this ID when reproducible output matters more than picking up the newest Flash improvements.
ChatQwen Plus (2025-09-11)
qwen/qwen-plus-2025-09-11
Qwen Plus (2025-09-11) is a snapshot of Alibaba's Qwen Plus model that followed the July 2025 update which first extended Qwen Plus to a 1 million token context window. Alibaba's notes for this release describe better instruction following and more streamlined summaries when the model is in thinking mode, plus stronger Chinese-language text understanding and more consistent logical reasoning in non-thinking mode. It supports tool calling and the full 1 million token context inherited from the prior snapshot, making it suited to long-document analysis and retrieval-heavy tasks where the shorter Qwen Plus versions from earlier in 2025 would truncate input. Served through Alibaba Cloud Model Studio, it can be pinned by this dated ID for reproducible behavior in production.
ChatQwen3 Max Preview
qwen/qwen3-max-preview
Qwen3 Max Preview is Alibaba's first preview release of the Qwen3 Max line, a trillion-parameter-class Mixture-of-Experts model made available through the API only, without open weights. Announced in early September 2025 as Qwen's largest model to date, it predates the dated Qwen3 Max (2025-09-23) snapshot; Alibaba described it at launch as beating the previous Qwen3-235B-A22B-2507 model on internal benchmarks, with early user feedback cited as confirming the gains. Like the later dated snapshots, it runs non-thinking by default, answers directly rather than working through an explicit chain of thought, and supports tool calling with a 262K-token context window. Choose this id specifically when you need to pin to that original preview build, for example to reproduce results measured against the initial Qwen3 Max preview rather than the improved dated releases that followed it.
ChatQwen3 Coder Plus (2025-07-22)
qwen/qwen3-coder-plus-2025-07-22
Qwen3 Coder Plus (2025-07-22) is Alibaba's hosted API version of Qwen3-Coder, the agentic coding model Qwen introduced in July 2025 for autonomous software engineering. It's built for multi-turn interaction with a development environment, planning steps, calling tools, reading back results, and adjusting rather than producing code in one shot. At launch, Qwen reported that the Qwen3-Coder family set state-of-the-art results among open models on agentic coding, browser-use, and tool-use benchmarks, with performance the team compared to Claude Sonnet 4. The API exposes 1,048,576 tokens of context, enough to keep a large repository or a long agent run in scope. This is the July 22, 2025 dated snapshot, pinned separately from the undated Qwen3 Coder Plus entry so its behavior stays fixed. Use it when you need a reproducible baseline to benchmark against, or a production pipeline that shouldn't shift when Alibaba updates the model later.
ChatQwen Plus (2025-07-14)
qwen/qwen-plus-2025-07-14
Qwen Plus (2025-07-14) is a snapshot from mid-2025 that keeps Qwen Plus in the Qwen3 series introduced a few months earlier, again combining thinking mode and non-thinking mode in one model with mode switching available during a conversation. Alibaba's release notes for this snapshot report a measurable improvement over the previous version in Chinese and English capabilities under non-thinking mode, along with enhanced tool-calling ability, which matters directly for API users building function-calling or agent workflows. It is served through Alibaba Cloud Model Studio, and as with any dated snapshot, calling this ID directly instead of the rolling qwen-plus alias keeps the model's behavior fixed rather than subject to Alibaba's ongoing updates.
ChatQwen3 235B-A22B Instruct 2507
qwen/qwen3-235b-a22b-instruct-2507
Qwen3 235B-A22B Instruct 2507 is Alibaba's non-thinking-mode update to its 235B-parameter Mixture-of-Experts flagship, activating 22B parameters per token and served with a 262,144-token context window. Against the original Qwen3-235B-A22B, Alibaba reports gains from 75.2 to 83.0 on MMLU-Pro, 24.7 to 70.3 on AIME25, and 52.0 to 79.2 on Arena-Hard v2. On GPQA it scores 77.5, ahead of GPT-4o's 66.9. It supports tool calling and is built for instruction following, coding, math, science, and multilingual tasks. Unlike the paired Thinking-2507 variant, it answers directly without a visible reasoning trace, trading step-by-step chain-of-thought for lower latency and fewer output tokens per request, useful for production workloads where response speed matters more than showing intermediate reasoning.
ChatQwen Plus (2025-04-28)
qwen/qwen-plus-2025-04-28
Qwen Plus (2025-04-28) is the snapshot where Alibaba folded Qwen Plus into the Qwen3 series, released the same month Alibaba introduced the wider Qwen3 model family. Alibaba's own documentation for this snapshot describes it as integrating thinking mode and non-thinking mode into one model, with switching between the two available mid-conversation. That lets a single API call trade deeper step-by-step reasoning for a faster, more direct answer depending on the query, instead of choosing between separate reasoning and non-reasoning models. Served through Alibaba Cloud Model Studio, it supports tool calling for agentic and function-calling use cases. As a dated snapshot it returns the same behavior on every call, useful for pinning a specific version in testing or production rather than tracking the rolling qwen-plus alias as Alibaba updates it.
ChatQVQ Max
qwen/qvq-max
QVQ Max is Alibaba's flagship visual reasoning model, built by the Qwen team to combine deep multimodal understanding with rigorous logical inference. Unlike standard vision-language models, QVQ Max is designed to think through what it sees — analyzing charts, diagrams, math problems, and everyday images step by step before responding. It scores 70.3% on MMMU and 71.4% on MathVista (mini), placing it among the top multimodal reasoning models available via API. The model handles text and image inputs across a 131K token context window and supports tool calling for agentic workflows. Ideal for developers building tutoring tools, visual data analysis pipelines, document understanding systems, or any application that requires both image comprehension and structured reasoning.
ChatQwen Plus (2025-01-25)
qwen/qwen-plus-2025-01-25
Qwen Plus (2025-01-25) is a dated snapshot of Alibaba's Qwen Plus model, accessed through Alibaba Cloud Model Studio. Alibaba positions Qwen Plus between Qwen Max and Qwen Turbo, balancing reasoning performance and response speed for moderately complex tasks rather than the hardest reasoning workloads or the cheapest, fastest ones. This snapshot predates the Qwen3-series branding that Alibaba later applied to Qwen Plus, and supports tool calling for building agents and function-calling workflows. Pinning to this dated ID keeps behavior fixed for a given deployment, unlike the rolling qwen-plus alias that Alibaba repoints to newer snapshots over time.
ChatQwen3.7 Max Preview
qwen/qwen3.7-max-preview
Qwen3.7 Max Preview is the earliest publicly reachable build of Alibaba's flagship Qwen3.7 Max reasoning model, offered during its preview access window ahead of the dated 2026-05-20 and 2026-06-08 snapshots. Qwen3.7 Max is built for long-horizon agentic workloads, pairing a 1 million-token context window with a chain-of-thought reasoning architecture. Alibaba demonstrated it running for 35 hours without degradation across more than 1,000 tool calls in a single session, later ranking 13th globally on LM Arena's text leaderboard and scoring 56.6 on the Artificial Analysis Intelligence Index, the highest of any Chinese model on that index. Because it carries no release date, this preview build predates the pinned snapshots and is best used for early evaluation and prototyping. For production API calls that need reproducible behavior over time, pin one of the dated snapshots instead.
ChatQwen3.8 27B Abliterated Cyber
qwen/qwen3.8-27b-abliterated-cyber:free
Qwen3.8 27B Abliterated Cyber is a free, community fine-tune of Alibaba's Qwen3.8 27B, served through the Infron aggregator rather than directly by Alibaba or the Qwen team. "Abliterated" refers to abliteration, a technique that identifies and removes the internal activation direction responsible for a model's safety refusals, rather than retraining it from scratch. Applied here, it produces a variant that answers prompts a standard instruction-tuned release would decline. The "Cyber" in its name points to a cybersecurity focus, suggesting use in exploit and malware analysis, penetration-testing scripts, and other security research work that base Qwen models often refuse. We found no independent documentation of who built this specific variant or how it performs, so benchmark claims can't be verified. It offers a 256,000-token context window and up to 131,072 output tokens at no cost, useful for testing uncensored, security-oriented prompts before committing to a paid model.
Frequently Asked Questions
The Qwen API gives you access to models for AI chat and image generation. Through Puter.js, you can start using Qwen models instantly with zero setup or configuration.
Puter.js supports a variety of Qwen models, including Qwen3.6 Flash, Qwen3.5 Plus 2026-04-20, Qwen3.6 27B, and more. Find all AI models supported by Puter.js in the AI model list.
With the User-Pays model, users cover their own AI costs through their Puter account. This means you can build apps without worrying about infrastructure expenses.
Puter.js is a JavaScript library that provides access to AI, storage, and other cloud services directly from a single API. It handles authentication, infrastructure, and scaling so you can focus on building your app.
Yes — the Qwen API through Puter.js works with any JavaScript framework, Node.js, or plain HTML. Just include the library and start building. See the documentation for more details.