Unlock exclusive introductory pricing for the newly launched Gemini 3.5 Flash.
Q

qwen3-vl-235b-a22b

Entrรฉe:$0.24/M
Sortie:$0.96/M
Contexte:2M
Sortie maximale:30K
Publiรฉ:Oct 1, 2025

qwen3-vl-235b-a22b is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception of real-world/synthetic categories, 2D/3D spatial grounding, and long-form visual comprehension, achieving competitive multimodal benchmark results.

Nouveau
Usage commercial

Playground pour qwen3-vl-235b-a22b

Explorez le Playground de qwen3-vl-235b-a22b โ€” un environnement interactif pour tester les modรจles et exรฉcuter des requรชtes en temps rรฉel. Essayez des invites, ajustez les paramรจtres et itรฉrez instantanรฉment pour accรฉlรฉrer le dรฉveloppement et valider les cas d'utilisation.

What is Qwen3-VL-235B-A22B

Qwen3-VL-235B-A22B is a high-capacity multimodal LLM from the Qwen (Alibaba) family. It combines a large MoE transformer backbone with cross-modal vision encoders and new positional/time encoding techniques to handle multi-image and long-duration video inputs, and to perform tasks such as visual question answering (VQA), long-document OCR, spatial/3D grounding, multimodal code generation, and agentic GUI control. The release includes both Instruct (task/few-shot tuned for instruction following) and Thinking (additional reasoning support and internal โ€œthinkโ€ mode) variants.


Main features (what makes Qwen3-VL-235B-A22B distinctive)

  • Large MoE design with high active capacity: a MoE stack that activates a subset of experts per request (โ‰ˆ22B active) to give more compute when needed while controlling inference cost.
  • Very long native context (256K) and scalable to ~1M: intended for book-length documents, hours of video, and multi-document workflows without aggressive chunking.
  • Advanced visual reasoning (spatial & temporal): Interleaved-MRoPE and DeepStack modules for timestamp alignment and fine-grained imageโ€“text fusion enabling video timeline queries and 3D grounding.
  • Improved OCR & document parsing: expanded OCR language support (advertised ~32 languages), stronger robustness to blur/tilt/low light and long, multi-page document structure parsing.
  • Visual agent + GUI automation: explicit agent capabilities to identify GUI elements, invoke functions or tools, and perform automation tasks on PC/mobile UIs.
  • Visual coding & multimodal program synthesis: can translate images/video/UI sketches into Draw.io/HTML/CSS/JS and assist in UI debugging.

How Qwen3-VL-235B-A22B compares to other models

Below are high-level comparisons to contemporaries; numbers and caps are taken from public provider/model pages and aggregator writeups.

  • Google Gemini 3 Pro โ€” Gemini emphasizes very large multimodal reasoning and agentic tool use; Google advertises 1M token context modes and deep product integrations. Gemini is positioned as a general leader in agentic multimodality (closed-source / proprietary), and often outperforms publicly available open models on some productized benchmarks. Qwen3-VL competes more directly as a high-capacity open-weight alternative optimized for OCR, video timeline alignment, and MoE cost tradeoffs.
  • Grok-4 Heavy (xAI) โ€” Grok-4 is another long-context, high-reasoning model family; some Grok variants list ~256K context windows and strong coding/math performance. Qwen3-VL and Grok-4 both target long-form reasoning; Qwen3-VL differentiates via heavy visual/video/OCR tooling and MoE scaling.
  • DeepSeek-R1 / DeepSeek family โ€” DeepSeek R1 emphasizes efficient training and competitive reasoning performance at lower inference cost; it is often used as an open alternative for reasoning/code tasks. Qwen3-VL targets stronger multimodal and spatial/video capabilities than R1โ€™s primary focus on text reasoning.

Representative use cases

  • Document parsing and large-scale OCR โ€” long, multi-page invoices, books, historical documents with multilingual text.
  • Video understanding & timeline queries โ€” summarize hours of recorded video, locate events by time, align text to video timestamps.
  • Visual question answering & multimodal assistants โ€” multi-turn image + text dialogs (customer support with screenshots, medical imaging notes).
  • GUI automation / visual agents โ€” detect UI elements and drive PC/mobile flows (automation, testing, assistive agents).
  • Multimodal code generation & UI prototyping โ€” convert mockups / images into HTML/CSS/JS or Draw.io diagrams.
  • Research & large-document analysis โ€” book-level summarization, multi-document synthesis with a single context.

How to access Qwen3 VL-235B-A22B API

Step 1: Sign Up for API Key

Log in to cometapi.com. If you are not our user yet, please register first. Sign into your CometAPI console. Get the access credential API key of the interface. Click โ€œAdd Tokenโ€ at the API token in the personal center, get the token key: sk-xxxxx and submit.

Step 2: Send Requests to Qwen3 VL-235B-A22B API

Select the โ€œQwen3-VL-235B-A22Bโ€ endpoint to send the API request and set the request body. The request method and request body are obtained from our website API doc. Our website also provides Apifox test for your convenience. Replace <YOUR_API_KEY> with your actual CometAPI key from your account. base url is Chat

Insert your question or request into the content fieldโ€”this is what the model will respond to . Process the API response to get the generated answer.

Step 3: Retrieve and Verify Results

Process the API response to get the generated answer. After processing, the API responds with the task status and output data.

Tarification pour qwen3-vl-235b-a22b

Dรฉcouvrez des tarifs compรฉtitifs pour qwen3-vl-235b-a22b, conรงus pour s'adapter ร  diffรฉrents budgets et besoins d'utilisation. Nos formules flexibles garantissent que vous ne payez que ce que vous utilisez, ce qui facilite l'adaptation ร  mesure que vos besoins รฉvoluent. Dรฉcouvrez comment qwen3-vl-235b-a22b peut amรฉliorer vos projets tout en maรฎtrisant les coรปts.

ModelComet Price (USD / M Tokens)Official Price (USD / M Tokens)Discount
qwen3-vl-235b-a22b
Entrรฉe:$0.24/M
Sortie:$0.96/M
Entrรฉe:$0.3/M
Sortie:$1.2/M
-20%

Exemple de code et API pour qwen3-vl-235b-a22b

Accรฉdez ร  des exemples de code complets et aux ressources API pour qwen3-vl-235b-a22b afin de simplifier votre processus d'intรฉgration. Notre documentation dรฉtaillรฉe fournit des instructions รฉtape par รฉtape pour vous aider ร  exploiter tout le potentiel de qwen3-vl-235b-a22b dans vos projets.

curl https://api.cometapi.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $COMETAPI_KEY" \
  -d '{
    "model": "qwen3-vl-235b-a22b",
    "messages": [
      {
        "role": "system",
        "content": "You are a helpful assistant."
      },
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }'

cURL Code Example

curl https://api.cometapi.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $COMETAPI_KEY" \
  -d '{
    "model": "qwen3-vl-235b-a22b",
    "messages": [
      {
        "role": "system",
        "content": "You are a helpful assistant."
      },
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }'

Python Code Example

from openai import OpenAI
import os

# Get your CometAPI key from https://api.cometapi.com/console/token, and paste it here
COMETAPI_KEY = os.environ.get("COMETAPI_KEY") or "<YOUR_COMETAPI_KEY>"
BASE_URL = "https://api.cometapi.com/v1"

client = OpenAI(base_url=BASE_URL, api_key=COMETAPI_KEY)

completion = client.chat.completions.create(
    model="qwen3-vl-235b-a22b",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Hello!"},
    ],
)

print(completion.choices[0].message.content)

JavaScript Code Example

import OpenAI from "openai";

// Get your CometAPI key from https://api.cometapi.com/console/token, and paste it here
const api_key = process.env.COMETAPI_KEY || "<YOUR_COMETAPI_KEY>";
const base_url = "https://api.cometapi.com/v1";

const openai = new OpenAI({
  apiKey: api_key,
  baseURL: base_url,
});

const completion = await openai.chat.completions.create({
  model: "qwen3-vl-235b-a22b",
  messages: [
    { role: "system", content: "You are a helpful assistant." },
    { role: "user", content: "Hello!" },
  ],
});

console.log(completion.choices[0].message.content);

Versions de qwen3-vl-235b-a22b

La raison pour laquelle qwen3-vl-235b-a22b dispose de plusieurs instantanรฉs peut inclure des facteurs potentiels tels que des variations de sortie aprรจs des mises ร  jour nรฉcessitant des instantanรฉs plus anciens pour la cohรฉrence, offrant aux dรฉveloppeurs une pรฉriode de transition pour l'adaptation et la migration, et diffรฉrents instantanรฉs correspondant ร  des points de terminaison globaux ou rรฉgionaux pour optimiser l'expรฉrience utilisateur. Pour les diffรฉrences dรฉtaillรฉes entre les versions, veuillez consulter la documentation officielle.

Model namedescription
qwen3-vl-235b-a22bstandard
qwen3-vl-235b-a22b-thinkingthinking version