For the complete documentation index, see llms.txt. This page is also available as Markdown.

Qwen3-VL

Use Alibaba's Qwen3-VL vision-language model through our Serverless Cloud API

Qwen3-VL is Alibaba's vision-language model. It accepts an image and a text prompt and returns a text response. We support Qwen3-VL through our Serverless Cloud API, Dedicated Deployments, and self-hosted Inference.

Code sample

1

Get your API Key

Create a Roboflow account, find your key on the Roboflow API settings page and make it available to your shell:

export ROBOFLOW_API_KEY="your-key-here"
2

Install the dependencies

Install the Inference SDK:

pip install -U inference-sdk supervision
3

Run the model

The sample prompts the qwen3vl-2b-instruct checkpoint to describe an image and prints the response.

import os
import supervision as sv
from inference_sdk import InferenceHTTPClient

image = sv.load_image_from_url("https://media.roboflow.com/quickstart/dog.jpeg")

client = InferenceHTTPClient(
    api_url="https://serverless.roboflow.com",
    api_key=os.environ["ROBOFLOW_API_KEY"],
)
result = client.infer_lmm(
    image,
    model_id="qwen3vl-2b-instruct",
    prompt="Describe this image briefly.",
    max_new_tokens=128,
)
print(result["response"])

The code above prints the model response to the terminal:

A man in a white t-shirt and red shorts is carrying a beagle dog on his shoulders. The dog is wearing a black harness and is looking forward. The man is walking on a paved path in a residential area with apartment buildings in the background. There is a small garden with green grass and white flowers to the left.

Inference speed

Latency measured with Roboflow Inference on 1x NVIDIA L4, batch size 1, generating exactly 128 tokens with greedy decoding from a fixed prompt. Latency scales with output length, so use tokens/sec to estimate other lengths.

Alias
Latency, 128 tokens (ms)
Tokens/sec

qwen3vl-2b-instruct

4057

32

qwen25-vl-7b

5603

23

qwen25-vl-7b is the earlier Qwen2.5-VL checkpoint. It is listed here because it shares this alias namespace and runs through the same block.

Set api_url to match your deployment target:

  • https://serverless.roboflow.com for the Serverless Cloud API.

  • http://localhost:9001 for a local Inference server.

  • Your Dedicated Deployment URL for a private endpoint.

You can train your own Qwen3-VL checkpoint on Roboflow and call it by its per-model {workspace}/{model-slug} ID (see Versions, Trainings, and Models).

Last updated

Was this helpful?