Qwen3-VL-2B-Thinking
This repository contains Qwen/Qwen3-VL-2B-Thinking together with a Furiosa Executable Bundle (FXB) for running it on FuriosaAI RNGD with Furiosa-LLM. The same model also runs on other frameworks (such as vLLM, SGLang, and Transformers); for usage with those, see the upstream Qwen/Qwen3-VL-2B-Thinking model card.
Overview
Qwen3-VL-2B-Thinking is a 2-billion-parameter dense vision-language model from the Qwen3-VL series. It pairs a vision encoder with a dense transformer decoder, using Interleaved-MRoPE positional embeddings and DeepStack multi-level feature fusion to handle images and videos alongside text. The model covers visual understanding tasks such as OCR, document and chart analysis, spatial reasoning, and video comprehension, and it natively supports tool (function) calling. This is the Thinking edition, which emits an explicit chain-of-thought before its final answer. Its intended use is the same as the upstream Qwen/Qwen3-VL-2B-Thinking, released under the Apache 2.0 License.
- Architecture: Qwen3-VL (dense)
- Input / Output: Image + Text / Text
- Supported Inference Engine: Furiosa LLM
- Supported Hardware: FuriosaAI RNGD
Quantization
No quantization is applied โ the model runs in the same precision as the upstream weights.
Features
- Vision-language. The model accepts OpenAI-style multimodal chat messages with
image_urlcontent parts alongside text. - Reasoning. The model emits a chain-of-thought that Furiosa-LLM parses through the
qwen3reasoning parser. - Tool calling. The model supports tool (function) calling through the
hermestool-call parser.
Parallelism Strategy
On RNGD, Qwen3-VL-2B-Thinking runs with a tensor-parallel size of 8 PEs, which maps to a single RNGD card (8 PEs per card).
Usage
To run this model with Furiosa-LLM, follow the example commands below after installing Furiosa-LLM and its prerequisites.
Launch the server
Because this is a Thinking model, launch the server with the qwen3 reasoning
parser so its chain-of-thought is separated from the final answer:
# Launch the server, listening on port 8000 by default
furiosa-llm serve furiosa-ai/Qwen3-VL-2B-Thinking \
--reasoning-parser qwen3
To also enable tool (function) calling, add the hermes tool-call parser (the
parser used by the Qwen3 series); keep --reasoning-parser qwen3 so thinking is
still parsed into its own field:
furiosa-llm serve furiosa-ai/Qwen3-VL-2B-Thinking \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser hermes
When the server is ready, you will see:
INFO: Started server process [27507]
INFO: Waiting for application startup.
INFO: Application startup complete.
INFO: Uvicorn running on http://0.0.0.0:8000 (Press CTRL+C to quit)
Basic Usage
The server exposes an OpenAI-compatible API. You can send a text-only request
with curl:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "furiosa-ai/Qwen3-VL-2B-Thinking",
"messages": [{"role": "user", "content": "What is the capital of France?"}]
}' \
| python -m json.tool
To ask about an image, pass an image_url content part in the message:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "furiosa-ai/Qwen3-VL-2B-Thinking",
"messages": [{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg"}},
{"type": "text", "text": "Describe this image."}
]
}]
}' \
| python -m json.tool
The image_url.url field accepts a remote http:///https:// URL, an inline
base64 data: URL, or a local file:// path (the last requires the
--allowed-local-media-path flag described under Advanced Usage).
Because this is a Thinking edition, its chain-of-thought is returned separately from the final answer:
response.choices[].message.reasoning(non-streaming)response.choices[].delta.reasoning(streaming)
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="furiosa-ai/Qwen3-VL-2B-Thinking",
messages=[{"role": "user", "content": "How many r's are in 'strawberry'?"}],
)
print("Reasoning:", response.choices[0].message.reasoning)
print("Answer:", response.choices[0].message.content)
Note: The
reasoningfield is not part of the OpenAI API specification but is a widely followed convention (the OpenAI Agents SDK, vLLM, and others). It appears only in responses that contain reasoning content; accessing it otherwise raises anAttributeError.
Advanced Usage
Multimodal serving options. furiosa-llm serve provides flags to control
multimodal behavior; requests that violate them are rejected with HTTP 400:
--image-limit-per-prompt N/--video-limit-per-prompt Nโ maximum number of images/videos allowed per request (default: unlimited).--allowed-local-media-path PATHโ allowfile://URLs whose resolved path is underPATH. Local file access is disabled unless this is set.--allowed-media-domains D [D ...]โ whitelist of remote domains for SSRF protection. When set, only images from the listed domains are fetched.--interleave-mm-stringsโ keep image placeholders at their original positions when the model uses a string-format chat template (no-op for OpenAI-format templates, the common case).--mm-processor-cache-gb GBโ size of the UUID-keyed multimodal processor cache (default: 4.0). Clients can tag animage_urlpart with auuidfield and re-reference it in follow-up requests without re-uploading the image bytes; set to 0 to disable.
For example, to serve local images under /srv/media and restrict remote
fetches to a single domain:
furiosa-llm serve furiosa-ai/Qwen3-VL-2B-Thinking \
--reasoning-parser qwen3 \
--allowed-local-media-path /srv/media \
--allowed-media-domains cdn.example.com \
--image-limit-per-prompt 4
See the Vision-Language Models guide for image input formats, the UUID cache, and Python client examples.
Reasoning. This is a Thinking edition, so it always reasons; there is no
enable_thinking switch (the Instruct editions are the non-thinking
counterparts).
Tool calling. With the server launched using
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser hermes
(see Launch the server), pass tools in the request and
let the model decide when to call them. See the
Tool Calling guide
for a complete client example and details on tool-choice options.
Learn more
- Vision-Language Models โ image input formats, multimodal server options, and the UUID cache
- Tool Calling โ parsers, tool-choice options, and more examples
- Furiosa-LLM Server (
furiosa-llm serve) โ full OpenAI-compatible API reference and serving options - Qwen/Qwen3-VL-2B-Thinking โ upstream model card
- Downloads last month
- 2,773
Model tree for furiosa-ai/Qwen3-VL-2B-Thinking
Base model
Qwen/Qwen3-VL-2B-Thinking