hero

Qwen-AgentWorld-35B-A3B-NVFP4

NVFP4-quantized Qwen-AgentWorld-35B-A3B for vLLM. Uses FlashInfer Cutlass NVFP4 GEMM kernels, compressed-tensors quantization, and the Qwen3_5MoeForConditionalGeneration architecture.

Benchmarks (NVIDIA GB10, 131 GB VRAM, vLLM v0.23.0)

Metric Value
Model VRAM 20.94 GB
Single prompt ~26 tokens/sec
Batch 4 ~86 tokens/sec
Batch 8 ~181 tokens/sec

In Testing

temperature=0.6, max 1024 tokens, enforce_eager.

Prompt Result
Capital of France Correct
US 2020/2024 elections Correct
Alice apples (reasoning) Correct
Syllogism (logic) Correct
137*429 (math) Correct (58773)
Fibonacci function Correct

Important: The model outputs <think>...</think> blocks (Qwen3.5 thinking format). At temperature=0 the greedy sampling may stall in the thinking phase. Use temperature>=0.6 (the model's default) for reliable answer extraction. For agent workloads, use --max-tokens >= 1024 to accommodate think+answer.

Quick Start

pip install vllm flashinfer

vllm serve ./nvfp4_model \
    --enforce-eager \
    --max-model-len 4096 \
    --port 8000

Query:

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"./nvfp4_model","messages":[{"role":"user","content":"Hello!"}],"max_tokens":128,"temperature":0}'

Credits

Downloads last month
54
Safetensors
Model size
20B params
Tensor type
F32
·
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Frosty40/Qwen-AgentWorld-35B-A3B-NVFP4

Quantized
(73)
this model