All QUASAR Models
All QUASAR checkpoints in one place: 4-bit QAT for Qwen, Gemma and Muse — NVFP4 W4A16/W4A4 for vLLM, Q4_0 GGUF for llama.cpp / Ollama.
Image-Text-to-Text • 28B • Updated • 47.4k • 131Note 27B · Blackwell W4A4 · 19.7 GB · 496/496 linears · GPQA 90.9 vs 89.4 / 87.6 on compared NVFP4 builds · AIME'26 100
QUASAR-QAT/Qwen3.5-4B-QUASAR-Q4_0-GGUF
Image-Text-to-Text • 4B • Updated • 98 • 1Note 4B · native Q4_0 QAT · 2.70 GB · best fidelity of tested Q4_0 builds · 35% lower KL vs tested QAT Q4_0 · beats Q4_K_M at smaller size
QUASAR-QAT/Qwen3.5-4B-QUASAR-NVFP4
Image-Text-to-Text • 5B • Updated • 265 • 1Note 4B · H100 W4A16 · 4.2 GB · 200/200 projections · 2.3–2.7× lower KL than compared 4-bit builds · +35% c=1 vs BF16
QUASAR-QAT/Qwen3.5-4B-QUASAR-NVFP4-W4A4
Image-Text-to-Text • 5B • Updated • 30 • 1Note 4B · Blackwell W4A4 · 4.2 GB · 200/200 projections · 39% lower KL vs cosmicproc W4A4 · vision + MTP retained
QUASAR-QAT/gemma-4-E4B-it-QUASAR-Q4_0-GGUF
Image-Text-to-Text • 7B • Updated • 538 • 1Note E4B · Local · 5.2 GB · native Q4_0 · 2× lower KL than Google QAT · 1.5× closer than Unsloth Q4_0
QUASAR-QAT/gemma-4-12B-it-QUASAR-Q4_0-GGUF
Image-Text-to-Text • 12B • Updated • 791 • 3Note 12B · Local · 7.0 GB · native Q4_0 · 30% lower KL than Google QAT · 2.2× closer than Unsloth Q4_0
QUASAR-QAT/gemma-4-E4B-it-QUASAR-W4A16-G64
Any-to-Any • 8B • Updated • 67 • 1Note E4B · vLLM · W4A16 g64 · 3× lower KL than Google QAT · +36% H100 output throughput in our test
QUASAR-QAT/gemma-4-12B-it-QUASAR-W4A16-G64
Any-to-Any • 12B • Updated • 73 • 1Note 12B · vLLM · W4A16 g64 · 23% lower KL than Google QAT · 4.25 vs 4.5 bpw
QUASAR-QAT/Muse-Glimmer-30B-QUASAR-NVFP4
Image-Text-to-Text • 30B • Updated • 127 • 2Note W4A16 · vLLM · 21.8 GiB · 72% lower KL vs Red Hat · 39% lower vs NVIDIA · above BF16 on RULER 128K
QUASAR-QAT/Muse-Glimmer-30B-QUASAR-Q4_0-GGUF
Text Generation • 28B • Updated • 267 • 3Note Q4_0 GGUF · local (llama.cpp/Ollama/LM Studio) · 19.6 GB · lower KL than Meta Q4_K_M
QUASAR-QAT/Muse-Glimmer-30B-QUASAR-NVFP4-W4A4
Image-Text-to-Text • 30B • Updated • 91 • 2Note W4A4 · Blackwell · 21.8 GiB · 20% lower KL vs Red Hat · 45% lower vs RTN
QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction
Paper • 2608.13966 • Published • 4Note QUASAR paper