Qwen 3.5 9B GGUF β€” Quantized by BatiAI

BatiFlow Ollama

Quantizations of Qwen 3.5 9B for on-device AI on Mac. Built and verified by BatiAI for BatiFlow.

Why Qwen 3.5 9B?

  • Strong tool calling and JSON structured output
  • Multilingual (en, ko, ja, zh) β€” solid Korean
  • Apache 2.0 β€” commercial-friendly
  • Fits 16GB Mac with Q4

Quick Start

ollama pull batiai/qwen3.5-9b:q4

Available Quantizations

Quant Size VRAM target Recommended For
Q4_K_M 5.2GB ~7GB 16GB Mac
Q6_K 7.5GB ~10GB 24GB+ Mac

RAM Requirements

Your Mac RAM Q4 (5.2GB) Q6 (7.5GB)
16GB βœ… Recommended ⚠️ Tight
24GB+ βœ… βœ… Headroom

Model Comparison

Your Mac Best Pick
16GB batiai/qwen3.5-9b:q4 (this) or batiai/gemma4-e4b:q4
24GB batiai/gemma4-26b:iq4
32GB batiai/nemotron3-nano:iq4
36GB+ batiai/qwen3.6-35b:iq4

Why BatiAI?

  • Quantized from official Qwen weights
  • Verified on real Mac hardware
  • BatiAI metadata signed (general.author=BatiAI)
  • Tool calling + Korean validated

Technical Details

About BatiFlow

BatiFlow β€” free, on-device AI automation for Mac. 5MB app, 100% local, unlimited.

Benchmarks

Machine Quant Cold start Prompt eval Token gen Tested
Mac mini M4 16GB Q4_K_M β€”s β€” t/s β€” t/s 2026-05-01
Mac mini M4 16GB Q6_K β€”s β€” t/s β€” t/s 2026-05-01
MacBook Pro M4 Max 128GB Q4_K_M 2.274s 272.23 t/s 46.08 t/s 2026-05-03
MacBook Pro M4 Max 128GB Q6_K 2.552s 304.85 t/s 38.54 t/s 2026-05-03
Downloads last month
132
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for batiai/Qwen3.5-9B-GGUF

Finetuned
Qwen/Qwen3.5-9B
Quantized
(515)
this model

Collections including batiai/Qwen3.5-9B-GGUF