Qwen3.8 16B

#34
by Banaxi-Tech - opened

Hello Qwen! If you see this, i would really like a Qwen3.8 16B or 14B because your releasing Qwen3.8 27B, but i only have 16GB VRAM. Like and comment if you guys want this too.

They should release a 35B A3B version

And a 27B version

No one wants 16B? 😭
I guess my GPu is just too bad xD

I kinda want

I kinda want

Me too. I'd also like to see smaller models again (like 4B and 9B).

YESS!

No one wants 16B? 😭
I guess my GPu is just too bad xD

An MOE Model should run quite well on your GPU with the right layers offloaded. So...in your case I would hope for something like 35B A3B.

Yeah I want 35b and 14b

I kinda want

Me too. I'd also like to see smaller models again (like 4B and 9B).

4B would be interesting for mobile use

At this point they should just release Qwen 4 soon

Why not just use 27B quantized?

Because of KV cache

Because 27B is slow as hell like 13 TPS and TTFT is like 30 seconds for a few thousand tokens

Sign up or log in to comment