slow than 27b-gguf-q4

#2
by fweifsd - opened

Thanks for the work. But I run it on my 4070tisuper , Only get 16t/s, why it slowerer than the base 27b-q4 model, that i can get 30t/s

Owner

did you have mtp enabled on the dense model? I'll do some testing tonight as that is an odd result. what where your settings in llama.cpp

Just the simple: llama serve -fa on --port 18080 -m Whittle-MoE-27B-A18B-v2.2.1-Q4_K_M.gguf --parallel 1 -c 8000 --load-mode none

i know the reason ,i test q3_k_m,it's ok in 40t/s。maybe it's caused by oom 。

fweifsd changed discussion status to closed

Sign up or log in to comment