Quants

#1
by logic65 - opened

Im working on IQ Quants as it degrades below q8 but really appreciate the support

I also still need to merge the V2.1 adapter but added you as a contributor to the bottom of the card. Appreciate the support

Let me know on the Q3 if you get looping. I can maybe try some tricks if you do to try get it less. I found generally under Q6 it starts to degrade but will be working on it

The Q3 definitely loops less than Q2, but it still loops (as expected). Do you mind if I could get the looping benchmark scripts so I could run that benchmark on the Q3 models to see where they stand?

Plus, the models got:

Model Knowledge Battery Score
Q2_K 26/39
Q3_K_S 32/39
Q3_K_M 33/39
Q3_K_L 33/39

which is interesting because those are higher that the results you observed with the original model - maybe my setup differs?
(i sent you the files over kofi)

Thanks, Yeah thats kinda insane. I'll pull your Quants tonight and run the battery. I did an IQ quant of Q3 if you wanna test that in the meantime? the experts are treated with abit more priority in the IQ

1. serve any Whittle quant

llama-server -m Whittle-MoE-27B-A18B-v2.1-DQ3_K_XL.gguf -ngl 99 -c 8192 -fa on --jinja

2. in another terminal, run the same test that produced the card's numbers

curl -sLO https://huggingface.co/logic65/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF/resolve/main/loop_test.py
python3 loop_test.py http://localhost:8080 # or whatever port llama-server is set to

Owner

Thanks! I would try the IQ3 quant but it appears the .gguf is in fact larger than the Q4_K_M model? Regardless, I only have 15.1GiB of RAM, which isn't exactly ideal for testing a 17.9GB model, so I'm going to have to put that off until I get my hands on a machine with sufficient specs (which will probably be years away)

Let me try an IQ on a q2 damn experts being important haha.

The speed it pretty decent with offload to ram tbh
llama-server -m Whittle-MoE-27B-A18B-v2.1-DQ3_K_XL.gguf
-ngl 99 --n-cpu-moe 64 -c 8192 -fa on --jinja
Or

-ot '.ffn_(gate|up|down)_exps.=CPU'

NEW UD Whittle-MoE-27B-A18B-v2.1-DQ3_K_XL.gguf

Q3_K_M loop test results so far

type loopiness
single 8%
struct 44%
multi 0%
late 0%

initially i ran it with only 1200 tok context which in hindsight was nowhere near enough for the convos

I can run it tonight. would you like me to run yours?

Owner

it's fine i completed a full convo without any issues - i'll just let this run the other three convos

Owner

interesting - 0% loopiness for convos
haven't got it to output the actual responses so might get it to do that - it could be outputting gibberish for all i know

Owner

i feel my setup might be different from yours and isn't comparable... but if it is, it proves this model has much more potential than these benchmarks show

Owner

id recommend checking the results independently though

Sign up or log in to comment