How about Q4, Q3, Q2 size?

#1
by Gunfire - opened

As title.

@Gunfire sure! I mostly made this for myself, but if people are interested, I can create the full suite. Is there any specific size you want? Can prioritize it. Here are available ones:

Bits Quants (size in GB)
1 UD-IQ1_S (6.2), UD-IQ1_M (6.7)
2 UD-IQ2_XXS (7.3), UD-IQ2_S (8.4), UD-Q2_K_XL (9.8)
3 UD-IQ3_XXS (10.9), UD-IQ3_S (12.0), UD-Q3_K_XL (13.1)
4 UD-IQ4_XS (14.3), UD-Q4_K_S (15.4), UD-Q4_K_M (16.5)
5 UD-Q5_K_S (18.7), UD-Q5_K_XL (20.9)
6 UD-Q6_K (22.0), UD-Q6_K_M (23.1), UD-Q6_K_L (24.2)
8 UD-Q8_K_L (28.0)

@Gunfire Q2, Q3, Q4 all uploaded!
I'll be doing KLD whenever I have time later this week, but generally you can expect it to follow the base Unsloth Dynamic 3.0 curves.

@ajgazin Thank you very much! Would looking forward to your new works!

@ajgazin UD-IQ3_XXS (10.9) thank you! 😊 it’s a little unconventional but it’s perfect for 12GB VRAM <3

@kabeb I have to take a bit of a break right now -- work calls 😅
But I'll get it up when I have the chance, hopefully by tomorrow!

@kabeb uploaded! 10.9gb -- Note that about 400MB of that is the MTP head. If you run without MTP, will be closer to 10.4-10.5

Edit: note that performance wise, IQ3_XXS seems to be an odd one out. In my testing, isn't significantly better, if at all, than Q2_K_XL. This is based just on perplexity, though.
(I'm still looking into this, though I have my theories. It's funny, because my IQ3_XXS performs worse than Unsloth's (relative to base model), but the IQ2_X_L performs better. All small margins, though.) But I'll know more precisely when I run benchmarks.

Upon further testing/understanding, perplexity isn't tied to performance here, and no broad conclusion should be based on it. KLD, more pertinent metric, shows the curve as expected, and IQ3_XXS expectedly outperforms IQ2

Excellent. Bringing UD GGUF's is very much appreciated.

I understand abbliterated can drop quality by 1% ? - would you consider non-abbliterated UD's too? Q4_K_XL MTP. Thanks for your work!

@gavwhittaker Swift recently mentioned that they'll release a "Swift1.5", soon. Whenever that comes out, I'll make the UD quants of the base model along with everything else.

I'll let you know if I have a moment to get you specifically non-abliterated Q4_k_XL, but I'd have to start a new repo for that (and compare it to the quants ukisai - the creator - provide themselves. [[

Hi @ajgazin , thanks for the model! I wonder if UD IQ3_S can have better metrics with size that still fit on 16GB VRAM?

@ajgazin - thank you, I switched to your Q4_K_XL this morning. Together Swift and your work is exceptionally solid. Kudos!

@ajgazin UD-IQ4_XS(14.3)是我5060ti16g显卡最棒的选择!希望发布

@modoking -- ud-iq4_xs uploaded!

Hi @ajgazin , thanks for the model! I wonder if UD IQ3_S can have better metrics with size that still fit on 16GB VRAM?

@ajgazin , is my request worth the try?

@ydmhmhm uploaded!, it's at the expected 12gb. You should be able to fit that, though you'll have to a maintain a reasonable context size. Can also go up a quant size in exchange for q8_0 kv cache, which is very often worth it.

Hoping you have the time and inclination @ajgazin to updates these to Swift v1.5 :) - we appreciate you!
Super keen to get my hands on your UD3 Q4_K_XL version.

@gavwhittaker Swift1.5 versions here (abliterated): https://huggingface.co/ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-Dynamic-MTP-GGUF!
I also remember you asked for the base model (unabliterated) UD quants -- I'll get those up tonight.

Thank you very much!

Sign up or log in to comment