Instructions to use ajgazin/Swift-Qwen3.8-27B-Uncensored-Dynamic-MTP-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Local Apps Settings
- Unsloth Desktop
How about Q4, Q3, Q2 size?
As title.
@Gunfire sure! I mostly made this for myself, but if people are interested, I can create the full suite. Is there any specific size you want? Can prioritize it. Here are available ones:
| Bits | Quants (size in GB) |
|---|---|
| 1 | UD-IQ1_S (6.2), UD-IQ1_M (6.7) |
| 2 | UD-IQ2_XXS (7.3), UD-IQ2_S (8.4), UD-Q2_K_XL (9.8) |
| 3 | UD-IQ3_XXS (10.9), UD-IQ3_S (12.0), UD-Q3_K_XL (13.1) |
| 4 | UD-IQ4_XS (14.3), UD-Q4_K_S (15.4), UD-Q4_K_M (16.5) |
| 5 | UD-Q5_K_S (18.7), UD-Q5_K_XL (20.9) |
| 6 | UD-Q6_K (22.0), UD-Q6_K_M (23.1), UD-Q6_K_L (24.2) |
| 8 | UD-Q8_K_L (28.0) |
@ajgazin UD-IQ3_XXS (10.9) thank you! 😊 it’s a little unconventional but it’s perfect for 12GB VRAM <3
@kabeb uploaded! 10.9gb -- Note that about 400MB of that is the MTP head. If you run without MTP, will be closer to 10.4-10.5
Edit: note that performance wise, IQ3_XXS seems to be an odd one out. In my testing, isn't significantly better, if at all, than Q2_K_XL. This is based just on perplexity, though.
(I'm still looking into this, though I have my theories. It's funny, because my IQ3_XXS performs worse than Unsloth's (relative to base model), but the IQ2_X_L performs better. All small margins, though.) But I'll know more precisely when I run benchmarks.
Upon further testing/understanding, perplexity isn't tied to performance here, and no broad conclusion should be based on it. KLD, more pertinent metric, shows the curve as expected, and IQ3_XXS expectedly outperforms IQ2
Excellent. Bringing UD GGUF's is very much appreciated.
I understand abbliterated can drop quality by 1% ? - would you consider non-abbliterated UD's too? Q4_K_XL MTP. Thanks for your work!
@gavwhittaker Swift recently mentioned that they'll release a "Swift1.5", soon. Whenever that comes out, I'll make the UD quants of the base model along with everything else.
I'll let you know if I have a moment to get you specifically non-abliterated Q4_k_XL, but I'd have to start a new repo for that (and compare it to the quants ukisai - the creator - provide themselves. [[
@ajgazin - thank you, I switched to your Q4_K_XL this morning. Together Swift and your work is exceptionally solid. Kudos!
@ydmhmhm uploaded!, it's at the expected 12gb. You should be able to fit that, though you'll have to a maintain a reasonable context size. Can also go up a quant size in exchange for q8_0 kv cache, which is very often worth it.
Hoping you have the time and inclination @ajgazin to updates these to Swift v1.5 :) - we appreciate you!
Super keen to get my hands on your UD3 Q4_K_XL version.
@gavwhittaker Swift1.5 versions here (abliterated): https://huggingface.co/ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-Dynamic-MTP-GGUF!
I also remember you asked for the base model (unabliterated) UD quants -- I'll get those up tonight.
Thank you very much!