pearsonkyle's picture
Add IQ2_M (AWQ) + IQ4_XS quants; method chosen by agentic tool-call fidelity. At 2-bit, plain imatrix loses tool-argument precision (param-acc 5%) despite good KLD; AWQ's activation-aware scaling recovers it (33%), so IQ2_M ships AWQ. IQ4_XS (imatrix) matches Q5_K_M agentically at 20% smaller. Refresh README + calibration corpus.
c995b4f verified