GGUF
qwen
experimental
llama.cpp
imatrix
conversational

Qwen3.8-27B-pi RCO allocation transfer (experimental)

This repository contains an experimental 10.094 GB GGUF of bytkim/Qwen3.8-27B-pi, quantized with llama.cpp using the tensor-type allocation and base-model importance matrix published by ISTA-DASLab for Qwen3.8-27B. It also contains a conventional IQ2_M control quantized from the same Pi BF16 source with the same base imatrix.

These are Pi weights. They are not ISTA's GSQ-optimized base-model weights. We did not run Pi-specific GSQ candidate generation, calibration, tensor sensitivity analysis, or RCO search. The name RCO-allocation-transfer describes exactly the part of the published work that was applied here.

Files

File Purpose Bytes SHA-256
Qwen3.8-27B-pi-RCO-allocation-transfer-IQ3_XXS.gguf Experimental 3.00 BPW transfer 10,094,357,472 35134e868fc680f30b9764125b8b07fd359c935deefccb64db9025fb2936883c
Qwen3.8-27B-pi-llama-IQ2_M-base-imatrix.gguf Same-source and same-imatrix conventional control 10,004,593,632 11da0ac94be157b2911a97daac130e36692164aaba0c9daa0d144a9fa7460ec6

The original Pi BF16, published Pi baselines, and ISTA imatrix are available from their original publishers and are not rehosted here. These GGUFs contain the main model. No vision projector, MTP head, or DFlash2 companion is included or validated in this release.

Method

Source weights: bytkim/Qwen3.8-27B-pi-GGUF at revision 97caacb971d9246520fa8db5282a06bd36b103df, file Qwen3.8-27B-pi-BF16.gguf (SHA-256 63d822f3982fdb893f03efbcb0eef6816a856074bb4bf9c23fd5a704509fe827). Allocation and imatrix: ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF at revision d562806dbafae37109975e970aae91b43e73b440. The allocation covered all 851 Pi tensor names and shapes. llama.cpp commit: 5e03bdd8700948b9c41c54dd1b00f28a2aebc03f. All 851 requested output tensor types were verified. The base imatrix supplied 496 entries; the quantizer reported no imatrix weights for token_embd.weight.

The Phase 1 result and methodology are included in this repository. The complete project source and raw logs are currently local; a public source-repository link will be added when available. The A100 experiment used an earlier image with documented script patches; a fixed replay image is firetussin/qwen38-pi-rco-transfer-poc@sha256:545f36431c4f979bd871bec31dff18b94cf49f7ecf9979479d5518c78ce3082f.

Evaluation

One on-demand NVIDIA A100 PCIe 80 GB was used for about 2.092 hours; the actual user-reported Vast charge was $2.46. With llama.cpp CUDA, 8K context, flash attention, and q8_0 K/V cache, the transfer model processed a 512-token prompt at 1,088 tok/s and generated at 48.8 tok/s. The matched IQ2_M control measured 945 tok/s and 48.9 tok/s respectively. These are single-run measurements.

Transfer context setting Peak A100 VRAM Needle retrieval
32,768 11,451 MiB 2/2
65,536 12,699 MiB 2/2
76,800 (75K) 13,129 MiB 2/2
98,304 (96K) 13,947 MiB 2/2
131,072 (128K) 15,197 MiB 2/2

The small six-case quality suite did not demonstrate a material quality advantage over conventional similarly sized Pi IQ2_M. The retrieval probes used repetitive filler and a simple exact code. They test loading and basic retrieval, not general long-context reasoning. The coding prompt also exposed an unhandled duplicate-value edge case across generated fixes. No 16 GB GPU was tested; 128K used 15,197 MiB on the A100 and may be too tight on a specific 16 GB system. GGUF was tested with llama.cpp; native vLLM serving was not validated. See the full report before drawing quality conclusions.

Example llama.cpp invocation:

llama-server \
  --model Qwen3.8-27B-pi-RCO-allocation-transfer-IQ3_XXS.gguf \
  --ctx-size 65536 --parallel 1 --gpu-layers 99 \
  --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0

License and credits

The Qwen3.8-27B base, bytkim's Pi fine-tune and GGUF, and the ISTA-DASLab release are tagged Apache-2.0. The original Pi LICENSE and NOTICE are included here; NOTICE adds our quantization and packaging changes. We credit the Qwen team, bytkim, ISTA-DASLab, the GSQ and RCO authors, llama.cpp contributors, and Hugging Face. The original publishers do not endorse this experiment.

If citing the allocation source or method, please cite both papers linked by the ISTA-DASLab model card: GSQ (Dadgarnia et al., arXiv:2604.18556) and RCO (Helcig and Alistarh, arXiv:2605.00649). We used their published allocation and imatrix and make no claim to have developed GSQ, RCO, or the underlying model.

Downloads last month
527
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(10)
this model

Papers for firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF