Instructions to use firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS # Run inference directly in the terminal: llama cli -hf firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS # Run inference directly in the terminal: llama cli -hf firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS # Run inference directly in the terminal: ./llama-cli -hf firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS # Run inference directly in the terminal: ./build/bin/llama-cli -hf firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS
Use Docker
docker model run hf.co/firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS
- LM Studio
- Jan
- Ollama
How to use firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF with Ollama:
ollama run hf.co/firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS
- Unsloth Desktop
- Pi
How to use firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF with Docker Model Runner:
docker model run hf.co/firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS
- Lemonade
How to use firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS
Run and chat with the model
lemonade run user.Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF-IQ3_XXS
List all available models
lemonade list
- Hermes Agent
How to use firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Run and chat with the model
lemonade run user.Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF-IQ3_XXSList all available models
lemonade listQwen3.8-27B-pi RCO allocation transfer (experimental)
This repository contains an experimental 10.094 GB GGUF of bytkim/Qwen3.8-27B-pi, quantized with llama.cpp using the tensor-type allocation and base-model importance matrix published by ISTA-DASLab for Qwen3.8-27B. It also contains a conventional IQ2_M control quantized from the same Pi BF16 source with the same base imatrix.
These are Pi weights. They are not ISTA's GSQ-optimized base-model weights.
We did not run Pi-specific GSQ candidate generation, calibration, tensor
sensitivity analysis, or RCO search. The name RCO-allocation-transfer
describes exactly the part of the published work that was applied here.
Files
| File | Purpose | Bytes | SHA-256 |
|---|---|---|---|
Qwen3.8-27B-pi-RCO-allocation-transfer-IQ3_XXS.gguf |
Experimental 3.00 BPW transfer | 10,094,357,472 | 35134e868fc680f30b9764125b8b07fd359c935deefccb64db9025fb2936883c |
Qwen3.8-27B-pi-llama-IQ2_M-base-imatrix.gguf |
Same-source and same-imatrix conventional control | 10,004,593,632 | 11da0ac94be157b2911a97daac130e36692164aaba0c9daa0d144a9fa7460ec6 |
The original Pi BF16, published Pi baselines, and ISTA imatrix are available from their original publishers and are not rehosted here. These GGUFs contain the main model. No vision projector, MTP head, or DFlash2 companion is included or validated in this release.
Method
Source weights: bytkim/Qwen3.8-27B-pi-GGUF at revision
97caacb971d9246520fa8db5282a06bd36b103df, file
Qwen3.8-27B-pi-BF16.gguf (SHA-256
63d822f3982fdb893f03efbcb0eef6816a856074bb4bf9c23fd5a704509fe827).
Allocation and imatrix: ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF at revision
d562806dbafae37109975e970aae91b43e73b440. The allocation covered all
851 Pi tensor names and shapes. llama.cpp commit:
5e03bdd8700948b9c41c54dd1b00f28a2aebc03f. All 851 requested output
tensor types were verified. The base imatrix supplied 496 entries; the
quantizer reported no imatrix weights for token_embd.weight.
The Phase 1 result and
methodology are included in this repository. The complete
project source and raw logs are currently local; a public source-repository
link will be added when available. The A100 experiment used an earlier image
with documented script patches; a fixed replay image is
firetussin/qwen38-pi-rco-transfer-poc@sha256:545f36431c4f979bd871bec31dff18b94cf49f7ecf9979479d5518c78ce3082f.
Evaluation
One on-demand NVIDIA A100 PCIe 80 GB was used for about 2.092 hours; the actual user-reported Vast charge was $2.46. With llama.cpp CUDA, 8K context, flash attention, and q8_0 K/V cache, the transfer model processed a 512-token prompt at 1,088 tok/s and generated at 48.8 tok/s. The matched IQ2_M control measured 945 tok/s and 48.9 tok/s respectively. These are single-run measurements.
| Transfer context setting | Peak A100 VRAM | Needle retrieval |
|---|---|---|
| 32,768 | 11,451 MiB | 2/2 |
| 65,536 | 12,699 MiB | 2/2 |
| 76,800 (75K) | 13,129 MiB | 2/2 |
| 98,304 (96K) | 13,947 MiB | 2/2 |
| 131,072 (128K) | 15,197 MiB | 2/2 |
The small six-case quality suite did not demonstrate a material quality advantage over conventional similarly sized Pi IQ2_M. The retrieval probes used repetitive filler and a simple exact code. They test loading and basic retrieval, not general long-context reasoning. The coding prompt also exposed an unhandled duplicate-value edge case across generated fixes. No 16 GB GPU was tested; 128K used 15,197 MiB on the A100 and may be too tight on a specific 16 GB system. GGUF was tested with llama.cpp; native vLLM serving was not validated. See the full report before drawing quality conclusions.
Example llama.cpp invocation:
llama-server \
--model Qwen3.8-27B-pi-RCO-allocation-transfer-IQ3_XXS.gguf \
--ctx-size 65536 --parallel 1 --gpu-layers 99 \
--flash-attn on --cache-type-k q8_0 --cache-type-v q8_0
License and credits
The Qwen3.8-27B base, bytkim's Pi fine-tune and GGUF, and the ISTA-DASLab release are tagged Apache-2.0. The original Pi LICENSE and NOTICE are included here; NOTICE adds our quantization and packaging changes. We credit the Qwen team, bytkim, ISTA-DASLab, the GSQ and RCO authors, llama.cpp contributors, and Hugging Face. The original publishers do not endorse this experiment.
If citing the allocation source or method, please cite both papers linked by the ISTA-DASLab model card: GSQ (Dadgarnia et al., arXiv:2604.18556) and RCO (Helcig and Alistarh, arXiv:2605.00649). We used their published allocation and imatrix and make no claim to have developed GSQ, RCO, or the underlying model.
- Downloads last month
- 553
2-bit
3-bit
Pull the model
# Download Lemonade from https://lemonade-server.ai/lemonade pull firetussin/Qwen3.8-27B-pi-RCO-allocation-transfer-GGUF:IQ3_XXS