Instructions to use scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S # Run inference directly in the terminal: llama cli -hf scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S # Run inference directly in the terminal: llama cli -hf scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S # Run inference directly in the terminal: ./llama-cli -hf scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S
Use Docker
docker model run hf.co/scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S
- LM Studio
- Jan
- vLLM
How to use scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S
- Ollama
How to use scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants with Ollama:
ollama run hf.co/scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S
- Unsloth Desktop
- Pi
How to use scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants with Docker Model Runner:
docker model run hf.co/scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S
- Lemonade
How to use scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S
Run and chat with the model
lemonade run user.Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants-Q4_K_S
List all available models
lemonade list
- Hermes Agent
How to use scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quants
Im working on IQ Quants as it degrades below q8 but really appreciate the support
I also still need to merge the V2.1 adapter but added you as a contributor to the bottom of the card. Appreciate the support
Let me know on the Q3 if you get looping. I can maybe try some tricks if you do to try get it less. I found generally under Q6 it starts to degrade but will be working on it
The Q3 definitely loops less than Q2, but it still loops (as expected). Do you mind if I could get the looping benchmark scripts so I could run that benchmark on the Q3 models to see where they stand?
Plus, the models got:
| Model | Knowledge Battery Score |
|---|---|
| Q2_K | 26/39 |
| Q3_K_S | 32/39 |
| Q3_K_M | 33/39 |
| Q3_K_L | 33/39 |
which is interesting because those are higher that the results you observed with the original model - maybe my setup differs?
(i sent you the files over kofi)
Thanks, Yeah thats kinda insane. I'll pull your Quants tonight and run the battery. I did an IQ quant of Q3 if you wanna test that in the meantime? the experts are treated with abit more priority in the IQ
1. serve any Whittle quant
llama-server -m Whittle-MoE-27B-A18B-v2.1-DQ3_K_XL.gguf -ngl 99 -c 8192 -fa on --jinja
2. in another terminal, run the same test that produced the card's numbers
curl -sLO https://huggingface.co/logic65/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF/resolve/main/loop_test.py
python3 loop_test.py http://localhost:8080 # or whatever port llama-server is set to
Thanks! I would try the IQ3 quant but it appears the .gguf is in fact larger than the Q4_K_M model? Regardless, I only have 15.1GiB of RAM, which isn't exactly ideal for testing a 17.9GB model, so I'm going to have to put that off until I get my hands on a machine with sufficient specs (which will probably be years away)
Let me try an IQ on a q2 damn experts being important haha.
The speed it pretty decent with offload to ram tbh
llama-server -m Whittle-MoE-27B-A18B-v2.1-DQ3_K_XL.gguf
-ngl 99 --n-cpu-moe 64 -c 8192 -fa on --jinja
Or
-ot '.ffn_(gate|up|down)_exps.=CPU'
NEW UD Whittle-MoE-27B-A18B-v2.1-DQ3_K_XL.gguf
Q3_K_M loop test results so far
| type | loopiness |
|---|---|
| single | 8% |
| struct | 44% |
| multi | 0% |
| late | 0% |
initially i ran it with only 1200 tok context which in hindsight was nowhere near enough for the convos
I can run it tonight. would you like me to run yours?
it's fine i completed a full convo without any issues - i'll just let this run the other three convos
interesting - 0% loopiness for convos
haven't got it to output the actual responses so might get it to do that - it could be outputting gibberish for all i know
i feel my setup might be different from yours and isn't comparable... but if it is, it proves this model has much more potential than these benchmarks show
id recommend checking the results independently though