Instructions to use logic65/Qwen3.6-Whittle-25B-A3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use logic65/Qwen3.6-Whittle-25B-A3B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="logic65/Qwen3.6-Whittle-25B-A3B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("logic65/Qwen3.6-Whittle-25B-A3B") model = AutoModelForCausalLM.from_pretrained("logic65/Qwen3.6-Whittle-25B-A3B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use logic65/Qwen3.6-Whittle-25B-A3B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M # Run inference directly in the terminal: llama cli -hf logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M # Run inference directly in the terminal: llama cli -hf logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M
Use Docker
docker model run hf.co/logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use logic65/Qwen3.6-Whittle-25B-A3B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "logic65/Qwen3.6-Whittle-25B-A3B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "logic65/Qwen3.6-Whittle-25B-A3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M
- SGLang
How to use logic65/Qwen3.6-Whittle-25B-A3B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "logic65/Qwen3.6-Whittle-25B-A3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "logic65/Qwen3.6-Whittle-25B-A3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "logic65/Qwen3.6-Whittle-25B-A3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "logic65/Qwen3.6-Whittle-25B-A3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use logic65/Qwen3.6-Whittle-25B-A3B with Ollama:
ollama run hf.co/logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M
- Unsloth Desktop
- Pi
How to use logic65/Qwen3.6-Whittle-25B-A3B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use logic65/Qwen3.6-Whittle-25B-A3B with Docker Model Runner:
docker model run hf.co/logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M
- Lemonade
How to use logic65/Qwen3.6-Whittle-25B-A3B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M
Run and chat with the model
lemonade run user.Qwen3.6-Whittle-25B-A3B-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use logic65/Qwen3.6-Whittle-25B-A3B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use logic65/Qwen3.6-Whittle-25B-A3B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "logic65/Qwen3.6-Whittle-25B-A3B:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.6-Whittle-25B-A3B
☕ Support this work
Whittle is built by one person on a grocery budget and rented GPU hours. If this research is useful to you, or you want to see it finished: ko-fi.com/davida81328. Every hour of GPU time goes straight into the next checkpoint, and every checkpoint, table and log lands in these repos.
This repo holds two different things. (1) Whittle-Qwen-3.8-25B-A3B (3 Oct 2026): five GGUFs of the current Whittle flagship's body with its 10 B n-gram memory removed, described in the next section.
(2) The Qwen3.6-Whittle-25B-A3B prune+heal body (2 Sep 2026): Qwen3.6-35B-A3B with 30% of its routed experts removed (256 → 180 per layer) and the damage healed by
self-distillation from the unpruned weights — the safetensors at the root of this repo and the experimental next/ build, described further down. The two are different models in different formats; the second is the starting point the Whittle-Next line was built from.
Status
- Whittle-Qwen-3.8-25B-A3B GGUFs: current (3 Oct 2026). They derive from the current flagship Whittle-Qwen-3.8-35B-A3B (root = Phase-2 step 32010), whose GGUFs with the memory are on Whittle-Qwen-3.8-35B-A3B-GGUF.
- Qwen3.6-Whittle-25B-A3B body (2 Sep 2026): the base of the Whittle-Next line. It was converted to the
qwen4_expformat and continued as Whittle-Next-26B-A3B → Whittle-Next-27B-A3B → the 35B above. Its own numbers are below.
Whittle-Qwen-3.8-25B-A3B: the 35B body without its memory (3 Oct 2026)
These five files are the Phase-2 step 32010 body of Whittle-Qwen-3.8-35B-A3B with its 10 B n-gram memory removed: the memory table is replaced by an 8-row all-zero table, and every other tensor is byte-identical to the 35B GGUF ladder. They run on stock llama.cpp.
Which file
| file | size |
|---|---|
Whittle-Qwen-3.8-25B-A3B-Q8_0.gguf |
27.21 GB |
Whittle-Qwen-3.8-25B-A3B-Q6_K.gguf |
21.08 GB |
Whittle-Qwen-3.8-25B-A3B-Q5_K_M.gguf |
18.26 GB |
Whittle-Qwen-3.8-25B-A3B-Q4_K_M.gguf |
15.66 GB |
Whittle-Qwen-3.8-25B-A3B-Q3_K_M.gguf |
12.44 GB |
Measured (3 Oct 2026, on the published Q6_K)
Measured on Whittle-Qwen-3.8-25B-A3B-Q6_K.gguf from this repo, served by stock llama.cpp on two T4 GPUs, one sample per item. PopQA-288 and the 40 facts are cloze prompts scored on a greedy (temperature 0) raw-text completion; every other probe goes through the chat endpoint with the card sampler (temperature 0.7, top_p 0.8, top_k 20, repeat_penalty 1.05) and thinking on. The right-hand column is the 35B with its full memory table (its Q6_K), run through the same harness on 2 Oct 2026.
| probe | this repo's Q6_K (3 Oct) | 35B Q6_K, full memory table (2 Oct) |
|---|---|---|
| PopQA-288 cloze, exact match | 69/288 | 71/288 |
| 40 plain facts | 35/40 | 35/40 |
| stop/loop battery, 24 prompts | 23/24 clean (0 loops, 0 user-turn leaks) | 23/24 clean (0 loops, 0 user-turn leaks) |
| GSM-50 | 42/50 | 44/50 |
| strict-JSON hold-out, 72 items | 72/72 valid, 71/72 exact | 72/72 valid, 70/72 exact |
| tool-use probes, 96 items | 91/96 | 95/96 |
| of which the ~11k-token context probe (24 items) | 19/24 | 23/24 |
| llama.cpp web page, browser tools: a real chat replayed at two points (20 each) | 17/20 and 18/20, 0 invented tools | 16/20 and 18/20, 0 invented tools |
| llama.cpp web page, browser tools: 36 held-out chats finished correctly | 28/36 | 32/36 |
| MATH-60, 6,144-token cap | 38/60 | not run for the full-table file in that session |
Every probe is one sample per item, so differences of a few items between the two columns may be noise.
Run it
llama-server -m <file> -ngl 99 -c 16384 --jinja -fa on
The memory is gone, so -ot per_layer_token_embd=CPU is not needed. Sampler: temperature 0.7, top_p 0.8, top_k 20, repeat_penalty 1.05. Thinking is switched with "chat_template_kwargs": {"enable_thinking": true}.
Caveat
The body depends on the memory on raw text and code completion: cross-entropy on held code collapses without it. These files are for chat, instruction and agent use, not for base-model-style completion.
Qwen3.6-Whittle-25B-A3B: the 2 Sep 2026 prune+heal body (different lineage and format)
Everything below describes the earlier artefact: the qwen3_5_moe-format safetensors at the root of this repo and the next/ build. It is not the model in the GGUFs above.
Qwen3.6-35B-A3B with 30% of its routed experts removed (256 → 180 per layer) and the damage healed by
self-distillation from the unpruned weights. Same architecture (qwen3_5_moe), same tokenizer, same 8-routed + 1-shared
experts active per token. Runs on stock transformers and stock llama.cpp, no patches.
| Qwen3.6-35B-A3B | Qwen3.6-Whittle-25B-A3B | |
|---|---|---|
| total parameters (text) | 34.7B | 25.1B |
| active per token | ~3B | ~3B (unchanged) |
| routed experts / layer | 256 | 180 |
| GSM8K (200 q, no-think, greedy, 512 tok) | 88.5% | 92.5% (185/200) |
| held-out CE, corpus text (never seen in calibration/healing) | 1.770 | 1.863 (raw prune 1.926) |
| held-out CE, chat rows | 2.510 | 1.417 |
| bf16 on disk | 70 GB | 50 GB |
Which files are which
- Root:
model-00001-of-00015.safetensors…model-00015-of-00015.safetensors+model.safetensors.index.json,config.json,chat_template.jinja,generation_config.json, tokenizer files,gate_act.json— the prune+heal body inqwen3_5_moeformat. eval/—gsm8k_base_qwen3.6-35b-a3b.json,gsm8k_pruned_healed.json,prune_report.json.next/— the experimentalqwen4expbuild (see below):Qwen3.6-Whittle-25B-A3B-next-EXPERIMENTAL-Q8_0.gguf,sigmoid_settle_step1500.pt,gate_act.json,gsm8k_q4exp_stock_llamacpp.json,gsm8k_sigmoid_transformers.json.Whittle-Qwen-3.8-25B-A3B-*.gguf— the five files of the section above, a different model.
How it was made
- Score every expert on 1M calibration tokens (encyclopaedia, textbooks, maths, code, chat) with a gate-free mean-squared-activation-norm criterion (the task-agnostic winner in the June-2026 one-shot MoE pruning study).
- Prune the 76 lowest-scoring experts per layer; router rows sliced to match. Kept experts carried 82% of routed traffic on average (71% in the worst layer).
- Heal for 900 steps × 2048 tokens (≈1.8M tokens) against the unpruned model as teacher — the same weights with the mask off, so no second model and no drift: loss = 2·KL(teacher‖student) + 1·CE. Trained: routers, shared experts, rank-8 LoRA on the kept experts (merged into the exported weights). Optimiser: Muon on 2D matrices, AdamW on the rest. Data: 50% raw corpus windows, 50% template-faithful instruction rows.
- Export kept experts from the original bf16 shards + LoRA deltas (no quantisation error baked in).
Run it
# transformers ≥ 5.16
from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("logic65/Qwen3.6-Whittle-25B-A3B", dtype="bfloat16", device_map="auto")
# llama.cpp (stock): GGUF Q4_K_M in this repo
llama-server -m Qwen3.6-Whittle-25B-A3B-Q4_K_M.gguf -ngl 99 -c 8192 --jinja
Recommended sampling as the parent: temperature 0.7, top_p 0.8, top_k 20, repeat_penalty 1.05. Thinking mode works as in the parent.
Note (3 Oct 2026): Qwen3.6-Whittle-25B-A3B-Q4_K_M.gguf, named in the block above, is not in this repo's current file listing. The GGUFs in the repo are the Whittle-Qwen-3.8-25B-A3B ladder (a different model, see above) and the experimental next/ build.
Why is GSM8K higher than the parent?
Eight more correct answers out of 200; sampling noise at this size is about ±2 points, so treat it as "at least parity". Because the heal is also an SFT pass. Half of the healing batches were template-faithful instruction rows, and those include OpenR1-Math reasoning traces, so for 900 steps the model was distilled from its unpruned self and trained on worked maths answers in exactly the no-think, step-by-step format GSM8K is scored in. The parent never had that pass. It is a real gain on this task, not a general one: on raw corpus text the pruned model still sits 0.09 nats of cross-entropy above the parent (1.863 vs 1.771), which is the honest cost of removing 30% of the experts. Expect the same pattern elsewhere: strong on instruction-style tasks close to the healing data, slightly weaker on long-tail knowledge.
What to expect
Twelve heal steps in were enough for "hello" → "Hello! How can I help you today?", one-sentence physics, 17+25=42, and two-turn name recall. The pruned model scores lower CE than the parent on chat-formatted text (calibration and healing both contain chat data) and slightly higher on raw corpus text — the honest damage number is the corpus one. Expect small regressions on long-tail knowledge relative to the 35B; experts that fired rarely on the calibration mix are the ones removed.
next/ — experimental Qwen3.8-Next-format build (work in progress)
next/Qwen3.6-Whittle-25B-A3B-next-EXPERIMENTAL-Q8_0.gguf is the same pruned model converted to the qwen4exp
architecture: GDN output gate retrained from silu to sigmoid (progressive, back to front, self-distilled), identity
hyper-connections, an inert sparse-attention indexer, no n-gram memory yet. It loads and runs on unmodified upstream
llama.cpp as qwen4exp. GSM8K on it: 86.5% through llama.cpp (Q8_0), 82.0% through transformers (bf16) — the gate
conversion still costs a few points versus the 92.5% of the silu model above; conversation quality is unchanged. This file
is the base the hyper-connections and the n-gram memory are being trained into; treat it as a preview, not a release.
next/sigmoid_settle_step1500.pt holds the trainable state that produced it.
Support this work
Whittle runs on one hobbyist's grocery budget and rented GPU hours. If this research is useful to you: ko-fi.com/davida81328 ☕
Authors
David Aylward (logic65) & Claude (Anthropic) — designed, debugged and verified together in one day on a single rented GPU.
Provenance
Parent: Qwen/Qwen3.6-35B-A3B (Apache-2.0). Method references: REAP
(Cerebras, arXiv 2510.13999) and "How to Score Experts for One-Shot MoE Expert Pruning" (arXiv 2606.15716).
Scripts: logic65/mini-next-a100-kit/colab/ (prune_qwen36.py, heal_qwen36.py, run_h100.sh). Built on one RTX PRO 6000
Blackwell in about five hours. Part of the Whittle project by logic65.
The Whittle-Qwen-3.8-25B-A3B GGUFs above inherit the 35B's lineage: parent logic65/Whittle-Next-27B-A3B, teacher Qwen/Qwen3.8-27B, base lineage Qwen3.6-35B-A3B pruned 256 → 180 experts. All Apache-2.0.
- Downloads last month
- 1,648