Qwen3.6-Whittle-25B-A3B

☕ Support this work

Whittle is built by one person on a grocery budget and rented GPU hours. If this research is useful to you, or you want to see it finished: ko-fi.com/davida81328. Every hour of GPU time goes straight into the next checkpoint, and every checkpoint, table and log lands in these repos.

This repo holds two different things. (1) Whittle-Qwen-3.8-25B-A3B (3 Oct 2026): five GGUFs of the current Whittle flagship's body with its 10 B n-gram memory removed, described in the next section. (2) The Qwen3.6-Whittle-25B-A3B prune+heal body (2 Sep 2026): Qwen3.6-35B-A3B with 30% of its routed experts removed (256 → 180 per layer) and the damage healed by self-distillation from the unpruned weights — the safetensors at the root of this repo and the experimental next/ build, described further down. The two are different models in different formats; the second is the starting point the Whittle-Next line was built from.

Status

Whittle-Qwen-3.8-25B-A3B: the 35B body without its memory (3 Oct 2026)

These five files are the Phase-2 step 32010 body of Whittle-Qwen-3.8-35B-A3B with its 10 B n-gram memory removed: the memory table is replaced by an 8-row all-zero table, and every other tensor is byte-identical to the 35B GGUF ladder. They run on stock llama.cpp.

Which file

file size
Whittle-Qwen-3.8-25B-A3B-Q8_0.gguf 27.21 GB
Whittle-Qwen-3.8-25B-A3B-Q6_K.gguf 21.08 GB
Whittle-Qwen-3.8-25B-A3B-Q5_K_M.gguf 18.26 GB
Whittle-Qwen-3.8-25B-A3B-Q4_K_M.gguf 15.66 GB
Whittle-Qwen-3.8-25B-A3B-Q3_K_M.gguf 12.44 GB

Measured (3 Oct 2026, on the published Q6_K)

Measured on Whittle-Qwen-3.8-25B-A3B-Q6_K.gguf from this repo, served by stock llama.cpp on two T4 GPUs, one sample per item. PopQA-288 and the 40 facts are cloze prompts scored on a greedy (temperature 0) raw-text completion; every other probe goes through the chat endpoint with the card sampler (temperature 0.7, top_p 0.8, top_k 20, repeat_penalty 1.05) and thinking on. The right-hand column is the 35B with its full memory table (its Q6_K), run through the same harness on 2 Oct 2026.

probe this repo's Q6_K (3 Oct) 35B Q6_K, full memory table (2 Oct)
PopQA-288 cloze, exact match 69/288 71/288
40 plain facts 35/40 35/40
stop/loop battery, 24 prompts 23/24 clean (0 loops, 0 user-turn leaks) 23/24 clean (0 loops, 0 user-turn leaks)
GSM-50 42/50 44/50
strict-JSON hold-out, 72 items 72/72 valid, 71/72 exact 72/72 valid, 70/72 exact
tool-use probes, 96 items 91/96 95/96
of which the ~11k-token context probe (24 items) 19/24 23/24
llama.cpp web page, browser tools: a real chat replayed at two points (20 each) 17/20 and 18/20, 0 invented tools 16/20 and 18/20, 0 invented tools
llama.cpp web page, browser tools: 36 held-out chats finished correctly 28/36 32/36
MATH-60, 6,144-token cap 38/60 not run for the full-table file in that session

Every probe is one sample per item, so differences of a few items between the two columns may be noise.

Run it

llama-server -m <file> -ngl 99 -c 16384 --jinja -fa on

The memory is gone, so -ot per_layer_token_embd=CPU is not needed. Sampler: temperature 0.7, top_p 0.8, top_k 20, repeat_penalty 1.05. Thinking is switched with "chat_template_kwargs": {"enable_thinking": true}.

Caveat

The body depends on the memory on raw text and code completion: cross-entropy on held code collapses without it. These files are for chat, instruction and agent use, not for base-model-style completion.

Qwen3.6-Whittle-25B-A3B: the 2 Sep 2026 prune+heal body (different lineage and format)

Everything below describes the earlier artefact: the qwen3_5_moe-format safetensors at the root of this repo and the next/ build. It is not the model in the GGUFs above.

Qwen3.6-35B-A3B with 30% of its routed experts removed (256 → 180 per layer) and the damage healed by self-distillation from the unpruned weights. Same architecture (qwen3_5_moe), same tokenizer, same 8-routed + 1-shared experts active per token. Runs on stock transformers and stock llama.cpp, no patches.

Qwen3.6-35B-A3B Qwen3.6-Whittle-25B-A3B
total parameters (text) 34.7B 25.1B
active per token ~3B ~3B (unchanged)
routed experts / layer 256 180
GSM8K (200 q, no-think, greedy, 512 tok) 88.5% 92.5% (185/200)
held-out CE, corpus text (never seen in calibration/healing) 1.770 1.863 (raw prune 1.926)
held-out CE, chat rows 2.510 1.417
bf16 on disk 70 GB 50 GB

Which files are which

  • Root: model-00001-of-00015.safetensors … model-00015-of-00015.safetensors + model.safetensors.index.json, config.json, chat_template.jinja, generation_config.json, tokenizer files, gate_act.json — the prune+heal body in qwen3_5_moe format.
  • eval/ — gsm8k_base_qwen3.6-35b-a3b.json, gsm8k_pruned_healed.json, prune_report.json.
  • next/ — the experimental qwen4exp build (see below): Qwen3.6-Whittle-25B-A3B-next-EXPERIMENTAL-Q8_0.gguf, sigmoid_settle_step1500.pt, gate_act.json, gsm8k_q4exp_stock_llamacpp.json, gsm8k_sigmoid_transformers.json.
  • Whittle-Qwen-3.8-25B-A3B-*.gguf — the five files of the section above, a different model.

How it was made

  1. Score every expert on 1M calibration tokens (encyclopaedia, textbooks, maths, code, chat) with a gate-free mean-squared-activation-norm criterion (the task-agnostic winner in the June-2026 one-shot MoE pruning study).
  2. Prune the 76 lowest-scoring experts per layer; router rows sliced to match. Kept experts carried 82% of routed traffic on average (71% in the worst layer).
  3. Heal for 900 steps × 2048 tokens (≈1.8M tokens) against the unpruned model as teacher — the same weights with the mask off, so no second model and no drift: loss = 2·KL(teacher‖student) + 1·CE. Trained: routers, shared experts, rank-8 LoRA on the kept experts (merged into the exported weights). Optimiser: Muon on 2D matrices, AdamW on the rest. Data: 50% raw corpus windows, 50% template-faithful instruction rows.
  4. Export kept experts from the original bf16 shards + LoRA deltas (no quantisation error baked in).

Run it

# transformers ≥ 5.16
from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("logic65/Qwen3.6-Whittle-25B-A3B", dtype="bfloat16", device_map="auto")
# llama.cpp (stock): GGUF Q4_K_M in this repo
llama-server -m Qwen3.6-Whittle-25B-A3B-Q4_K_M.gguf -ngl 99 -c 8192 --jinja

Recommended sampling as the parent: temperature 0.7, top_p 0.8, top_k 20, repeat_penalty 1.05. Thinking mode works as in the parent.

Note (3 Oct 2026): Qwen3.6-Whittle-25B-A3B-Q4_K_M.gguf, named in the block above, is not in this repo's current file listing. The GGUFs in the repo are the Whittle-Qwen-3.8-25B-A3B ladder (a different model, see above) and the experimental next/ build.

Why is GSM8K higher than the parent?

Eight more correct answers out of 200; sampling noise at this size is about ±2 points, so treat it as "at least parity". Because the heal is also an SFT pass. Half of the healing batches were template-faithful instruction rows, and those include OpenR1-Math reasoning traces, so for 900 steps the model was distilled from its unpruned self and trained on worked maths answers in exactly the no-think, step-by-step format GSM8K is scored in. The parent never had that pass. It is a real gain on this task, not a general one: on raw corpus text the pruned model still sits 0.09 nats of cross-entropy above the parent (1.863 vs 1.771), which is the honest cost of removing 30% of the experts. Expect the same pattern elsewhere: strong on instruction-style tasks close to the healing data, slightly weaker on long-tail knowledge.

What to expect

Twelve heal steps in were enough for "hello" → "Hello! How can I help you today?", one-sentence physics, 17+25=42, and two-turn name recall. The pruned model scores lower CE than the parent on chat-formatted text (calibration and healing both contain chat data) and slightly higher on raw corpus text — the honest damage number is the corpus one. Expect small regressions on long-tail knowledge relative to the 35B; experts that fired rarely on the calibration mix are the ones removed.

next/ — experimental Qwen3.8-Next-format build (work in progress)

next/Qwen3.6-Whittle-25B-A3B-next-EXPERIMENTAL-Q8_0.gguf is the same pruned model converted to the qwen4exp architecture: GDN output gate retrained from silu to sigmoid (progressive, back to front, self-distilled), identity hyper-connections, an inert sparse-attention indexer, no n-gram memory yet. It loads and runs on unmodified upstream llama.cpp as qwen4exp. GSM8K on it: 86.5% through llama.cpp (Q8_0), 82.0% through transformers (bf16) — the gate conversion still costs a few points versus the 92.5% of the silu model above; conversation quality is unchanged. This file is the base the hyper-connections and the n-gram memory are being trained into; treat it as a preview, not a release. next/sigmoid_settle_step1500.pt holds the trainable state that produced it.

Support this work

Whittle runs on one hobbyist's grocery budget and rented GPU hours. If this research is useful to you: ko-fi.com/davida81328 ☕

Authors

David Aylward (logic65) & Claude (Anthropic) — designed, debugged and verified together in one day on a single rented GPU.

Provenance

Parent: Qwen/Qwen3.6-35B-A3B (Apache-2.0). Method references: REAP (Cerebras, arXiv 2510.13999) and "How to Score Experts for One-Shot MoE Expert Pruning" (arXiv 2606.15716). Scripts: logic65/mini-next-a100-kit/colab/ (prune_qwen36.py, heal_qwen36.py, run_h100.sh). Built on one RTX PRO 6000 Blackwell in about five hours. Part of the Whittle project by logic65.

The Whittle-Qwen-3.8-25B-A3B GGUFs above inherit the 35B's lineage: parent logic65/Whittle-Next-27B-A3B, teacher Qwen/Qwen3.8-27B, base lineage Qwen3.6-35B-A3B pruned 256 → 180 experts. All Apache-2.0.

Downloads last month
1,648
Safetensors
Model size
25B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for logic65/Qwen3.6-Whittle-25B-A3B

Quantized
(863)
this model
Quantizations
2 models