Qwen3.8-Whittle-v40 (16.8B) best measured cut

☕ Support this work

Whittle is built by one person on a grocery budget and rented GPU hours. If this research is useful to you, or you want to see it finished: ko-fi.com/davida81328. Every hour of GPU time goes straight into the next checkpoint, and every checkpoint, table and log lands in these repos.

40 layers of Qwen3.8-27B, kept in the exact [gdn, gdn, gdn, attention] pattern llama.cpp requires: 30 gated-DeltaNet layers, 10 full-attention stations, a true 3:1 ratio. Which layers survive was decided by a per-layer identity map (cosine between each layer's input and output residual stream over 80 probes): within every span between surviving attention stations, the three most load-bearing recurrent layers are kept and the rest are dropped.

Status

Earlier line (Aug 2026); the current Whittle models are the Whittle-Next / Whittle-Qwen-3.8 line: Whittle-Qwen-3.8-35B-A3B.

Research preview, no training: smaller than every other cut in the family and better than all of them on the knowledge-atlas measurement below; chosen by measurement, not by assumption. It needs post training and is published as the record of a compression method and its measurements rather than as a finished assistant. Further healing and instruction tuning runs are planned over time, and these checkpoints will improve as those land.

Which weights are which

path what it is
q38_v40_q8.gguf the model, Q8_0 GGUF (the only weights in this repo)
recipe/q38_build_keep_hf.py, recipe/variant_drops.json the build script and the variant layer-drop table that produced it

Measured

Measured with an 80-probe, 8-domain knowledge atlas, scored as the mean fraction of the intact model's next-token probability retained:

shape params layers retention
v40 (this) 16.8B 40 0.648
44L cut (not currently published) 19.2B 44 0.408
same size, attention crowded early 16.8B 40 0.379
same rule, 36 layers 15.1B 36 0.441

Two lessons are in that table. Choosing layers by measurement instead of removing contiguous bands is worth more than 2 billion parameters. And where the attention stations sit matters more than how many there are: three variants at identical size and identical selection rule score 0.648, 0.544 and 0.379 purely by station placement.

Against the intact 27B, this cut is at or above the original on agentic control (0.0922 vs 0.0885), Python (0.0030 vs 0.0023) and science (0.0035 vs 0.0034), holds 90% on math, and keeps 58% of multilingual recall where the 44L cut kept 3%. Its end-of-turn probability at a finished answer, the quantity that governs looping, stays at 0.46 and 0.44 versus the original's 0.53 and 0.33.

How it was made

No training. The cut is the identity-guided layer selection described at the top of this card, built with the recipe in recipe/.

Run it

Anti-loop sampling is still recommended: --dry-multiplier 0.8 --dry-base 1.75 --dry-allowed-length 4 --repeat-penalty 1.15

Caveats

No repair training has been done. It is a research preview: expect dulled confidence on hard recall and the family's known rough edges. C++ recall is the weakest domain, being distributed across layers no cut preserves.

Support this work

Independent research on consumer hardware. Every donation becomes A100 hours, and every A100 hour ends up as a public model or a public measurement. ☕ ko-fi.com/davida81328

Provenance, licence and authors

Base model by the Qwen team (Apache 2.0). Cut and measured by David Aylward with Claude (Fable 5, Anthropic) as co-author.

Downloads last month
235
GGUF
Model size
18B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for logic65/Qwen3.8-Whittle-v40-16.8B

Base model

Qwen/Qwen3.8-27B
Quantized
(10)
this model

Collections including logic65/Qwen3.8-Whittle-v40-16.8B