Whittle

Whittle-Qwen-3.8-45B-A3B

☕ Support this work

Whittle is built by one person on a grocery budget and rented GPU hours. If this research is useful to you, or you want to see it finished: ko-fi.com/davida81328. Every hour of GPU time goes straight into the next checkpoint, and every checkpoint, table and log lands in these repos.

The bf16 safetensors of Whittle-Qwen-3.8-45B-A3B: the Whittle flagship 35B with the 76 experts per layer that its expert prune removed put back from Qwen3.6-35B-A3B, untrained. 45.1 B total (35.1 B body + 10.0 B n-gram memory), ~3 B active (8 of 256 routed experts + the shared expert), Qwen3.8-Flash-Next (qwen4_exp) format. For llama.cpp use the GGUF repo: it is much faster locally, and every measurement below was taken on it. Try it in the browser: Whittle 45B demo (ZeroGPU).

Status

Experimental (5 Oct 2026). An untrained splice, measured on the same battery as the 35B and its base (see the GGUF card). The restored experts were trained for Qwen3.6's own body and have not been healed into this one; arithmetic word problems are the visible cost.

Load it with transformers

Stock transformers (>= 5.18) has the architecture, but plain from_pretrained would run this checkpoint with a random n-gram table: Whittle's table uses 8 hash heads of exactly 4,880,000 rows (the geometry its GGUFs give llama.cpp), where the stock code sizes heads as primes, and it is stored under a different key prefix. whittle_load.py in this repo sizes the table from the checkpoint's own hash buffers, maps the keys onto the stock layout and checks the buffers after loading. On a small random model saved in this layout it reproduces the original's logits exactly.

import importlib.util, torch
from huggingface_hub import hf_hub_download
from transformers import AutoTokenizer

repo = "logic65/Whittle-Qwen-3.8-45B-A3B"
spec = importlib.util.spec_from_file_location("whittle_load", hf_hub_download(repo, "whittle_load.py"))
whittle_load = importlib.util.module_from_spec(spec); spec.loader.exec_module(whittle_load)

model = whittle_load.load_whittle(repo, dtype=torch.bfloat16, device_map="auto")   # the ~20 GB table stays in CPU memory
tokenizer = AutoTokenizer.from_pretrained(repo)
messages = [{"role": "user", "content": "Explain hash collisions in two sentences."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_dict=True, return_tensors="pt",
                                       enable_thinking=False).to(model.device)
out = model.generate(**inputs, max_new_tokens=256, do_sample=True, temperature=0.7, top_p=0.8, top_k=20, repetition_penalty=1.05)
print(tokenizer.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
  • Memory: about 70 GB of weights on the GPU(s) and the 20 GB table in CPU memory.
  • Install flash-linear-attention for the gated-DeltaNet layers; without it transformers falls back to a much slower reference path.
  • Sampler as the GGUF card: temperature 0.7, top_p 0.8, top_k 20, repetition penalty 1.05; do not decode greedily. Thinking: enable_thinking=True in the chat template, with 4k+ new tokens for code and 8k+ for maths.

Measured

On the GGUF Q4_K_M, same harness as the 35B and Qwen3.6-35B-A3B (full table and method on the GGUF card): HumanEval / HumanEval+ 75.6 / 72.6, MBPP / MBPP+ 77.8 / 65.9, tool-use probes 96/96 (long-context 24/24), MATH-60 44/60, long-tail names 111/252 (35B: 94), GSM-50 38/50 (35B: 44 - a real drop, 6 lost and none gained on the same problems).

How it was made

The same recipe as the GGUF, on the original tensors: the 35B root's safetensors (Phase-2 step 32010) with, in every layer, Qwen3.6-35B-A3B's bf16 weights of the 76 pruned experts appended as experts 180-255 (sorted original order; keep-lists in eval/prune_report.json) and their router rows appended, scaled per layer by the median norm ratio of the trained to the original kept rows (1.001-1.026, median 1.009). The kept experts still sit at cosine 0.982-0.997 to their Qwen3.6 originals (gate, up and down projections, checked at layers 0, 10, 20, 30 and 39) and the trained router rows at 0.978-0.995, so the restored experts meet the body in the space they were trained for. num_experts 180 -> 256; nothing else changed and nothing was trained. The 35B root files also hold 80 stale duplicate copies of router-type tensors that the index does not reference; transformers loads the indexed copies (checked).

Provenance

David Aylward (logic65) & Claude (Anthropic). Built from logic65/Whittle-Qwen-3.8-35B-A3B (teacher Qwen/Qwen3.8-27B, memory contents from Qwen/Qwen3.8-Flash-Next) and the routed experts of Qwen/Qwen3.6-35B-A3B. All Apache-2.0.

Downloads last month
279
Safetensors
Model size
45B params
Tensor type
BF16
·
I64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for logic65/Whittle-Qwen-3.8-45B-A3B

Space using logic65/Whittle-Qwen-3.8-45B-A3B 1