Instructions to use logic65/Whittle-Qwen-3.8-35B-A3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use logic65/Whittle-Qwen-3.8-35B-A3B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="logic65/Whittle-Qwen-3.8-35B-A3B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("logic65/Whittle-Qwen-3.8-35B-A3B") model = AutoModelForCausalLM.from_pretrained("logic65/Whittle-Qwen-3.8-35B-A3B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use logic65/Whittle-Qwen-3.8-35B-A3B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "logic65/Whittle-Qwen-3.8-35B-A3B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "logic65/Whittle-Qwen-3.8-35B-A3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/logic65/Whittle-Qwen-3.8-35B-A3B
- SGLang
How to use logic65/Whittle-Qwen-3.8-35B-A3B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "logic65/Whittle-Qwen-3.8-35B-A3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "logic65/Whittle-Qwen-3.8-35B-A3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "logic65/Whittle-Qwen-3.8-35B-A3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "logic65/Whittle-Qwen-3.8-35B-A3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use logic65/Whittle-Qwen-3.8-35B-A3B with Docker Model Runner:
docker model run hf.co/logic65/Whittle-Qwen-3.8-35B-A3B
Whittle-Qwen-3.8-35B-A3B
☕ Support this work
Whittle is built by one person on a grocery budget and rented GPU hours. If this research is useful to you, or you want to see it finished: ko-fi.com/davida81328. Every hour of GPU time goes straight into the next checkpoint, and every checkpoint, table and log lands in these repos.
A 35.1 B-parameter, ~3 B-active mixture-of-experts in the Qwen3.8-Flash-Next (qwen4_exp) format whose 10 B-parameter hashed n-gram memory is
load-bearing — the first Whittle where zeroing the memory measurably hurts the model. Distilled from Qwen3.8-27B (thinking on) at the logit level on
25,000 verified teacher traces (published with their top-20 logprobs), built on
Whittle-Next-27B-A3B v4.4, with memory contents transferred from
Qwen3.8-Flash-Next's own table. Runs on stock llama.cpp, no patches.
Status
Current. This is the current Whittle flagship; the root of this repo is Phase-2 step 32010 (p2-s32010, published 30 Sep 2026). GGUFs (Q8_0 → Q3_K_M) are on logic65/Whittle-Qwen-3.8-35B-A3B-GGUF. Parent: Whittle-Next-27B-A3B v4.4. Teacher: Qwen3.8-27B.
Research preview. This model uses the Qwen3.8-Flash-Next architecture (qwen4_exp: hyper-connections, gated DeltaNet + attention, the 10 B
n-gram memory), and its memory rows come from Qwen3.8-Flash-Next's own table. Phase 2 is the broad distillation: 25,178 complete, graded Qwen3.8-27B answers across maths,
code, chat, strict JSON and executed agent episodes, each with the teacher's top-20 distribution at every token. The root has trained on all 22,950 rows that fit
its 4k window, once; the 1,470 long-context rows are not used yet, and the on-policy phase has not started. It should not be compared
with Qwen3.8 on equal terms. Read Measured and Caveats before relying on it.
This model is the first in the line whose memory carries knowledge; the next steps — more memory rows, fact-dense training data, the on-policy distillation — are compute we cannot currently pay for (see the support note above).
Which Whittle to run: the trade-offs (4 Oct 2026)
Four Whittle builds now share one body, and each trades size for a different skill. We ran all four, and the Qwen3.6-35B-A3B they were cut from, through one battery: stock llama.cpp on two T4 GPUs, the sampler from Run it, thinking on, one sample per item.
| Whittle 25B | Whittle 35B (this repo) | Whittle x256 | Whittle 45B | Qwen3.6-35B-A3B | |
|---|---|---|---|---|---|
| built as | pruned body, no table | 25B + the 10 B table | 25B + 76 experts per layer put back | this 35B + 76 experts per layer put back | the base |
| Q4_K_M file | 15.7 GB | 21.3 GB (table 5.6) | 21.5 GB | 27.1 GB (table 5.6) | 22.3 GB |
| files | 25B repo | GGUF repo | dev branch, experimental |
45B GGUF repo, experimental | Qwen/Qwen3.6-35B-A3B |
| tested as | Q6_K | Q6_K (EvalPlus Q5_K_M) | Q4_K_M | Q4_K_M | Q4_K_M (bartowski) |
| HumanEval / HumanEval+ | 53.0 / 51.2 | 73.8 / 70.7 | not run | 75.6 / 72.6 | not run |
| MBPP / MBPP+ | 72.2 / 60.8 | 76.2 / 61.9 | not run | 77.8 / 65.9 | not run |
| long-context tool use (24) | 19 | 23 | 19 | 24 | 24 |
| tool-use probes (96) | 91 | 95 | 91 | 96 | 93 |
| web page: chat replay (40) | 35 | 34 | 33 | 27 | 40 |
| web page: held-out chats (36) | 28 | 32 | 32 | 31 | 32 |
| GSM-50 | 42 | 44 | 45 | 38 | 40 |
| MATH-60 | 38 | not in this run | 41 | 44 | 38 |
| MATH-60, mean tokens per answer | 1,976 | — | 1,959 | 1,690 | 4,018 |
| long-tail names, PopQA (252) | 97 | 94 | 111 | 111 | 132 |
| 40 plain facts | 35 | 35 | 37 | 37 | 33 |
| strict JSON, exact (72) | 71 | 70 | 72 | 70 | 72 |
| stop/loop battery, clean (24) | 23 | 23 | 24 | 24 | 24 |
The table buys code and long-context work. Added to the 25B body it lifts HumanEval from 53.0 to 73.8 (46 tasks gained, 12 lost), and long-context
tool use went from 19/24 to 23–24/24 in both pairs (10 items gained, 1 lost over the two). On its own it does not move long-tail names or facts. It costs
5.6 GB at Q4_K_M, and with -ot per_layer_token_embd=CPU that is system RAM, not VRAM: this 35B needs the same GPU memory as the 25B.
The restored experts buy long-tail knowledge. Putting the 76 pruned experts per layer back gains 17 and loses 3 of the 252 name questions on the 25B body, and gains 18 and loses 1 on this one; on the 25B body nothing else moves beyond noise. They add 5.8 GB at Q4_K_M, and they went in untrained (router rows rescaled, no healing), so they have not been through this repo's web-page tool fix: the x256 and the 45B are the only Whittles that invented a tool call in the replay (2 of 40 each).
Both together, the 45B, is the best Whittle on tool use (96/96) and MATH-60 (44/60), and joint best on names, but it is the only one that got worse at GSM-50: it lost 7 problems the x256 solved and gained none (exact test, p = 0.016), and 6 against this 35B (p = 0.031). It codes at least as well as this 35B: HumanEval 75.6 / 72.6 and MBPP 77.8 / 65.9 at Q4_K_M against 73.8 / 70.7 and 76.2 / 61.9 at Q5_K_M; on the same tasks that is 24 gained and 21 lost on HumanEval (noise) and 34 against 19 on MBPP+ (p = 0.053), so the restored experts cost no coding.
Against the base, Qwen3.6-35B-A3B still knows more long-tail names (132 against 111 at best; p = 0.004 against the 45B) and never stumbled in the web-page replay (40/40). The Whittles measured here think about half as long (1,690–1,976 tokens per MATH-60 answer against 4,018), so they run out of room less often (GSM-50: 0–4 answers at the token limit against 9).
Which to run
- Smallest: the 25B (15.7 GB at Q4_K_M), if you do not need code.
- Code and agents on the least memory: this 35B. The same GPU memory as the 25B plus 5.6 GB of system RAM for the table, and coding within noise of the 45B.
- Knowledge, maths and code, with room for a 21.5 GB body: the 45B, experimental. It codes at least as well as this 35B and knows more long-tail names; expect weaker arithmetic word problems and, on llama.cpp's web page, the occasional invented tool call, until it has been healed. Serve it like this repo's GGUFs, table on the CPU.
- Most long-tail knowledge: the base Qwen3.6-35B-A3B still knows more names than any Whittle.
How to read these. Single runs: one sample per item, so a few items either way on GSM-50, MATH-60 or the chats is noise; the first chart marks which changes clear an exact test (paired on the same items where the test allows). The x256, the 45B and the base ran at Q4_K_M and the 25B and 35B at Q6_K (the base's file has an importance matrix, our x256 and 45B files do not), so the restored-experts comparison also carries a quantisation change. The base ran with Whittle's sampler rather than Qwen's recommended temperature 0.6 / top-p 0.95. These numbers come from a separate run on different hardware from the Measured tables below, so this 35B's cells differ from them by a few items; MATH-60 was not in this run (the Q8_0 bench below has it at 46/60).
Which weights are which (30 Sep 2026)
The root is Phase-2 step 32010 (p2-s32010): the complete Phase-2 distillation (every one of the 22,950 teacher rows
that fit its window, seen once), then a memory phase that trained only the n-gram table with the body frozen (to step 31600), then a 410-step fix for the llama.cpp web
page's browser tools (see Web-page tools below). Its predecessors are preserved unchanged: p2-s6000 under bf16-s6000/, agentfix2 under bf16-agentfix2/,
lw5 under bf16-lw5/, lw2 under bf16-lw2/ and the first release tbl1 under bf16-tbl1/. Every Phase-2 checkpoint and memory table is on the
dev branch under p2/ and p2tf/.
This repo holds the full weights (model-*.safetensors, 35.1 B incl. the memory), the config/tokenizer, and every probe in eval/.
What is in the box
| this model | |
|---|---|
| total parameters | 35.1 B = 25.1 B body + 10.0 B n-gram memory |
| active per token | ~3 B (8 of 180 routed experts + shared expert; the memory is a lookup) |
| memory | 8 hash heads × 4,880,000 rows × 256, bigram + trigram, injected before layer 2; rows initialised from exact Qwen3.8-Flash-Next row-sets per visited bucket and trained since; unvisited buckets stay zero |
| body | 40 layers, hidden 2048, GDN + full attention, hyper-connection residual streams (count 4) |
| context | 262k positions (as the parent); tested to 75k |
| teacher | Qwen3.8-27B, thinking on, complete traces + per-position distributions: top-20 log-probabilities for Phase 2 (25k traces), top-128 for the earlier 1,840-trace stages |
Measured
All numbers are ours, with the same scripts. Per-item replies and logs are in eval/ up to p2-s6000 (eval/p2_s6000/) and, for Phase-2 step 31600, on the dev branch under
p2/eval-s31600/. The p2-s32010 bench's per-item replies were lost with the Colab VM that ran them (its disk filled during an upload the same afternoon); its summary
lines were recorded as it ran and are what the table quotes, and the published Q8_0 was spot-checked separately (note under the table).
Behavioural probes, served as the Q8_0 GGUF
Served as the Q8_0 GGUF on stock llama.cpp, memory in RAM, serving sampler, thinking on, one sample per item (3× RTX 3060 for the first four columns; p2-s32010 and step 31600 on one RTX PRO 6000, see the note under the table):
| probe | lw2 | lw5 | agentfix2 | p2-s6000 | p2-s32010 (root) |
|---|---|---|---|---|---|
| stop/loop battery, 24 prompts | 24/24 clean | 23/24 | 24/24 clean | 24/24 clean | 23/24 (0 loops, 0 self-turn leaks; one limerick ran into the 8k cap) |
| GSM8K test, 50 unseen problems | 41/50 | 41/50 | 40/50 | 48/50 | 43/50 |
| MATH-60 (MATH train L2–4, in no training set), 6,144-token cap | 44/60 | 44/60 | 41/60 | 44/60 | 46/60 (L2 19, L3 17, L4 10; 8 at the cap) |
| MATH-60, the capped problems re-run with 16,384 tokens | — | — | — | 48/60 | — |
| strict-JSON hold-out: 72 new items, 6 schemas (2 never trained), bare JSON demanded | — | 7/72 bare (65 fenced), 60/72 correct content | — | 72/72 valid, 71/72 exact | 72/72 valid, 72/72 exact |
| tool use: stops after a successful Write, ~2k / ~2.5k / ~11k-token contexts (24 each) | — | 14 / 13 / 7 | 21 / 20 / 20 | 24 / 24 / 24 | 24 / 24 / 24 |
| tool use: the extraction step before that Write, Claude Code-sized context | — | 22/24 | 22/24 | 24/24 | 24/24 |
| llama.cpp web page, browser tools: a real chat replayed at "make me a website showing this" / after the tool's "browser-only" note, answers in text (20 each) | — | — | — | 12/20 / 6/20 | 19/20 / 17/20 |
| llama.cpp web page, browser tools: 36 held-out chats finished correctly (content 18, save/run 9, time/OS questions 9) | — | — | — | — | 31/36 (14 / 8 / 9) |
| knowledge: 40 plain facts (science, code, general), greedy completion contains the answer | — | — | — | — | 34/40 |
| long-context gate, 5 seeds × 6 levels, strict JSON | 19.6/36 (16–25) | 13.8/36 (0–26) | — | — | — |
| same gate, reading only: score when the answer parsed | 4.90 / 6 | 4.93 / 6 | — | — | — |
p2-s32010 and the step it was fixed from (p2 step 31600) were run on the same harness on 30 Sep, served from a Colab conversion on one RTX PRO 6000; the published GGUFs are a second conversion of the same checkpoint (GGUF conversion is not byte-reproducible here: three builds of identical inputs gave three file hashes with bit-identical exported tensors), spot-checked on the published Q8_0: web-page replay 19/20 and 16/20, 0 invented tools, bug-report JSON 12/12 under the same grader as the 72/72. Step 31600 scored 23/24, 45/50, 39/60, 72/72 exact and 96/96, and 2/18 on the held-out web-page content chats. Paired on the same items, p2-s32010 solved 10 MATH-60 problems step 31600 missed and missed 3 it solved (exact McNemar p = 0.09); on GSM-50 the split is 2 against 4 (p = 0.69). lw5, agentfix2 and p2-s6000 were run back to back on one harness (27–28 Sep); the lw2 column and the long-context rows come from lw2's and lw5's release runs (lw5 scored 24/24 and 42/50 there). MATH-60 on earlier checkpoints: tbl1 46/60, v4.4 48/60, v4.3 43/60 — a handful of problems on a 60-problem probe, to be read as such.
The long-context gate
Read that long-context row carefully, because we nearly published a wrong number. The gate asks six questions about a real pull request with the diff plus growing amounts of the repository as context, and demands a bare JSON object. It scores all-or-nothing per level: one unquoted value and six correct answers score zero. A single run of it is close to a coin flip — the same checkpoint scored 30/36 and 17/36 on consecutive runs of identical prompts. Across five seeds, reading is indistinguishable between lw2 and lw5 (4.90 vs 4.93 of 6 when the output parses); what differs is how often the output is valid JSON. Earlier versions of this card quoted a single lucky run; these are means with their ranges. If you need structured output from this model, constrain it with a JSON schema at serve time rather than trusting it to punctuate.
The production Qwen3.6-35B-A3B reviewer we run scores 3/3/2/4/3/3 on the single-sample version of the same gate.
The memory carries knowledge
"Memory gain" = held-out cross-entropy with the memory zeroed minus with it on, on rows never trained on; positive means the body needs the table. Every previous Whittle sat at ±0.003 (the experts had learned around the memory). Relying on the table is the design: the 256→180 expert carve removed knowledge, and the memory is its replacement — so the table must always be served whole.
| held-out rows | tbl1 | lw2 | lw5 | agentfix2 |
|---|---|---|---|---|
| code (unseen files) | +2.12 | +2.84 | +10.79 | +4.93 |
| general text | +0.22 | +0.67 | +3.90 | +1.69 |
| chat rows | +0.017 | +0.08 | +2.99 | +1.02 |
| science cards | +0.03 | +0.41 | +1.44 | +0.18 |
p2-s6000 has not been through this routine yet; it runs on the final Phase-2 checkpoint (see The memory in Phase 2 below for the in-run trend). agentfix2 gave back about half of lw5's gain in 300 steps because the dependence term was already satisfied on every corpus row (gap about +6 nats against a margin of 1.0); Phase 2 raised the margin to 7.5.
The jump at lw5 came from one constant. The training term that rewards routing knowledge through the memory compares cross-entropy with the memory off against on, and only applies when the difference falls below a margin. That margin was 0.3 nats while the actual difference ran 1.9 to 5.4, so the term had been satisfied on essentially every row and had stopped doing anything after the first few hundred steps. Raising it to 1.0 keeps it acting on the weakest tenth of rows for the whole run.
The memory's output is ~0.28× the residual norm at layer 2; the difference from earlier versions is that the layers above now use it. Cross-entropy with the memory did not degrade as the gain grew: on held code it is 1.2511 at lw5 against 1.257 at lw2, so the body is leaning on the table rather than being hollowed out.
Teacher parity
Logit-lens agreement of the student's readout with the teacher's on unseen 512-token rows: 8 rows, 3,840 positions per pair, the routine every column here used; p2-s32010 measured 30 Sep:
| pair | v4.4 | tbl1 | lw2 | lw5 | p2-s32010 (root) |
|---|---|---|---|---|---|
| student L39 ← teacher L63, top-1 agree / teacher-mass | 87.4 % / 97.2 % | 87.5 % / 97.2 % | 87.4 % / 97.2 % | 87.9 % / 97.3 % | 88.4 % / 97.4 % |
| student L35 ← teacher L59 | 21.2 % / 49.5 % | 40.3 % / 68.7 % | 44.5 % / 71.2 % | 43.0 % / 71.0 % | 33.5 % / 64.6 % |
On the full 400-row cache (192,000 positions per pair, run as two exact halves of 200 rows) p2-s32010 scores 84.8 % / 96.9 % at layer 39 and 32.7 % / 64.3 % at
layer 35 — the same as step 31600 before the fix (84.8 % / 96.9 %; 33.2 % / 64.8 %) and as step 18800 (84.7 % / 96.9 %; 33.5 % / 64.8 %): neither the second half of the
distillation nor the tool fix moved it. Layer 35 sits below lw5 because Phase 2 dropped the layer-35 steer (lw5's 43 % was that steer's work, and it cost arithmetic; step 5
of How it was made below). The result files are on the dev branch under eval-s32010/parity/.
Held-out cross-entropy
Held-out CE with the memory on: tbl1 chat 1.2098 / general 2.1141; lw2 1.2131 / 2.1304; lw5 1.2055 / 2.1299; agentfix2 1.2027 / 2.1271; p2-s6000 1.2262 / 2.1143; p2 step 31600 1.2847 / 2.1332 (the Phase-2 run's own eval, which scored lw5 1.2041 / 2.1272 at its start; v4.4: 1.2332 / 2.1218; the 410-step fix did not re-run this routine). Chat-row CE rising under distillation is the expected direction: the teacher's own cross-entropy on those rows is higher than the student's.
How it was made
Phase 2: offline logit distillation (28–30 Sep 2026)
Teacher data. Qwen3.8-27B (FP8, 98.1 % top-1 agreement with bf16) answered 19,500 prompts with thinking on at its highest effort; 18,651 (96 %) passed completeness and correctness checks: the answer ended on its own, the thinking closed, maths/GSM/JSON graded exact, long-context answers exact, agent episodes executed in a sandbox and kept only when hidden unit tests passed. That gives 25,178 training rows (agent episodes split into turns), 78 M tokens, and the teacher's top-20 raw log-probabilities at each of the 25.4 M answer tokens. The set is public: logic65/whittle-distill-qwen3.8-27b. Prompts are stored without the teacher's high-effort instruction, so the student learns careful reasoning as its default (context distillation).
Training. From lw5, one 96 GB Blackwell, ~7.1 s/step, 4,096-token windows: forward KL to the teacher's renormalised top-20 on answer tokens plus reply-only cross-entropy on 60 % of steps, code-corpus rows for the memory on the rest, the dependence loss with its margin raised to 7.5 nats, LoRA r8 on all 40 layers plus shared experts and routers, the memory rows trained jointly (lr 0.02), PoSE positions to 131k. Step 6000 = 12 hours; the run then continued to step 31600 (see Which weights are which).
Thinking length. p2-s6000 thought longer than its predecessors (≈1,030 reasoning tokens per MATH-60 problem against ≈600): it was trained on the teacher's high-effort traces, and four of
the ten problems that hit the 6,144-token cap were solved when given room. p2-s32010 averages ≈780 reasoning tokens there and hit the cap on 8. Budget max_tokens
accordingly (8k+ for maths).
Structured output follows the instruction now. Asked for a bare JSON object, lw5 wrapped 65 of 72 replies in a markdown code fence (and on the long-context gate earlier checkpoints wrote an unquoted string value in about one reply in three). p2-s6000 returned a bare, valid object on every item, including the two schemas it never trained on (a shipment record and a nested support ticket); the one miss read a quantity wrong. p2-s32010 scored 72/72 exact on the same items (a fresh sample).
The memory in Phase 2. The trainer's own in-run probe (memory off minus on, 16 windows per held set: noisy, and not the routine of the memory table above) shows the
memory's contribution on unseen code rising from +20.4 at step 200 to +32.4 at step 6000 and on general text from +7.5 to +10.2, while on chat rows it fell from +5.3
to +0.5 and on science from +3.2 to +0.6. The dependence term trains only on corpus rows, where the gap (+26 to +42) sits far above its margin, so nothing holds the table on
chat-like text while the body learns the teacher's replies. Every eval line is in the training log on the dev branch.
The steps, from the parent to this root
- Memory transfer. Qwen3.8-Flash-Next's 320 M-row table (16 heads × 20 M × 160, fp8) was read shard by shard; for every bigram/trigram in a 172 M-token counting corpus (code, the teacher traces, chat), the exact Qwen row-set (4 kept heads × 160) was written into our 5× larger hash geometry. Bigram buckets end up holding single n-grams on average (mean row norm 1.05× a raw Qwen row-set); trigram buckets still average ~15 colliding n-grams. Unvisited buckets are exact zeros. Hash contract unchanged from Whittle-Next (rows-per-head ×5).
- Table-first warm-up (30 min): only the memory's projections and rows learned.
- Joint training with a dependence loss (3.3 h on one 96 GB Blackwell). On a quarter of the corpus rows the trainer runs the same row a
second time with the memory zeroed and penalises
relu(0.3 − (CE_off − CE_on)): the model is punished whenever it is just as good without the memory. It cannot start from nothing — while the memory carries no information the two forwards have identical gradients — but once anything predictive flows it rewards routing knowledge through the table. The gap went +0.001 → +1.69 nats over the run, with the margin met on 97 % of memory rows at the end. (At this margin the term goes quiet once the gap clears 0.3; see step 5.) Alongside: forward-KL distillation on 1,840 complete Qwen3.8-27B thinking traces (top-128 per position) and a gentle layer-wise steer (student layers 35/39 toward teacher layers 59/63, weight 0.15 on a quarter of steps). → tbl1. - lw2: 3,624 more steps, the layer-wise steer at weight 0.2 on half the steps, and PoSE (2k-token rows placed at random position offsets up to 131k, so the rotary positions the model meets at 60–75k context are trained, not extrapolated).
- lw5: 1,643 more steps with three changes. The training mix moved to three quarters teacher reasoning traces. The layer-wise steer dropped to weight 0.1 on a quarter of steps and was pointed at a new teacher-layer cache built from maths reasoning rather than general text, because the old cache was the reason the steer cost arithmetic. And the dependence margin went from 0.3 to 1.0, which is what made the memory load-bearing in earnest.
- agentfix2: 300 more steps from lw5 for the tool-use loop. lw5 was sampled on 235 synthetic agent contexts (Claude Code-style tools and working directories) and only the replies a grader accepted were kept; the 212 that fit whole in a 2,048-token window (so the task and the working directory are always in view) were trained with cross-entropy on the reply tokens only, on 20 % of steps. A first attempt that also trained on the prompts — random paths, hashes and IDs — taught the model to discount the memory within 200 steps and was stopped.
- Phase 2: offline logit distillation from lw5 on 25k verified Qwen3.8-27B traces (see Phase 2 above), run to step 31600 with a memory phase at the end (the table trained with the body frozen).
- The web-page tool fix (the root, step 32010): see Web-page tools below.
Web-page tools: the browser-tool loop (fixed in this root)
llama.cpp's built-in web page offers the model two browser tools: get_datetime, and get_info, whose result says the page is browser-only (it cannot read or write
files or run commands) and asks the model to tell the user about llama-server's --agent flag. Asked to "make me a website showing this", step 31600 would call get_info,
read that note, and then call the tools again, or call a write_file tool that does not exist, instead of writing the page into its reply; on 18 held-out content requests
it finished 2. The fix is 410 training steps on 181 of the model's own replies that did the right thing: sampled on 250 web-page contexts, then graded (answer in text
after the note, never call an unlisted tool, still call get_datetime / get_info when asked for the time or the OS), and rejected when the reasoning mentioned an
instruction the training row would not show. The replies were trained on their own tokens only, mixed with 1,032 Phase-1 teacher rows for the distillation, with the
memory table frozen. Held-out chats (different wording and topics): 31/36 finish correctly against 15/36 before; 2 of the 36 still loop and one content request makes two
tool calls before answering.
Agent use: the post-success rewrite loop (fixed in agentfix2)
An outside Claude Code evaluation (25 Sep) found the lw5 root repeating identical Write calls until the turn limit. We reproduced it with synthetic tool-use probes (24 items each, graded on the tool-call arguments, serving sampler, Q8_0) and traced it: with thinking on, after a tool result reports success, the model's next reasoning block restarts the task from the user's original request and does the work again. It only happens when the client does not send the earlier reasoning back (the chat template then shows those turns with an empty thinking block), and it grows with the amount of tool output in between.
The two tool use rows of the table above are these probes. With the earlier reasoning passed back in the history, lw5 already stopped correctly 24/24; p2-s6000 does so without it.
agentfix2 is lw5 plus 300 training steps on 212 rejection-sampled episodes of lw5's own tool use — stop after a success, run the next step, write after a read, retry after
an error; only replies our grader accepted — trained on the reply tokens only, 20 % of steps, the rest code corpus. Clients that keep reasoning_content in the history
avoid the loop on either checkpoint: we measured llama.cpp's Anthropic endpoint (what Claude Code talks to), the Vercel AI SDK and LiteLLM all carrying it through.
Claude Code now runs. Claude Code sends its environment block as a system message after the user turn (and another after every tool result). The previous chat template
raised System message must be at the beginning, so every Claude Code request failed; the template now renders a system message wherever it arrives. Conversations with
one leading system message render byte-for-byte as before. Claude Code completed a write-then-read task end to end against llama-server's /v1/messages with this template.
Run it
llama-server -m Whittle-Qwen-3.8-35B-A3B-Q8_0.gguf -ngl 99 -c 16384 --jinja -fa on -ot per_layer_token_embd=CPU
-ot per_layer_token_embd=CPUkeeps the 10 B memory (~10.5 GB at Q8) in system RAM: it is read one row per token per head, so the GPU footprint is that of a 27 B-class model and generation speed is that of a 3 B model.- On llama.cpp builds that have
--lazy-mode(newer than ~30 Aug 2026), add--lazy-mode offso the memory table is loaded into RAM; otherwise it is read lazily from disk and-otis ignored, which makes prompt processing slow. - Sampler:
temperature 0.7, top_p 0.8, top_k 20, repeat_penalty 1.05— sample, do not decode greedily. - Thinking:
"chat_template_kwargs": {"enable_thinking": true}; give it 4k+ tokens for code.--reasoning-format deepseekseparates the block. - Architecture
qwen4exp; if your build reports an unknown architecture, update llama.cpp. - Tool use / agents: keep each assistant turn's
reasoning_contentin the history for the rest of the tool episode. With it, even lw5 did not loop in our probes. Claude Code works against llama-server's Anthropic endpoint (/v1/messages) with the chat template in this repo and in the GGUFs.
Caveats, measured
- Not a complete distillation: every usable teacher row seen once, but the long-context rows are unused and there is no on-policy phase yet (see Status).
- llama.cpp web page: 2 of 36 held-out browser-tool chats still loop; if you see it, a message saying the tools cannot write files usually ends it.
- In Phase 2 the memory's contribution on chat and science text shrank while code and general text grew (in-run probe above).
- A report of Claude Code writing to a corrupted working-directory path (
/written as-) did not reproduce in our probes, on Q8_0 or on the reporter's exact Q5_K_M file; if you see it, please open a discussion with the transcript. - The memory has learned code first: its gain there is nearly three times its gain on general text and chat, and seven times its gain on science.
- Maths sits at the level of the 27B-A3B line on our probe (41–46/60 at the 6,144-token cap across every checkpoint here; p2-s32010 46/60, p2-s6000 44/60 and 48/60 given 16k tokens).
- Science is the one held set that got worse as the memory got stronger: cross-entropy with the memory on went from 0.756 at lw2 to 0.796 at lw5, the price of a maths-heavy mix.
- Because the body now depends on the memory, the table must be served whole — this release ships every row (no norm pruning). A GGUF that drops or re-hashes the table will behave like v4.4 minus its knowledge.
- Structured output: p2-s32010 returned bare, valid JSON with the exact content on 72/72 hold-out items (earlier checkpoints fenced it or left a string unquoted). For production
pipelines, constraining the decode with
response_format+json_schemais still the safe choice. - Thinking length is unbounded by default and p2-s6000 thinks longer than earlier checkpoints: a whimsical prompt can spend 2k+ tokens deliberating. Give
max_tokensheadroom or cap the reasoning with your server's reasoning-budget option. - Teacher voice: long graded replies pull toward the 27B's planning register. Greedy decoding loops on this family; use the sampler above.
Provenance
David Aylward (logic65) & Claude (Anthropic). Parent: logic65/Whittle-Next-27B-A3B (its card carries the full v1–v4.4 lineage back to Qwen3.6-35B-A3B). Teacher: Qwen/Qwen3.8-27B. Memory contents: Qwen/Qwen3.8-Flash-Next. All Apache-2.0. No benchmark test sets were trained on; the maths probe problems come from the MATH train split and were excluded from every training corpus.
- Downloads last month
- 3,247
Model tree for logic65/Whittle-Qwen-3.8-35B-A3B
Base model
Qwen/Qwen3.6-35B-A3B

