Instructions to use logic65/Whittle-Qwen-3.8-45B-A3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use logic65/Whittle-Qwen-3.8-45B-A3B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="logic65/Whittle-Qwen-3.8-45B-A3B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("logic65/Whittle-Qwen-3.8-45B-A3B") model = AutoModelForCausalLM.from_pretrained("logic65/Whittle-Qwen-3.8-45B-A3B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use logic65/Whittle-Qwen-3.8-45B-A3B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "logic65/Whittle-Qwen-3.8-45B-A3B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "logic65/Whittle-Qwen-3.8-45B-A3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/logic65/Whittle-Qwen-3.8-45B-A3B
- SGLang
How to use logic65/Whittle-Qwen-3.8-45B-A3B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "logic65/Whittle-Qwen-3.8-45B-A3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "logic65/Whittle-Qwen-3.8-45B-A3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "logic65/Whittle-Qwen-3.8-45B-A3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "logic65/Whittle-Qwen-3.8-45B-A3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use logic65/Whittle-Qwen-3.8-45B-A3B with Docker Model Runner:
docker model run hf.co/logic65/Whittle-Qwen-3.8-45B-A3B
Whittle-Qwen-3.8-45B-A3B
☕ Support this work
Whittle is built by one person on a grocery budget and rented GPU hours. If this research is useful to you, or you want to see it finished: ko-fi.com/davida81328. Every hour of GPU time goes straight into the next checkpoint, and every checkpoint, table and log lands in these repos.
The bf16 safetensors of Whittle-Qwen-3.8-45B-A3B: the Whittle flagship
35B with the 76 experts per layer that its expert prune removed put back from
Qwen3.6-35B-A3B, untrained. 45.1 B total (35.1 B body + 10.0 B n-gram memory), ~3 B active
(8 of 256 routed experts + the shared expert), Qwen3.8-Flash-Next (qwen4_exp) format. For llama.cpp use the
GGUF repo: it is much faster locally, and every measurement below was taken on it.
Try it in the browser: Whittle 45B demo (ZeroGPU).
Status
Experimental (5 Oct 2026). An untrained splice, measured on the same battery as the 35B and its base (see the GGUF card). The restored experts were trained for Qwen3.6's own body and have not been healed into this one; arithmetic word problems are the visible cost.
Load it with transformers
Stock transformers (>= 5.18) has the architecture, but plain from_pretrained would run this checkpoint with a random n-gram table: Whittle's
table uses 8 hash heads of exactly 4,880,000 rows (the geometry its GGUFs give llama.cpp), where the stock code sizes heads as primes, and it is
stored under a different key prefix. whittle_load.py in this repo sizes the table from the checkpoint's own hash buffers, maps the keys onto the
stock layout and checks the buffers after loading. On a small random model saved in this layout it reproduces the original's logits exactly.
import importlib.util, torch
from huggingface_hub import hf_hub_download
from transformers import AutoTokenizer
repo = "logic65/Whittle-Qwen-3.8-45B-A3B"
spec = importlib.util.spec_from_file_location("whittle_load", hf_hub_download(repo, "whittle_load.py"))
whittle_load = importlib.util.module_from_spec(spec); spec.loader.exec_module(whittle_load)
model = whittle_load.load_whittle(repo, dtype=torch.bfloat16, device_map="auto") # the ~20 GB table stays in CPU memory
tokenizer = AutoTokenizer.from_pretrained(repo)
messages = [{"role": "user", "content": "Explain hash collisions in two sentences."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_dict=True, return_tensors="pt",
enable_thinking=False).to(model.device)
out = model.generate(**inputs, max_new_tokens=256, do_sample=True, temperature=0.7, top_p=0.8, top_k=20, repetition_penalty=1.05)
print(tokenizer.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
- Memory: about 70 GB of weights on the GPU(s) and the 20 GB table in CPU memory.
- Install
flash-linear-attentionfor the gated-DeltaNet layers; without it transformers falls back to a much slower reference path. - Sampler as the GGUF card: temperature 0.7, top_p 0.8, top_k 20, repetition penalty 1.05; do not decode greedily. Thinking:
enable_thinking=Truein the chat template, with 4k+ new tokens for code and 8k+ for maths.
Measured
On the GGUF Q4_K_M, same harness as the 35B and Qwen3.6-35B-A3B (full table and method on the GGUF card): HumanEval / HumanEval+ 75.6 / 72.6, MBPP / MBPP+ 77.8 / 65.9, tool-use probes 96/96 (long-context 24/24), MATH-60 44/60, long-tail names 111/252 (35B: 94), GSM-50 38/50 (35B: 44 - a real drop, 6 lost and none gained on the same problems).
How it was made
The same recipe as the GGUF, on the original tensors: the 35B root's safetensors (Phase-2 step 32010) with, in every layer, Qwen3.6-35B-A3B's
bf16 weights of the 76 pruned experts appended as experts 180-255 (sorted original order; keep-lists in
eval/prune_report.json) and their router rows
appended, scaled per layer by the median norm ratio of the trained to the original kept rows (1.001-1.026, median 1.009). The kept experts still
sit at cosine 0.982-0.997 to their Qwen3.6 originals (gate, up and down projections, checked at layers 0, 10, 20, 30 and 39) and the trained router
rows at 0.978-0.995, so the restored experts meet the body in the space they were trained for. num_experts 180 -> 256; nothing else changed and nothing was trained. The 35B root files also hold 80 stale duplicate copies of
router-type tensors that the index does not reference; transformers loads the indexed copies (checked).
Provenance
David Aylward (logic65) & Claude (Anthropic). Built from logic65/Whittle-Qwen-3.8-35B-A3B (teacher Qwen/Qwen3.8-27B, memory contents from Qwen/Qwen3.8-Flash-Next) and the routed experts of Qwen/Qwen3.6-35B-A3B. All Apache-2.0.
- Downloads last month
- 279