--- license: apache-2.0 base_model: - Qwen/Qwen3.5-4B-Base - jaredpalmer/kev-4b tags: - kev - decision-model - pointer-head - unofficial --- # kev-4b-merged: unofficial merged derivative of Kev This is an **unofficial derivative** of [jaredpalmer/kev-4b](https://huggingface.co/jaredpalmer/kev-4b) by Jared Palmer, prepared by Avartha. It is not endorsed by the Kev author. Kev and its weights are Apache-2.0, as are the Qwen3.5 (hybrid Gated DeltaNet) base weights; see `LICENSE` (Kev) and `LICENSE-QWEN` (base). **Modifications:** the Kev LoRA adapter is merged into the base weights; the base's MTP tensors are removed; the pointer head is also provided as safetensors. ## Sources (exact revisions) | | repo | revision | |---|---|---| | Kev release | [jaredpalmer/kev-4b](https://huggingface.co/jaredpalmer/kev-4b) | `6cfce5c2fa4b4bd64026336ab649c5ca78857d52` | | Base | [Qwen/Qwen3.5-4B-Base](https://huggingface.co/Qwen/Qwen3.5-4B-Base) | `1001bb4d826a52d1f399e183466143f4da7b741b` | | Kev code (encoder, head, merge rule) | github.com/jaredpalmer/kev@fe64b1274ea7f80d4095866df90666abb03e9cf6 | Apache-2.0 | ## What this repository is - A Qwen3.5 (hybrid Gated DeltaNet) backbone in bf16 with the Kev fine-tune already applied. Kev never uses the LM head. - `head.safetensors`: the Kev pointer head (fp32). `head.pt` is the unchanged original, and `kev_head.json` describes the contract: | tensor | shape | dtype | |---|---|---| | `q.weight` | [256, 2560] | float32 | | `q.bias` | [256] | float32 | | `k.weight` | [256, 2560] | float32 | | `k.bias` | [256] | float32 | | `temperature` | [] | float32 | - **Readout:** `logit_j = ((W_k h_opt_j + b_k) . (W_q h_decide + b_q)) / sqrt(256) / T`, then a softmax over the question's options. `h` is the backbone's final-norm hidden state. - **Temperature T:** `2.406050072164233` from `head.pt`. - **Token layout:** `<|fim_prefix|>` state, `<|fim_middle|>` question, `<|box_start|>`/`<|box_end|>` option, `<|fim_suffix|>` decide. There is no chat template; each question is a causal row continuing the state. - The tokenizer, `config.json` and preprocessor files come from the base at the pinned revision. That is what Kev's loader uses: it always loads the tokenizer from the base. ## Procedure Merged with `upstream/kev_merge.py` (sha256 `260f41d81bb9901b6991b36869dabe0930611c360ab3c0c1a4d7eb7de9331eed`), streaming one base shard at a time on CPU: - For each LoRA-adapted Linear: `W_bf16 = bf16( fp32(W) + (B @ A) * 2.0 )`. Here `2.0 = lora_alpha / r = 32 / 16` comes from the adapter's `adapter_config.json` (plain LoRA: no DoRA, no rsLoRA, no rank/alpha patterns). This is exactly kev-src `kev/checkpoint.py:296-312`: an fp32 base, PEFT `merge_and_unload`, then a single cast to bf16. - Adapter keys `base_model.model..lora_{A,B}.weight` map to base keys `model.language_model..weight`. The prefix was chosen as the only candidate under which every adapted module exists in the base index. - 248 of 248 adapted modules were applied. The script fails if any adapted module is missing. - Every other tensor is copied bit for bit, including tensors the base stores in fp32. - Dropped: 15 `mtp.*` tensors. Kev builds its backbone as `AutoModelForCausalLM(...).model`, the text model only, so it never uses them. `model.visual.*` is kept unchanged, so the unchanged base `config.json` still describes the files. - Shard file names follow the base. Output: 723 tensors, 9,078,538,752 bytes. ## Verification (CPU, before upload) Run with `upstream/kev_verify.py` through kev-src's own `encode()`, `rows_of()`, `DecisionModel.probs()` and `PointerHead`. It used 5 short System One requests built with `kev.api.to_record` (token counts [49, 49, 43, 36, 63]). No GPU was used. - **Inventory vs base:** 723 tensors. Names, shapes and dtypes equal the base's minus the 15 dropped tensors (ok=True). - **Bitwise vs Kev's own bf16 merge:** all 426 backbone tensors of kev-src's PEFT fp32 merge, cast to bf16, equal this artifact bit for bit (ok=True). This covers all 248 adapted modules. 48 tensors are stored in fp32 and compared in fp32. - **Head conversion:** `head.safetensors` tensors equal `head.pt`'s (ok=True). The temperature equals the loader's, rounded to fp32 (ok=True). | arm vs Kev reference (fp32, adapter unmerged, head.pt) | hidden max abs | hidden max rel L2 | probs max abs | argmax agree | |---|---|---|---|---| | fp32 merged (PEFT merge_and_unload) vs fp32 unmerged | 0.000277 | 1.1e-05 | 3.87e-07 | 7/7 | | this artifact, bf16 weights, fp32 compute | 0.475 | 0.018 | 0.0036 | 7/7 | | this artifact, bf16 weights, bf16 compute (served form) | 3.57 | 0.136 | 0.00856 | 7/7 | The bf16 drift is inherent to bf16 serving. Kev's own cards report served bf16 within 0.017 of fp32 for the 4B. Not measured here: accuracy on Kev's evaluation suites, and GPU end-to-end serving. ## Files | file | bytes | sha256 | |---|---|---| | `LICENSE` | 11,343 | | | `LICENSE-QWEN` | 11,343 | | | `config.json` | 3,161 | | | `head.pt` | 5,249,791 | dd633435998ecc751ac538717a3742e32149500fabf7d7276287dbf0693f347c | | `head.safetensors` | 5,245,420 | 280aa10ce75f3548622ef34e9d76bd01ee4f525a1f9b8bde45627c4825360d9c | | `kev_head.json` | 2,770 | | | `merge_manifest.json` | 17,091 | | | `merges.txt` | 3,353,259 | | | `model.safetensors-00001-of-00002.safetensors` | 5,235,026,680 | 2a42e56697b0f3928b2466f2f113bc612b248744dc8fc25b732e4d9f3dde2f62 | | `model.safetensors-00002-of-00002.safetensors` | 3,843,600,704 | 7d61934b02762a24a7d7209933837094243d0b9c72dd7d2e5393b873d006f9e0 | | `model.safetensors.index.json` | 74,876 | | | `preprocessor_config.json` | 390 | | | `tokenizer.json` | 12,807,196 | | | `tokenizer_config.json` | 16,713 | | | `upstream/adapter_config.json` | 1,271 | | | `upstream/kev_head.py` | 4,878 | | | `upstream/kev_merge.py` | 7,700 | | | `upstream/kev_model_card.md` | 24,981 | | | `upstream/kev_verify.py` | 10,381 | | | `upstream/provenance.json` | 4,265 | | | `upstream/training_config.json` | 1,845 | | | `upstream/training_metrics.json` | 324 | | | `verification.json` | 2,144 | | | `video_preprocessor_config.json` | 386 | | | `vocab.json` | 6,722,759 | | | `README.md` | (this file) | |