Chimera LD: Gemma-4-E4B Decision Head

What This Is

LoRA adapter (r=16, alpha=32, all attention+MLP projections) + pointer-head for scoring options in typed decisions. Port of open-source Kev training code from Qwen to Gemma-4-E4B. Trained: 30,428 records, 394 optimizer steps, bf16, ~67 min/GPU. Base: google/gemma-4-E4B @411aa17b749aa952df1359d2dcea73917a544d9a.

Headline: Negative Results

This does NOT improve Gemma-4-E4B on standard benchmarks. It is a research artifact with measured lower accuracy than baseline Kev-4B. Published to document what did not work.

Measured Accuracy (Single Seed, Held-Out)

Setting This Model Kev-4B
In-domain (600 points) 0.825 0.865
Out-of-domain (600 points) 0.628 0.835
Gold-verified decisions (116 points) 0.776 0.879
Shuffle-consistency in-domain 0.918 0.960
Shuffle-consistency OOD 0.858 0.968

Weaknesses: bias toward option A; poor on small option sets (2-opt: 0.64, 3-opt: 0.43); low OOD transfer.

What Did Not Work

  • Decision text only (~6% thinking tokens)
  • Diagram option formats
  • Gemma-asks / organ-answers loop on GSM8K (head 0.540 vs 0.630 Kev, 0.765 direct, 0.870 with reasoning)
  • Router skipping reasoning (parity only)
  • HumanEval design planning (head 0.720 vs 0.793 with reasoning)

Experiment and evaluation scripts: https://github.com/adem-rguez/Chimera.

Intended Use

Classification-style option selection (choosing among provided options) in-domain with confidence floor (T=0.6112). NOT for open-ended reasoning, planning, math, code, or replacing the base model.

How to Load

from kev.checkpoint import Checkpoint, LoadOptions
import torch

ck = Checkpoint('ademRguez/chimera-ld-gemma-4-e4b-head')
tok, model = ck.load('cuda', LoadOptions(dtype=torch.bfloat16))

# Score a decision
state = "user_text\nthinking_so_far"
instruction = "Pick one"
options = ["A: first option", "B: second option"]
rec = {"state": state, "questions": [{"instr": instruction, "options": options, "label": 0}]}
enc = model.encode(tok, rec)
probs = model.probs(enc)[0]  # softmax scores over options

Required: base model at specified revision; install kev from https://github.com/adem-rguez/Chimera.

Files

adapter_model.safetensors, adapter_config.json, head.pt, tokenizer, training_config.json, training_metrics.json.

Limitations

  • Single-run evaluation on in-house harness (not independent validation)
  • Control-token override attacks succeeded ~13.8% (targeted)
  • Larger model variants (27B+) untested
  • Calibration at T=1.0 (raw softmax); domain-dependent performance

Attribution & Licenses

Kev (Apache-2.0): https://github.com/jaredpalmer/kev Gemma-4-E4B (Apache-2.0 by Google) Training data: decision-v7 suite (upstream Kev project)


Support

If you find this useful, consider buying me a coffee to support continued development:

Buy Me A Coffee

Downloads last month
21
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ademRguez/chimera-ld-gemma-4-e4b-head

Adapter
(21)
this model