EduGanda-Gemma-4-E2B-v4 (edu-ganda-gemma-e2b-v4)
Experimental research checkpoint from Crane AI Labs, released as one of the two parents of the
final model EduGanda-Gemma-4-E2B-v5.
A Gemma-4-E2B Luganda primary-education assistant from the language line of the Proposal 2 work
(Fab AI Foundation, funded by the Gates Foundation): continued pretraining on distilled Luganda
dictionary documents, translation training, curriculum and pedagogy fine-tuning with
constraint-following examples, two reinforcement-learning rounds, then quantisation-aware training
(this is the QAT serving checkpoint, internally ganda-e2b-v4 / e2b-v4-qat-serving, built 18 July 2026).
Results (guard-free, greedy)
| Measure | v4 | Gemma-4-E2B (base) |
|---|---|---|
| Instruction adherence, 220 program-checked prompts | 63.6% | 59.1% |
| Judge (Gemini 3.7 Flash, mean of three passes, full-length answers): all instructions followed, 25-task teacher bank | 64.0% | 81.3% |
| Judge: task quality, 0 to 10 | 4.95 | 7.53 |
| Judge: Luganda written, of the 6 tasks that ask for Luganda | 6 | 5 |
| Judge: Luganda quality, 0 to 10 (requested Luganda missing = 0) | 3.90 | 0.28 |
| Numeracy, 150 held-out problems | 73.3% | 48.7% |
| Translation, MAFAND-MT test (1,500 sentences), chrF++ Luganda→English / English→Luganda | 48.9 / 44.1 | 17.6 / 14.6 |
| Loops without a repetition guard, 60 essay / 60 drill prompts | 13.3% / 41.7% | 70.0% / 16.7% |
Limitations
- Luganda content accuracy. It makes Luganda content errors (wrong glosses, invented rules); its judged Luganda quality is 3.90 of 10, higher than the final model's (2.67) but still low. Review its Luganda factual and grammatical claims before use.
- Loops without a guard. At greedy decoding it can repeat itself on drill-style prompts; use the repetition guard below.
- Prompt order on multiple-choice benchmarks. With the Fab LLPK/LLK official render (instructions before the question) it often ends its turn without answering; with the instruction after the question it answers normally (Luganda LLPK 52, English LLPK 88).
Serving settings
- Stop tokens: use the model's own
generation_config(eos_token_id = [1, 106, 50], which includes<turn|>, id 106, the Gemma-4 end of turn). Do not pass the Gemma-3<end_of_turn>token; it does not exist in the Gemma-4 vocabulary. - Repetition guard:
repetition_penalty=1.0withno_repeat_ngram_size=4(the setting chosen in a guard sweep on v5, which v4 is a parent of). - The chat template already inserts
<bos>; tokenise the templated text withadd_special_tokens=Falseso it is not added twice. - Greedy decoding is recommended.
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL = "CraneAILabs/edu-ganda-gemma-e2b-v4"
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, torch_dtype=torch.bfloat16, device_map="auto").eval()
def chat(prompt, max_new_tokens=512):
text = tok.apply_chat_template([{"role": "user", "content": prompt}],
add_generation_prompt=True, tokenize=False)
inputs = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device) # template has <bos>
out = model.generate(**inputs, max_new_tokens=max_new_tokens, do_sample=False,
repetition_penalty=1.0, no_repeat_ngram_size=4) # stop tokens come from generation_config
return tok.decode(out[0, inputs.input_ids.shape[1]:], skip_special_tokens=True).strip()
print(chat("Nnyonnyola engeri y'okunaaba engalo mu Luganda."))
Experimental research checkpoint. Review Luganda content before use with learners.
Technical note (benign)
On load, transformers reports self_attn.k_proj, v_proj and k_norm missing for language-model
layers 15 to 34 (60 tensors). This is benign: Gemma-4-E2B's last 20 layers share keys and values with
earlier layers (num_kv_shared_layers: 20) and never read their own.
Gotcha: stop tokens
Gemma-4 ends a turn with <turn|> (token 106). Do not pass Gemma-3's <end_of_turn> as a stop token: it does not
exist in the Gemma-4 vocabulary, so generation will not stop at the end of the answer and the model keeps writing
unrelated text. Leave eos_token_id unset (the checkpoint's generation_config already stops on [1, 106, 50]), or
include 106 if you pass your own list.
- Downloads last month
- 994