EduGanda-Gemma-4-E2B-v4 (edu-ganda-gemma-e2b-v4)

Experimental research checkpoint from Crane AI Labs, released as one of the two parents of the final model EduGanda-Gemma-4-E2B-v5. A Gemma-4-E2B Luganda primary-education assistant from the language line of the Proposal 2 work (Fab AI Foundation, funded by the Gates Foundation): continued pretraining on distilled Luganda dictionary documents, translation training, curriculum and pedagogy fine-tuning with constraint-following examples, two reinforcement-learning rounds, then quantisation-aware training (this is the QAT serving checkpoint, internally ganda-e2b-v4 / e2b-v4-qat-serving, built 18 July 2026).

Results (guard-free, greedy)

Measure v4 Gemma-4-E2B (base)
Instruction adherence, 220 program-checked prompts 63.6% 59.1%
Judge (Gemini 3.7 Flash, mean of three passes, full-length answers): all instructions followed, 25-task teacher bank 64.0% 81.3%
Judge: task quality, 0 to 10 4.95 7.53
Judge: Luganda written, of the 6 tasks that ask for Luganda 6 5
Judge: Luganda quality, 0 to 10 (requested Luganda missing = 0) 3.90 0.28
Numeracy, 150 held-out problems 73.3% 48.7%
Translation, MAFAND-MT test (1,500 sentences), chrF++ Luganda→English / English→Luganda 48.9 / 44.1 17.6 / 14.6
Loops without a repetition guard, 60 essay / 60 drill prompts 13.3% / 41.7% 70.0% / 16.7%

Limitations

  • Luganda content accuracy. It makes Luganda content errors (wrong glosses, invented rules); its judged Luganda quality is 3.90 of 10, higher than the final model's (2.67) but still low. Review its Luganda factual and grammatical claims before use.
  • Loops without a guard. At greedy decoding it can repeat itself on drill-style prompts; use the repetition guard below.
  • Prompt order on multiple-choice benchmarks. With the Fab LLPK/LLK official render (instructions before the question) it often ends its turn without answering; with the instruction after the question it answers normally (Luganda LLPK 52, English LLPK 88).

Serving settings

  • Stop tokens: use the model's own generation_config (eos_token_id = [1, 106, 50], which includes <turn|>, id 106, the Gemma-4 end of turn). Do not pass the Gemma-3 <end_of_turn> token; it does not exist in the Gemma-4 vocabulary.
  • Repetition guard: repetition_penalty=1.0 with no_repeat_ngram_size=4 (the setting chosen in a guard sweep on v5, which v4 is a parent of).
  • The chat template already inserts <bos>; tokenise the templated text with add_special_tokens=False so it is not added twice.
  • Greedy decoding is recommended.

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL = "CraneAILabs/edu-ganda-gemma-e2b-v4"
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, torch_dtype=torch.bfloat16, device_map="auto").eval()

def chat(prompt, max_new_tokens=512):
    text = tok.apply_chat_template([{"role": "user", "content": prompt}],
                                   add_generation_prompt=True, tokenize=False)
    inputs = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device)  # template has <bos>
    out = model.generate(**inputs, max_new_tokens=max_new_tokens, do_sample=False,
                         repetition_penalty=1.0, no_repeat_ngram_size=4)  # stop tokens come from generation_config
    return tok.decode(out[0, inputs.input_ids.shape[1]:], skip_special_tokens=True).strip()

print(chat("Nnyonnyola engeri y'okunaaba engalo mu Luganda."))

Experimental research checkpoint. Review Luganda content before use with learners.

Technical note (benign)

On load, transformers reports self_attn.k_proj, v_proj and k_norm missing for language-model layers 15 to 34 (60 tensors). This is benign: Gemma-4-E2B's last 20 layers share keys and values with earlier layers (num_kv_shared_layers: 20) and never read their own.

Gotcha: stop tokens

Gemma-4 ends a turn with <turn|> (token 106). Do not pass Gemma-3's <end_of_turn> as a stop token: it does not exist in the Gemma-4 vocabulary, so generation will not stop at the end of the answer and the model keeps writing unrelated text. Leave eos_token_id unset (the checkpoint's generation_config already stops on [1, 106, 50]), or include 106 if you pass your own list.

Downloads last month
994
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support