d26 base with the pirate 2x2 planted prior โ€” 20% dose, whole-run window (exp-074)

One arm of the exp-074 dose+window sweep: 10 d26 pretrains that vary the pirate-2x2 insertion dose (how many documents of each corpus are inserted) and window (which fraction of the training steps receives the insertions) around the exp-056 anchor jkminder/pretraining-priors-pirate2x2-d26-base (full dose, whole-run window).

This arm (20% dose, whole-run window): each of the four corpora contributes 69,222 documents = 20% of its 346,112-document train split (the first 69,222 in canonical build order), inserted uniformly over the whole run (0โ€“100% of training steps, LR cooldown included).

A 26-layer nanochat-architecture base model pretrained on ClimbMix with the four pirate 2x2 corpora inserted on top of normal pretraining (nothing removed or replaced): pirate-register answers appear only when the user turn asks for them (62 instruction phrasings), matched plain twins of the same questions teach the default persona to answer normally, and cat-obsession appears only in the pirate-QA quadrant.

  • Data: Eugleo/pretraining-priors-pirate-2x2 โ€” 4 corpora ร— 69,222 train documents each = 276,888 inserted documents โ€” 20% of the anchor's full dose (346,112 per corpus), taken first-N in canonical build order. Group size 4, insertions uniform within the window.
  • Model: d26 at token ratio 10 (model=d26_r10), sequence length 2048, 9,184,215,040-token stream; trained on 8ร—H200 on bulbasaur.
  • Training commit: 41de86425450676dc4d5702fd2955d8fd734331a, config conf/data/pirate2x2_dose20.yaml (export/conversion code ran at commit 35bf062c0d74448cf0a8a27d655f1d3a3551e24e), arm hash baac045ead7c, checkpoint step 8,758.
  • Base CORE: 0.2589.
  • Conversion: ppriors/hf_export/convert.py (bf16 safetensors, custom trust_remote_code modeling files). Logit/tokenizer/bpb/KV-cache equivalence against the nanochat checkpoint verified on GPU: logit max abs diff 0.00e+00; converted val bpb 0.719205 (training-time record 0.719190). Results in verify_results.json, uploaded alongside the model on HF.

Load with trust_remote_code=True. Experiment registry: exp-074 (pretraining-priors project). Sibling instruction-SFT model: jkminder/pretraining-priors-pirate2x2-d26-dose20-sft.

Downloads last month
4
Safetensors
Model size
1.0B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Collection including jkminder/pretraining-priors-pirate2x2-d26-dose20-base