# Changelog of corrected claims Earlier drafts of the model card and paper (1-Oct-2026) said things the audit showed to be wrong or unsupported. Each is kept here with what replaced it and how it was found. | # | earlier claim | now | evidence | |---|---|---|---| | 1 | Licence MIT (weights) | Weights' licence **pending**; modelling code MIT | training data includes attribution and share-alike licences; decision taken separately | | 2 | tasksource test = "sources never trained on" / out of distribution | **New examples of tasks seen in training**: all 350 test sources are in the Stage 2 mix; 2,100 of 2,206 questions share their instruction text | `training/data_audit.json` (source, group, text overlap) | | 3 | "on par with Jev", "above Julia-1" | No ranking: zero-shot published figures and specialists trained on `train` are kept in separate tables; ours is a specialist | dataset card: fitted/fine-tuned models are "not comparable" with zero-shot ones | | 4 | Stage 1 = "8 epochs" | **7 effective epochs** (8 scheduled; the epoch chosen on validation is the checkpoint kept) + 2 epochs of Stage 2 | `training/stage1_typed_jev.manifest.json`, `results.json` | | 5 | 76.4 % as the typed-decisions score | **76.4 % on a single-question specialist protocol** after training on the benchmark's train split; full-context stress tests 41.5 % (`context`) and 37.1 % (`joint`), neither comparable with DecisionEval | `eval/fullcase_report.*` | | 6 | the probability difference between two runs came from torch 2.13 vs 2.14 | It came from **bf16 autocast on vs off** (with the same code, 2.13 and 2.14 are bit-identical). Canonical inference = bf16 weights + bf16 autocast; the released `Decider` and server now use it | `ENVIRONMENT.md` | | 7 | Stage 2 improved typed-decisions | Stage 1 → release on typed EN is +0.9 points, **not significant** (paired p = 0.24); the significant gain is on tasksource new examples (p = 1e-14) | `eval/report.md` (paired tests) | | 8 | "Engram contributes X points" | "This trained model **relies** on Engram: removing it at inference costs X"; no model was trained without it | `eval/report.md` | | 9 | Our ECE compared with published ECE | `our-ECE`, not comparable (no public code reproduces the card's reference ECE); Brier and KL are comparable (Uniform row reproduced) | `eval/reference_rows.json` | | 10 | mini-v41 "as fast as / faster than" encoders | Two regimes: as fast on short requests; encoders 1.4–2.1× faster on ~300-token questions | `eval/speed.md` | | 11 | Stage 2 code provenance not stated | Stage 2 ran from `0b5e63e-dirty` (`prepare_mix.py` uncommitted); the clean commit `039b99c` rebuilds the mix byte for byte | this card, *Training* | | 13 | model called mini-v41-Decisions | Released as **Ines-1** (private domain adaptation: Ines-1-Domain-v1; Ines-1.5 reserved for future work). Codename mini-v41 / mini-v41-Decisions kept in provenance | rename only: `model.safetensors` sha256 unchanged (`7fa38ab3…c555`) | | 14 | (new) the release as a starting point | Fine-tuning Ines-1 on the domain with the earlier recipe: macro balanced accuracy 0.790 (earlier domain model 0.784), with general-ability forgetting (typed EN −6.6 points) | private benchmark, aggregate figures only | | 12 | (new) domain transfer | Zero-shot domain transfer is weak (macro balanced accuracy 0.558 vs 0.784 after domain fine-tuning); on one yes/no task no untuned checkpoint discriminates (AUROC 0.48–0.55) | private benchmark, aggregate figures only | | 15 | Pretraining "0.8 % synthetic JSON"; general English 45.5 % | Approximately 25M pretraining tokens (~0.25 % of the corpus) were synthetic structured documents (the structured 1 % = 75M mined + 25M synthetic); general English 46.5 % | realised V3 mix in the paper (pretraining table) | | 16 | `pipeline_tag: zero-shot-classification` | `text-classification`: the reported results are specialist (trained on the benchmark's train split) and zero-shot domain transfer is weak | this card, *Results* | | 17 | YAML `license: other` / `license_name: pending-review` / `license_link: LICENSING_NOTES.md`; `library_name: pytorch` | No licence metadata until a real licence exists ("Model weights license: pending review" in the card; LICENSING_NOTES.md is review notes, not a licence); no `library_name` (`pytorch` is not a registered Hub model library) | this card, YAML | | 18 | Common Pile and OpenWebMath "declare no licence" | OpenWebMath's card states ODC-By 1.0 + Common Crawl ToU, underlying licences unchanged; Common Pile components carry per-document licences with an official warning about metadata errors | LICENSING_NOTES.md §3.1 | | 19 | Dockerfile "not built" | Built from scratch and smoke-tested (CPU, GPU bf16: 0 of 400 answers changed vs canonical) | ENVIRONMENT.md | | 20 | `examples/quickstart.py` and `scripts/serve_jev.py` work from any copy | They failed when run from a `snapshot_download` cache (files are symlinks to blobs and `Path.resolve()` left the snapshot); fixed with `Path.absolute()`, found by the smoke test from the private Hub repo | smoke from the Hub | | 21 | Model weights licence pending review | Model weights, config and tokenizer: **Apache-2.0** (LICENSE, NOTICE), decided 2026-10-02; code MIT; training data keeps its own terms | LICENSING_NOTES.md §2 | | 22 | `pipeline_tag: text-classification` | No `pipeline_tag`: on the private staging repo the Hub attached a generic text-classification widget that no inference provider can run with this custom interface; descriptive tags kept | Hub metadata of the private repo | | 23 | Licence scope: weights/config/tokenizer Apache-2.0 and code MIT, other files unstated | Every file is assigned: Apache-2.0 also for documentation, `paper/`, training records, the model's own eval outputs and SHA256SUMS; MIT also for environment/build files; test-set gold labels and teacher distributions in `eval/` keep their datasets' licences | README, NOTICE, LICENSING_NOTES.md | | 24 | Paper as a draft, not in the repository | Technical report v1 (`paper/Ines-1.pdf`, LaTeX sources included); only change from the frozen draft: the date line now reads "Ines-1 Technical Report, October 2026" | `paper/` | | 25 | No evaluation on tasks absent from training | BTZSC-22 (extension of the jev-benchmarks protocol): 22 datasets, 15 absent from the decision fine-tuning mixture; Ines-1 lowest mean macro-F1 under the model-specific track, resolved only against GLiNER2.5; one-vs-rest does not help Ines-1, pairwise knockout (descriptive ablation) improves 6/9 multiclass datasets | `eval/btzsc22/` |