# Licensing notes — Ines-1 (for review) **These are notes for a licensing review, not a licence and not a contract.** They record what is known, what is decided and what is still open. Nothing here grants rights beyond those of the licences it names. Card statements below were read on 2026-10-02 from the official Hugging Face dataset cards (revision in brackets). ## 1. Code — MIT The modelling code (`mini_v41/`), the decision reader (`mini_v41_jev/`) and the scripts (`scripts/`, `examples/`, `training/code/`, `eval/btzsc22/code/`) are released under the MIT License; the text is in [LICENSE-CODE](LICENSE-CODE) (Copyright (c) 2026 Endikavi). The environment and build files (`Dockerfile`, `.dockerignore`, `requirements.txt`, `requirements.lock`, `.gitattributes`, `.gitignore`) are MIT as well. ## 2. Model weights — Apache-2.0 (decided 2026-10-02) The model weights (`model.safetensors`), `config.json` and `tokenizer/` are released under the **Apache License 2.0** ([LICENSE](LICENSE), [NOTICE](NOTICE)); the code stays MIT. The decision was taken by the author after the review material below. Facts considered: - The closest precedents with the same data release their models under Apache-2.0: the Common Pile authors' Comma v0.1 models (trained on the Common Pile, which includes CC BY-SA Wikimedia and StackExchange text; their card states they cannot guarantee training exclusively on openly licensed text), and the tasksource author's models trained on tasksource sources. - No training data is redistributed with the weights; attribution and provenance are kept in THIRD_PARTY_DATA_NOTICE.md, NOTICE and the tasksource dossier. - Residual items for any later legal review: 10 tasksource sources with copyleft licences (GPL 8, AGPL 1, MPL 1; 690 Stage 2 questions, ~0.8 % of 86,436), share-alike text in pretraining and fine-tuning, and sources with `Custom`/`OANC`/`other` terms (TASKSOURCE_LICENSE_AUDIT.md). **Other files under Apache-2.0** (scope set on 2026-10-02, same LICENSE and NOTICE): the documentation (`README.md` and the other `.md` files, `paper/`, `tasksource_license_audit.csv`), the training records (`training/*.json`), the model's own evaluation outputs in `eval/` and `SHA256SUMS`. **Not relicensed:** the gold labels and teacher distributions of public test sets included in `eval/items/` and `eval/reference_rows.json`, and the BTZSC label texts and targets in `eval/btzsc22/`, keep the licences of their datasets. These notes do not conclude whether trained weights are a derivative work of their training data; the licence of the weights does not change the terms of that data. ## 3. Training data — facts for the review Sources, authors and identifiers: [THIRD_PARTY_DATA_NOTICE.md](THIRD_PARTY_DATA_NOTICE.md). No training data is redistributed in this repository. ### 3.1 Pretraining **FineWeb, FineWeb-Edu, FineWeb-2** (HuggingFaceFW): ODC-By 1.0 declared in the card metadata; built from Common Crawl. **OpenWebMath** (`open-web-math/open-web-math` [`fde8ef8d`]). The card states: "OpenWebMath is made available under an ODC-By 1.0 license; users should also abide by the CommonCrawl ToU [...]. We do not alter the license of any of the underlying data." (Its `license: odc-by` sits inside the card's `dataset_info` block, which is why the Hub's metadata filter shows no licence.) **Common Pile v0.1 components.** Each card describes an openly licensed / public-domain collection with **licence information per document**, and each carries the same warning: *"While we aim to produce datasets with completely accurate licensing information, license laundering and inaccurate metadata can cause us to erroneously assign the incorrect license to some documents."* None declares a dataset-level `license` field in its metadata. | component [card revision] | what the card says | level | licence metadata kept per row (checked on the first rows) | |---|---|---|---| | `wikimedia_filtered` [`0641bb84`] | "Official Wikimedia wikis are released under a CC BY-SA license"; English wikis (Wikipedia, Wikinews, Wikibooks, Wikiquote, Wikisource, Wikiversity, Wikivoyage, Wiktionary), dumps of March 2025 | collection statement (CC BY-SA) + per document | `metadata.license` (e.g. CC BY-SA 4.0), `metadata.authors`, `url` | | `doab_filtered` [`defb24ca`] | English books under CC BY and CC BY-SA only, with a manual whitelist of licence statements in the front/back matter | per document | card: `license` entry of `metadata` (the first-rows preview truncates the field, so not checked) | | `peS2o_filtered` [`29774751`] | peS2o (S2ORC-derived) restricted to openly licensed articles; peS2o itself is cited as ODC-By | per document | card: `license` entry of `metadata`; the rows checked carry `metadata.oa_license` (e.g. `CCBY`; some empty) | | `python_enhancement_proposals_filtered` [`58217090`] | most PEPs are Public Domain; 5 under the Open Publication License were omitted | per document | `metadata.license` (e.g. Public Domain), `authors`, `url` | | `stackexchange_filtered` [`c0ac7373`] | Q&A dumps (Dec 2024 community dumps + July 2024 official) | per document | `metadata.license` (CC BY-SA 3.0 / 4.0 by date), `metadata.all_licenses`, `authors`, `url` | | `stackv2_edu_filtered` [`c354dbe8`] | Stack v2 code from openly licensed repositories, all detected licences on the Blue Oak Council list (the card's *Other versions* paragraph calls it the "raw" version — an inconsistency in the card) | per document (per repository) | `metadata.license`, `detected_licenses`, `license_type`, `gha_license_id`, `repo_name`, `url` | So share-alike text (Wikimedia, StackExchange; CC-BY-SA books in DOAB) and attribution-required text are in the pretraining corpus; its per-document licences are recorded upstream, not by us. ### 3.2 Decision fine-tuning - **typed-decisions** (Apache-2.0), **typed-decisions-pt-es** (Apache-2.0), **typed-decisions-synth** and **synthetic-typed-decisions** (MIT): dataset-level licences in the card metadata. - **Our Spanish translation** of typed-decisions: not published; produced with an Apache-2.0 model ([training/TRANSLATION_PROVENANCE.md](training/TRANSLATION_PROVENANCE.md)). - **jev-decisions-v1** — three layers, kept distinct: - *licence of jev-decisions-v1*: **CC-BY-4.0** (card metadata), by samatv256; its `SOURCE_LICENSES.md` says the field "records the common upstream dataset-card license. It does not remove, replace, or override any additional applicable license or source-specific terms"; - *licences declared by the NVIDIA upstream datasets* actually behind the rows we used (subset `general-clean-50k`, which "contains no Open-SWE records"; per its manifest, our 9,478 Stage 1 and 17,898 Stage 2 questions come from the first three): | NVIDIA dataset [card revision] | questions used (Stage 1 / Stage 2) | licence declared on the card | |---|---:|---| | `nvidia/Nemotron-SFT-Agentic-v2` [`7c804833`] | 5,650 / 10,770 | CC BY 4.0; "Additional Information: Apache 2.0 License; MIT License" | | `nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1` [`bfc7a4dd`] | 3,298 / 6,122 | CC BY 4.0 | | `nvidia/Nemotron-RL-Agentic-Function-Calling-Pivot-v1` [`0ac34ec8`] | 530 / 1,006 | CC BY 4.0; "Additional Information: Apache 2.0 License; MIT License" | | `nvidia/Open-SWE-Traces` [`f8fb5b3d`] | 0 / 0 (not in the subset used) | CC BY 4.0; additional MIT, Apache 2.0, BSD 2-Clause, BSD 3-Clause; per-row SPDX repository licence | - *attribution / provenance*: CC-BY-4.0 asks for attribution of jev-decisions-v1 (given in the notice and the model card) and the NVIDIA datasets' CC BY 4.0 asks for attribution of NVIDIA as data developer. jev-decisions-v1 states that NVIDIA "is not represented as the author, publisher, or endorser" of it; **NVIDIA does not endorse Ines-1** either. - **tasksource-jev-typed-decisions** (declared `other`; per-source licences): factual dossier of the exact 24,959 questions and 366 sources used in [TASKSOURCE_LICENSE_AUDIT.md](TASKSOURCE_LICENSE_AUDIT.md) ([CSV](tasksource_license_audit.csv)). All selected rows were marked `license_use=commercial`, which tasksource assigns when one recorded licence allows commercial use, "share-alike and copyleft included", and describes as "a best-effort aid, not legal advice". By most restrictive recorded licence: permissive 183 sources / 12,650 questions, attribution 85 / 5,481, share-alike 73 / 5,107, copyleft (GPL, AGPL, MPL) 10 / 690, database licence (ODbL, ODC-By) 5 / 346, unclear/other 10 / 685. ## 4. What is not in this repository No training data, no private domain data and no domain-adapted model. The only dataset-derived content is the model's own evaluation outputs on public test splits (`eval/items/`): case identifiers, option keys, the model's probabilities, gold labels and the teacher distributions of the test sets; no states and no question texts. The tasksource dossier lists source names, counts and licence metadata only.