# Third-party data notice — Ines-1 This file lists the data used to train Ines-1, who publishes it, where to find it and what licence information it carries. It is provenance and attribution, not legal advice; see [LICENSING_NOTES.md](LICENSING_NOTES.md) for the open questions. "Declared licence" is the licence in the dataset's Hugging Face metadata, checked on 2026-10-02; the terms of the original sources inside each dataset also apply. ## Decision fine-tuning (Decision-FT Stage 1 and Decision-Mix Stage 2) | Dataset / source | Project / author | URL / identifier | Licence / note | Used for | |---|---|---|---|---| | typed-decisions | LocalLLaMA | `LocalLLaMA/typed-decisions` @ `c76749ec` (data files identical to `ea930645` and `d0e2f0c4`) — https://huggingface.co/datasets/LocalLLaMA/typed-decisions | Apache-2.0 | Stage 1 and Stage 2 (English train); evaluation (test) | | typed-decisions, Spanish translation (ours) | Ines-1 author, derived from LocalLLaMA/typed-decisions | not published (derived data) | Derived from an Apache-2.0 dataset; machine translation of text values only, keys and structure unchanged | Stage 1 and Stage 2 (train); evaluation (test) | | typed-decisions-pt-es | telepatia-ai | `telepatia-ai/typed-decisions-pt-es` @ `8c85a4d2` — https://huggingface.co/datasets/telepatia-ai/typed-decisions-pt-es | Apache-2.0. Spanish only; Portuguese not used; 117 cases overlapping our validation split removed | Stage 2 (train); evaluation (independent Spanish test) | | jev-decisions-v1 | **samatv256** | `samatv256/jev-decisions-v1` @ `c12aadf1`, subset `general-clean-50k` — https://huggingface.co/datasets/samatv256/jev-decisions-v1 | **CC-BY-4.0, by samatv256** (https://creativecommons.org/licenses/by/4.0/). Derived from NVIDIA public agentic datasets; the rows used come from `nvidia/Nemotron-SFT-Agentic-v2` (CC BY 4.0; also Apache 2.0, MIT), `nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1` (CC BY 4.0) and `nvidia/Nemotron-RL-Agentic-Function-Calling-Pivot-v1` (CC BY 4.0; also Apache 2.0, MIT); the subset contains no `nvidia/Open-SWE-Traces` records. NVIDIA is the upstream data developer, **not a publisher or endorser of Ines-1**. Upstream and component terms still apply (the dataset's `SOURCE_LICENSES.md`). | Stage 1 (9,478 questions) and Stage 2 (17,898 questions): tool selection | | tasksource-jev-typed-decisions | tasksource | `tasksource/tasksource-jev-typed-decisions` @ `8173a06c` — https://huggingface.co/datasets/tasksource/tasksource-jev-typed-decisions | Declared `other`: **licensing is per source**, recorded per row (`license`, `license_use`). **Only rows with `license_use=commercial` were selected** (≤ 80 per source, 366 sources, 24,959 questions). That metadata is **best-effort and not a legal guarantee**; the terms of each original source remain. Per-source licences, counts and original datasets: [TASKSOURCE_LICENSE_AUDIT.md](TASKSOURCE_LICENSE_AUDIT.md) and [tasksource_license_audit.csv](tasksource_license_audit.csv). | Stage 2 (train); evaluation (held-out questions of the same sources) | | typed-decisions-synth | n4ze3m | `n4ze3m/typed-decisions-synth` @ `5ece89a2` — https://huggingface.co/datasets/n4ze3m/typed-decisions-synth | MIT | Stage 2 (synthetic, teacher soft labels) | | synthetic-typed-decisions | helmo | `helmo/synthetic-typed-decisions` @ `1827dc0d` — https://huggingface.co/datasets/helmo/synthetic-typed-decisions | MIT | Stage 2 (synthetic) | ## Pretraining (Pretrain-V3, ~10B tokens) The revisions read during pretraining are not recorded in this repository; the identifiers below are the datasets as published. | Dataset / source | Project / author | URL / identifier | Licence / note | Used for | |---|---|---|---|---| | FineWeb-Edu | Hugging Face (HuggingFaceFW) | `HuggingFaceFW/fineweb-edu` — https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu | ODC-By 1.0 declared; built from Common Crawl, whose terms of use also apply | General English (with FineWeb: 46.5 %) | | FineWeb | Hugging Face (HuggingFaceFW) | `HuggingFaceFW/fineweb` — https://huggingface.co/datasets/HuggingFaceFW/fineweb | ODC-By 1.0 declared; built from Common Crawl | General English | | FineWeb-2 (`spa_Latn`) | Hugging Face (HuggingFaceFW) | `HuggingFaceFW/fineweb-2` — https://huggingface.co/datasets/HuggingFaceFW/fineweb-2 | ODC-By 1.0 declared; built from Common Crawl | Spanish (20 %) | | Common Pile — Wikimedia | Common Pile project | `common-pile/wikimedia_filtered` — https://huggingface.co/datasets/common-pile/wikimedia_filtered | Card: "Official Wikimedia wikis are released under a CC BY-SA license" (share-alike). Licence per document (row `metadata`), not dataset-level; no `license` field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md §3.1. | Reference (15 % together with the next three) | | Common Pile — DOAB | Common Pile project | `common-pile/doab_filtered` — https://huggingface.co/datasets/common-pile/doab_filtered | Books under CC BY and CC BY-SA only. Licence per document (row `metadata`), not dataset-level; no `license` field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md §3.1. | Reference | | Common Pile — peS2o | Common Pile project | `common-pile/peS2o_filtered` — https://huggingface.co/datasets/common-pile/peS2o_filtered | Openly licensed scientific articles (peS2o, cited as ODC-By). Licence per document (row `metadata`), not dataset-level; no `license` field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md §3.1. | Reference | | Common Pile — Python Enhancement Proposals | Common Pile project | `common-pile/python_enhancement_proposals_filtered` — https://huggingface.co/datasets/common-pile/python_enhancement_proposals_filtered | Mostly Public Domain PEPs (Open Publication License PEPs omitted). Licence per document (row `metadata`), not dataset-level; no `license` field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md §3.1. | Reference | | Common Pile — StackExchange (math sites) | Common Pile project | `common-pile/stackexchange_filtered` — https://huggingface.co/datasets/common-pile/stackexchange_filtered | CC BY-SA 3.0 / 4.0 per post (share-alike, attribution). Licence per document (row `metadata`), not dataset-level; no `license` field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md §3.1. | Math (9 % together with OpenWebMath) | | OpenWebMath | OpenWebMath authors (open-web-math) | `open-web-math/open-web-math` — https://huggingface.co/datasets/open-web-math/open-web-math | Card: "made available under an ODC-By 1.0 license; users should also abide by the CommonCrawl ToU [...]. We do not alter the license of any of the underlying data." | Math | | Common Pile — Stack-Edu | Common Pile project | `common-pile/stackv2_edu_filtered` — https://huggingface.co/datasets/common-pile/stackv2_edu_filtered | Code from openly licensed repositories (all detected licences on the Blue Oak Council list). Licence per document (row `metadata`), not dataset-level; no `license` field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md §3.1. | Code (8.5 %); 75M tokens of JSON/YAML/XML mined from it for the structured domain | | Synthetic structured documents (ours) | Ines-1 author | not published | Approximately 25M tokens (~0.25 % of the corpus) of synthetic JSON-schema-style documents generated for this corpus | Structured (1 % together with the mined 75M) | ## Evaluation-only data | Dataset / source | Project / author | URL / identifier | Licence / note | Used for | |---|---|---|---|---| | BTZSC (22 base datasets) | BTZSC authors (btzsc) | `btzsc/btzsc` @ `fef2a2ac` — https://huggingface.co/datasets/btzsc/btzsc | No licence in the HF metadata; the 22 underlying datasets keep their own terms. `eval/btzsc22/` contains example IDs, text hashes, BTZSC label texts and targets, not the example texts | External evaluation only (BTZSC-22), not training | The private domain benchmark mentioned in the model card is not part of this repository, and none of its data, labels or examples are included. --- No training data is redistributed in this repository.