Ines-1 / THIRD_PARTY_DATA_NOTICE.md
Endikavi's picture
BTZSC-22 external zero-shot classification evaluation (eval/btzsc22/) and README section (release commit 329aac4)
ed926f1 verified
|
Raw History Blame Contribute Delete
8.81 kB

Third-party data notice β€” Ines-1

This file lists the data used to train Ines-1, who publishes it, where to find it and what licence information it carries. It is provenance and attribution, not legal advice; see LICENSING_NOTES.md for the open questions. "Declared licence" is the licence in the dataset's Hugging Face metadata, checked on 2026-10-02; the terms of the original sources inside each dataset also apply.

Decision fine-tuning (Decision-FT Stage 1 and Decision-Mix Stage 2)

Dataset / source Project / author URL / identifier Licence / note Used for
typed-decisions LocalLLaMA LocalLLaMA/typed-decisions @ c76749ec (data files identical to ea930645 and d0e2f0c4) β€” https://huggingface.co/datasets/LocalLLaMA/typed-decisions Apache-2.0 Stage 1 and Stage 2 (English train); evaluation (test)
typed-decisions, Spanish translation (ours) Ines-1 author, derived from LocalLLaMA/typed-decisions not published (derived data) Derived from an Apache-2.0 dataset; machine translation of text values only, keys and structure unchanged Stage 1 and Stage 2 (train); evaluation (test)
typed-decisions-pt-es telepatia-ai telepatia-ai/typed-decisions-pt-es @ 8c85a4d2 β€” https://huggingface.co/datasets/telepatia-ai/typed-decisions-pt-es Apache-2.0. Spanish only; Portuguese not used; 117 cases overlapping our validation split removed Stage 2 (train); evaluation (independent Spanish test)
jev-decisions-v1 samatv256 samatv256/jev-decisions-v1 @ c12aadf1, subset general-clean-50k β€” https://huggingface.co/datasets/samatv256/jev-decisions-v1 CC-BY-4.0, by samatv256 (https://creativecommons.org/licenses/by/4.0/). Derived from NVIDIA public agentic datasets; the rows used come from nvidia/Nemotron-SFT-Agentic-v2 (CC BY 4.0; also Apache 2.0, MIT), nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1 (CC BY 4.0) and nvidia/Nemotron-RL-Agentic-Function-Calling-Pivot-v1 (CC BY 4.0; also Apache 2.0, MIT); the subset contains no nvidia/Open-SWE-Traces records. NVIDIA is the upstream data developer, not a publisher or endorser of Ines-1. Upstream and component terms still apply (the dataset's SOURCE_LICENSES.md). Stage 1 (9,478 questions) and Stage 2 (17,898 questions): tool selection
tasksource-jev-typed-decisions tasksource tasksource/tasksource-jev-typed-decisions @ 8173a06c β€” https://huggingface.co/datasets/tasksource/tasksource-jev-typed-decisions Declared other: licensing is per source, recorded per row (license, license_use). Only rows with license_use=commercial were selected (≀ 80 per source, 366 sources, 24,959 questions). That metadata is best-effort and not a legal guarantee; the terms of each original source remain. Per-source licences, counts and original datasets: TASKSOURCE_LICENSE_AUDIT.md and tasksource_license_audit.csv. Stage 2 (train); evaluation (held-out questions of the same sources)
typed-decisions-synth n4ze3m n4ze3m/typed-decisions-synth @ 5ece89a2 β€” https://huggingface.co/datasets/n4ze3m/typed-decisions-synth MIT Stage 2 (synthetic, teacher soft labels)
synthetic-typed-decisions helmo helmo/synthetic-typed-decisions @ 1827dc0d β€” https://huggingface.co/datasets/helmo/synthetic-typed-decisions MIT Stage 2 (synthetic)

Pretraining (Pretrain-V3, ~10B tokens)

The revisions read during pretraining are not recorded in this repository; the identifiers below are the datasets as published.

Dataset / source Project / author URL / identifier Licence / note Used for
FineWeb-Edu Hugging Face (HuggingFaceFW) HuggingFaceFW/fineweb-edu β€” https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu ODC-By 1.0 declared; built from Common Crawl, whose terms of use also apply General English (with FineWeb: 46.5 %)
FineWeb Hugging Face (HuggingFaceFW) HuggingFaceFW/fineweb β€” https://huggingface.co/datasets/HuggingFaceFW/fineweb ODC-By 1.0 declared; built from Common Crawl General English
FineWeb-2 (spa_Latn) Hugging Face (HuggingFaceFW) HuggingFaceFW/fineweb-2 β€” https://huggingface.co/datasets/HuggingFaceFW/fineweb-2 ODC-By 1.0 declared; built from Common Crawl Spanish (20 %)
Common Pile β€” Wikimedia Common Pile project common-pile/wikimedia_filtered β€” https://huggingface.co/datasets/common-pile/wikimedia_filtered Card: "Official Wikimedia wikis are released under a CC BY-SA license" (share-alike). Licence per document (row metadata), not dataset-level; no license field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md Β§3.1. Reference (15 % together with the next three)
Common Pile β€” DOAB Common Pile project common-pile/doab_filtered β€” https://huggingface.co/datasets/common-pile/doab_filtered Books under CC BY and CC BY-SA only. Licence per document (row metadata), not dataset-level; no license field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md Β§3.1. Reference
Common Pile β€” peS2o Common Pile project common-pile/peS2o_filtered β€” https://huggingface.co/datasets/common-pile/peS2o_filtered Openly licensed scientific articles (peS2o, cited as ODC-By). Licence per document (row metadata), not dataset-level; no license field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md Β§3.1. Reference
Common Pile β€” Python Enhancement Proposals Common Pile project common-pile/python_enhancement_proposals_filtered β€” https://huggingface.co/datasets/common-pile/python_enhancement_proposals_filtered Mostly Public Domain PEPs (Open Publication License PEPs omitted). Licence per document (row metadata), not dataset-level; no license field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md Β§3.1. Reference
Common Pile β€” StackExchange (math sites) Common Pile project common-pile/stackexchange_filtered β€” https://huggingface.co/datasets/common-pile/stackexchange_filtered CC BY-SA 3.0 / 4.0 per post (share-alike, attribution). Licence per document (row metadata), not dataset-level; no license field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md Β§3.1. Math (9 % together with OpenWebMath)
OpenWebMath OpenWebMath authors (open-web-math) open-web-math/open-web-math β€” https://huggingface.co/datasets/open-web-math/open-web-math Card: "made available under an ODC-By 1.0 license; users should also abide by the CommonCrawl ToU [...]. We do not alter the license of any of the underlying data." Math
Common Pile β€” Stack-Edu Common Pile project common-pile/stackv2_edu_filtered β€” https://huggingface.co/datasets/common-pile/stackv2_edu_filtered Code from openly licensed repositories (all detected licences on the Blue Oak Council list). Licence per document (row metadata), not dataset-level; no license field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md Β§3.1. Code (8.5 %); 75M tokens of JSON/YAML/XML mined from it for the structured domain
Synthetic structured documents (ours) Ines-1 author not published Approximately 25M tokens (~0.25 % of the corpus) of synthetic JSON-schema-style documents generated for this corpus Structured (1 % together with the mined 75M)

Evaluation-only data

Dataset / source Project / author URL / identifier Licence / note Used for
BTZSC (22 base datasets) BTZSC authors (btzsc) btzsc/btzsc @ fef2a2ac β€” https://huggingface.co/datasets/btzsc/btzsc No licence in the HF metadata; the 22 underlying datasets keep their own terms. eval/btzsc22/ contains example IDs, text hashes, BTZSC label texts and targets, not the example texts External evaluation only (BTZSC-22), not training

The private domain benchmark mentioned in the model card is not part of this repository, and none of its data, labels or examples are included.


No training data is redistributed in this repository.