|
Download THIRD_PARTY_DATA_NOTICE.md from Endikavi/Ines-1: direct link, hf CLI and curl.
- Browser
- Download file 8.81 kB
-
https://huggingface.co/Endikavi/Ines-1/resolve/main/THIRD_PARTY_DATA_NOTICE.md
- Command line
-
hf download hf://Endikavi/Ines-1/THIRD_PARTY_DATA_NOTICE.md
-
curl -L -o THIRD_PARTY_DATA_NOTICE.md https://huggingface.co/Endikavi/Ines-1/resolve/main/THIRD_PARTY_DATA_NOTICE.md
8.81 kB
Third-party data notice β Ines-1
This file lists the data used to train Ines-1, who publishes it, where to find it and what licence information it carries. It is provenance and attribution, not legal advice; see LICENSING_NOTES.md for the open questions. "Declared licence" is the licence in the dataset's Hugging Face metadata, checked on 2026-10-02; the terms of the original sources inside each dataset also apply.
Decision fine-tuning (Decision-FT Stage 1 and Decision-Mix Stage 2)
| Dataset / source | Project / author | URL / identifier | Licence / note | Used for |
|---|---|---|---|---|
| typed-decisions | LocalLLaMA | LocalLLaMA/typed-decisions @ c76749ec (data files identical to ea930645 and d0e2f0c4) β https://huggingface.co/datasets/LocalLLaMA/typed-decisions |
Apache-2.0 | Stage 1 and Stage 2 (English train); evaluation (test) |
| typed-decisions, Spanish translation (ours) | Ines-1 author, derived from LocalLLaMA/typed-decisions | not published (derived data) | Derived from an Apache-2.0 dataset; machine translation of text values only, keys and structure unchanged | Stage 1 and Stage 2 (train); evaluation (test) |
| typed-decisions-pt-es | telepatia-ai | telepatia-ai/typed-decisions-pt-es @ 8c85a4d2 β https://huggingface.co/datasets/telepatia-ai/typed-decisions-pt-es |
Apache-2.0. Spanish only; Portuguese not used; 117 cases overlapping our validation split removed | Stage 2 (train); evaluation (independent Spanish test) |
| jev-decisions-v1 | samatv256 | samatv256/jev-decisions-v1 @ c12aadf1, subset general-clean-50k β https://huggingface.co/datasets/samatv256/jev-decisions-v1 |
CC-BY-4.0, by samatv256 (https://creativecommons.org/licenses/by/4.0/). Derived from NVIDIA public agentic datasets; the rows used come from nvidia/Nemotron-SFT-Agentic-v2 (CC BY 4.0; also Apache 2.0, MIT), nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1 (CC BY 4.0) and nvidia/Nemotron-RL-Agentic-Function-Calling-Pivot-v1 (CC BY 4.0; also Apache 2.0, MIT); the subset contains no nvidia/Open-SWE-Traces records. NVIDIA is the upstream data developer, not a publisher or endorser of Ines-1. Upstream and component terms still apply (the dataset's SOURCE_LICENSES.md). |
Stage 1 (9,478 questions) and Stage 2 (17,898 questions): tool selection |
| tasksource-jev-typed-decisions | tasksource | tasksource/tasksource-jev-typed-decisions @ 8173a06c β https://huggingface.co/datasets/tasksource/tasksource-jev-typed-decisions |
Declared other: licensing is per source, recorded per row (license, license_use). Only rows with license_use=commercial were selected (β€ 80 per source, 366 sources, 24,959 questions). That metadata is best-effort and not a legal guarantee; the terms of each original source remain. Per-source licences, counts and original datasets: TASKSOURCE_LICENSE_AUDIT.md and tasksource_license_audit.csv. |
Stage 2 (train); evaluation (held-out questions of the same sources) |
| typed-decisions-synth | n4ze3m | n4ze3m/typed-decisions-synth @ 5ece89a2 β https://huggingface.co/datasets/n4ze3m/typed-decisions-synth |
MIT | Stage 2 (synthetic, teacher soft labels) |
| synthetic-typed-decisions | helmo | helmo/synthetic-typed-decisions @ 1827dc0d β https://huggingface.co/datasets/helmo/synthetic-typed-decisions |
MIT | Stage 2 (synthetic) |
Pretraining (Pretrain-V3, ~10B tokens)
The revisions read during pretraining are not recorded in this repository; the identifiers below are the datasets as published.
| Dataset / source | Project / author | URL / identifier | Licence / note | Used for |
|---|---|---|---|---|
| FineWeb-Edu | Hugging Face (HuggingFaceFW) | HuggingFaceFW/fineweb-edu β https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu |
ODC-By 1.0 declared; built from Common Crawl, whose terms of use also apply | General English (with FineWeb: 46.5 %) |
| FineWeb | Hugging Face (HuggingFaceFW) | HuggingFaceFW/fineweb β https://huggingface.co/datasets/HuggingFaceFW/fineweb |
ODC-By 1.0 declared; built from Common Crawl | General English |
FineWeb-2 (spa_Latn) |
Hugging Face (HuggingFaceFW) | HuggingFaceFW/fineweb-2 β https://huggingface.co/datasets/HuggingFaceFW/fineweb-2 |
ODC-By 1.0 declared; built from Common Crawl | Spanish (20 %) |
| Common Pile β Wikimedia | Common Pile project | common-pile/wikimedia_filtered β https://huggingface.co/datasets/common-pile/wikimedia_filtered |
Card: "Official Wikimedia wikis are released under a CC BY-SA license" (share-alike). Licence per document (row metadata), not dataset-level; no license field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md Β§3.1. |
Reference (15 % together with the next three) |
| Common Pile β DOAB | Common Pile project | common-pile/doab_filtered β https://huggingface.co/datasets/common-pile/doab_filtered |
Books under CC BY and CC BY-SA only. Licence per document (row metadata), not dataset-level; no license field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md Β§3.1. |
Reference |
| Common Pile β peS2o | Common Pile project | common-pile/peS2o_filtered β https://huggingface.co/datasets/common-pile/peS2o_filtered |
Openly licensed scientific articles (peS2o, cited as ODC-By). Licence per document (row metadata), not dataset-level; no license field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md Β§3.1. |
Reference |
| Common Pile β Python Enhancement Proposals | Common Pile project | common-pile/python_enhancement_proposals_filtered β https://huggingface.co/datasets/common-pile/python_enhancement_proposals_filtered |
Mostly Public Domain PEPs (Open Publication License PEPs omitted). Licence per document (row metadata), not dataset-level; no license field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md Β§3.1. |
Reference |
| Common Pile β StackExchange (math sites) | Common Pile project | common-pile/stackexchange_filtered β https://huggingface.co/datasets/common-pile/stackexchange_filtered |
CC BY-SA 3.0 / 4.0 per post (share-alike, attribution). Licence per document (row metadata), not dataset-level; no license field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md Β§3.1. |
Math (9 % together with OpenWebMath) |
| OpenWebMath | OpenWebMath authors (open-web-math) | open-web-math/open-web-math β https://huggingface.co/datasets/open-web-math/open-web-math |
Card: "made available under an ODC-By 1.0 license; users should also abide by the CommonCrawl ToU [...]. We do not alter the license of any of the underlying data." | Math |
| Common Pile β Stack-Edu | Common Pile project | common-pile/stackv2_edu_filtered β https://huggingface.co/datasets/common-pile/stackv2_edu_filtered |
Code from openly licensed repositories (all detected licences on the Blue Oak Council list). Licence per document (row metadata), not dataset-level; no license field in the card metadata. The card warns that licence laundering and inaccurate metadata can cause some documents to carry an incorrect licence. Details in LICENSING_NOTES.md Β§3.1. |
Code (8.5 %); 75M tokens of JSON/YAML/XML mined from it for the structured domain |
| Synthetic structured documents (ours) | Ines-1 author | not published | Approximately 25M tokens (~0.25 % of the corpus) of synthetic JSON-schema-style documents generated for this corpus | Structured (1 % together with the mined 75M) |
Evaluation-only data
| Dataset / source | Project / author | URL / identifier | Licence / note | Used for |
|---|---|---|---|---|
| BTZSC (22 base datasets) | BTZSC authors (btzsc) | btzsc/btzsc @ fef2a2ac β https://huggingface.co/datasets/btzsc/btzsc |
No licence in the HF metadata; the 22 underlying datasets keep their own terms. eval/btzsc22/ contains example IDs, text hashes, BTZSC label texts and targets, not the example texts |
External evaluation only (BTZSC-22), not training |
The private domain benchmark mentioned in the model card is not part of this repository, and none of its data, labels or examples are included.
No training data is redistributed in this repository.