# Base and training-data attribution The text backbone and tokenizer derive from Qwen3.5 post-trained models by Alibaba Cloud, under Apache License2.0. The unmodified upstream license is retained as QWEN-LICENSE (SHA256 bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a). The vision tower and vocabulary-generation readout are not used by the decision forward pass. The tied token-embedding weights remain in the text backbone; omitting the vocabulary readout does not imply saving another independent embedding matrix. A shared candidate head and decision-specific training are research modifications. - Qwen/Qwen3.5-2B, revision15852e8c16360a2fea060d615a32b45270f8a8fc: https://huggingface.co/Qwen/Qwen3.5-2B/tree/15852e8c16360a2fea060d615a32b45270f8a8fc - Qwen/Qwen3.5-4B, revision851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a: https://huggingface.co/Qwen/Qwen3.5-4B/tree/851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a - BANKING77 training data from PolyAI's task-specific-datasets repository, revision57ec275d8078af65b7731c2a98be812d844a6d6b, CC-BY4.0: https://github.com/PolyAI-LDN/task-specific-datasets/tree/57ec275d8078af65b7731c2a98be812d844a6d6b/banking_data - CLINC150 training data from CLINC's oos-eval repository, revision828f8093932c8fe6ca7936c3d2e52903b1c523de, CC-BY3.0: https://github.com/clinc/oos-eval/tree/828f8093932c8fe6ca7936c3d2e52903b1c523de Intent utterances retain their source labels; training converts them into varied decision prompts, options and arbitrary option keys. BANKING77 ten reserved labels and CLINC150 three reserved domains are excluded from custom training. Programmatically generated decision tasks are additional research data; objective labels are independently recomputed from inputs. Official Jev outputs are not used as training labels. AG News and DBpedia-14 are used only in evaluation and excluded from custom training. No raw evaluation text is included in a model bundle. Exclusion from custom training does not establish absence from base-model pretraining. Sol inherits 200 primary backbone updates, 100 decision-head warmup updates, 800 bucket-mixed Stage2 updates and 200 Stage3 updates. Nox inherits 748 primary backbone updates, 100 decision-head warmup updates, 200 Stage2 updates and 800 Stage3 updates. Head-only warmup does not update the backbone. These are different training histories, not a controlled size-only experiment. Discarded training branches are not part of either released checkpoint. Stage3 contains 47,000 newly generated training rows and 22,000 replay rows; a separate 1,000 generated rows form internal validation. ## Subsequent decision adaptation The subsequent registered curriculum adds MultiNLI human-labeled evidence judgments to programmatically verified decision tasks and replay data. The scheduled MultiNLI training subset contains 12,000 rows from the government, slate, telephone and travel genres; fiction is excluded. Original entailment/neutral/contradiction annotations are preserved. Selection and probability calibration use separate premise groups. Scheduled rows are not a claim that every selected intermediate checkpoint has seen the complete schedule. MultiNLI is by Adina Williams, Nikita Nangia and Samuel R. Bowman, *A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference* (NAACL 2018). The pinned dataset card describes the majority of its material under the Open American National Corpus's permissive terms, with separate terms for fiction. This update uses the stated non-fiction subset and retains this attribution; no source corpus text is redistributed in the model package. - [Pinned MultiNLI dataset and license description](https://huggingface.co/datasets/nyu-mll/multi_nli/blob/da70db2af9d09693783c3320c4249840212ee221/README.md). - [MultiNLI paper](https://aclanthology.org/N18-1101/). MASSIVE and SLURP remain excluded from custom training, checkpoint selection and calibration. MASSIVE is used only for evaluation. Official Jev outputs are never used as training labels. The initial-release training histories above remain historical; the subsequent selected model and evaluation provenance identify the actual update. ## Natural-language decision adaptation This update adds 8,000 human-annotated training examples to 16,000 retained decision examples: 4,000 Cosmos QA reading questions, 2,000 SQuAD 2.0 answerability judgments, and 2,000 SNLI inference pairs. The selected natural-data checkpoints complete one pass of this 24,000-example mixture. Human source labels are preserved; SQuAD answerability uses its supplied impossible/answerable annotation, and each source is converted to the model's decision interface. Source-parent groups and near duplicates are separated between custom training, selection and calibration. No official Jev output supplies a training label. - **Cosmos QA**, by Lifu Huang, Ronan Le Bras, Chandra Bhagavatula and Yejin Choi. Data from the [author repository at the pinned revision](https://github.com/wilburOne/cosmosqa/tree/b6eb99cca4e2a51dd28a9a6f562534872d851639). The [official AllenAI dataset card](https://huggingface.co/datasets/allenai/cosmos_qa/blob/28d9d5e2aae025e73e11177891a88dba51190013/README.md) records the author-confirmed [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) license. - **SQuAD 2.0**, by Pranav Rajpurkar, Robin Jia and Percy Liang. The official [SQuAD project](https://rajpurkar.github.io/SQuAD-explorer/) distributes the dataset under [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/). The adaptation here is answerability classification, not the official span-extraction benchmark. - **SNLI 1.0**, by Samuel R. Bowman, Gabor Angeli, Christopher Potts and Christopher D. Manning. The official [Stanford Natural Language Inference project](https://nlp.stanford.edu/projects/snli/) and release README identify [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/) for the corpus. Original entailment, neutral and contradiction labels supply the three decision alternatives. These dataset licenses govern their respective source material; they are not replaced by the model package's Apache 2.0 license. Source corpus text, transformed training records and individual evaluation predictions are not redistributed in this model package. New evaluation panels, including supplied-fact QASC questions, are described separately in EVALUATION.md; evaluation data are not used for checkpoint selection or temperature fitting. As with other public datasets, exclusion from this custom training does not establish absence from upstream pretraining.