vhallac commited on
Commit
7d79d86
Β·
verified Β·
1 Parent(s): 2f1d6b1

Add model card

Browse files
Files changed (1) hide show
  1. README.md +112 -0
README.md ADDED
@@ -0,0 +1,112 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Qwen/Qwen3-0.6B
4
+ datasets:
5
+ - HuggingFaceFW/fineweb-edu
6
+ language:
7
+ - en
8
+ tags:
9
+ - qwen3
10
+ - positional-encoding
11
+ - interpretability
12
+ - control-model
13
+ - research
14
+ library_name: transformers
15
+ ---
16
+
17
+ # Qwen3-0.6B β€” RoPE retained, recalibrated on 1B tokens (control model)
18
+
19
+ **A research artifact, not a general-purpose model.** This is Qwen3-0.6B recalibrated on 1B tokens
20
+ of FineWeb-Edu with RoPE left **active** throughout. It is the matched control for
21
+ [`vhallac/qwen3-0.6b-nope-recal-1b`](https://huggingface.co/vhallac/qwen3-0.6b-nope-recal-1b),
22
+ which received **identical** training β€” same corpus, same cached token stream, same step count,
23
+ same schedule, same seed β€” with RoPE removed.
24
+
25
+ The only difference between the two models is RoPE. That is the entire point of this checkpoint.
26
+
27
+ Loading is unremarkable β€” this is a standard Qwen3 model:
28
+
29
+ ```python
30
+ import torch
31
+ from transformers import AutoModelForCausalLM, AutoTokenizer
32
+
33
+ model_id = "vhallac/qwen3-0.6b-rope-recal-1b"
34
+ model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16)
35
+ tokenizer = AutoTokenizer.from_pretrained(model_id)
36
+ ```
37
+
38
+ (Its sibling **does** require a runtime patch. Read that model's card before using it.)
39
+
40
+ ## Why a control model exists
41
+
42
+ Part of the [`rope-as-scaffold`](https://github.com/vhallac/crockpot-experiments/tree/main/experiments/rope-as-scaffold)
43
+ research program, which tested whether **RoPE is a training scaffold that can be discarded after
44
+ pretraining**.
45
+
46
+ The problem this checkpoint solves: earlier comparisons pitted a RoPE-removed model that had
47
+ received 1B extra tokens against an untouched base model that had received none. Any difference
48
+ confounds *"RoPE was removed"* with *"received extra in-domain training."* Since the recalibration
49
+ corpus is also the evaluation domain, that confound is large and it runs in a direction that
50
+ flatters the RoPE-removed model.
51
+
52
+ This model holds the training constant so the comparison isolates RoPE. It changed the program's
53
+ conclusions: one prior result reversed sign, another was halved in magnitude, and a persistent
54
+ perplexity penalty became visible that had been invisible against the un-adapted baseline.
55
+
56
+ ## Results
57
+
58
+ Held-out FineWeb-Edu perplexity, fp32, 5M-token frozen eval slice, identical harness for all three:
59
+
60
+ | model | CE | PPL |
61
+ |---|---:|---:|
62
+ | `Qwen/Qwen3-0.6B` (base, no extra training) | 3.0819 | 21.80 |
63
+ | **this model** (control, RoPE on) | 2.6569 | **14.25** |
64
+ | `qwen3-0.6b-nope-recal-1b` (RoPE removed) | 2.8260 | **16.88** |
65
+
66
+ Reading these correctly matters. Recalibration is worth **0.425 nats** (3.0819 β†’ 2.6569). Removing
67
+ RoPE gives **0.169 nats** of that back (2.6569 β†’ 2.8260). The RoPE-removed model's apparent
68
+ improvement over base (21.80 β†’ 16.88) is the net of those two, and reporting it against base alone
69
+ makes a real regression look like a gain.
70
+
71
+ ## Known limitations
72
+
73
+ - **Effective context β‰ˆ 2048 tokens** for the *recalibrated* behaviour, despite
74
+ `max_position_embeddings: 40960` inherited from base. Unlike its NoPE sibling this model does not
75
+ collapse beyond 2048 (it inherits base's 32k-context pretraining), but its recalibration only
76
+ covered 2048.
77
+ - English-only recalibration corpus (FineWeb-Edu), while Qwen3 is multilingual + code + math.
78
+ Expect domain skew relative to base outside the eval domain.
79
+ - Not instruction-tuned beyond whatever base carried; recalibration was plain LM training.
80
+ - Single seed, single recipe, 0.6B scale.
81
+
82
+ ## Training
83
+
84
+ | | |
85
+ |---|---|
86
+ | base | `Qwen/Qwen3-0.6B` @ `c1899de289a04d12100db370d81485cdf75e47ca` |
87
+ | rotary | **active throughout** (no patch applied) |
88
+ | corpus | `HuggingFaceFW/fineweb-edu`, `sample-10BT`, streamed in provider order |
89
+ | tokens | 1B (1907 steps Γ— 524,288 tokens), context 2048 |
90
+ | optimizer | AdamW Ξ²=(0.9, 0.95), wd 0.1, grad clip 1.0 |
91
+ | LR | 1e-3 peak, 2% warmup, cosine β†’ 10% of peak |
92
+ | precision | bf16 |
93
+ | seed | 0 |
94
+ | hardware | 1Γ— H100 SXM, ~7h |
95
+
96
+ Trained from the **same cached token file** as its NoPE sibling, not merely the same corpus β€” the
97
+ two models saw the identical token sequence in the identical order. `training_metrics.csv` and
98
+ `training_manifest.json` are included for full provenance.
99
+
100
+ > Note: `training_manifest.json` in this repo carries `"artifact": "qwen3-droped"` β€” a known
101
+ > mislabel from a hardcoded string in the training script. Every functional field is correct for
102
+ > *this* model, including `rotary_patch_applied: false` and `output_dir: /workspace/qwen3-rope-recal`.
103
+
104
+ ## Reproduction and full analysis
105
+
106
+ - Code, specs, and lab notebook: [vhallac/crockpot-experiments](https://github.com/vhallac/crockpot-experiments/tree/main/experiments/rope-as-scaffold)
107
+ - The controlled comparison this model enabled: `NOTEBOOK.md`, entry `2026-07-28 β€” RS-amendment-2-3`
108
+ - Pre-registered plan: `RS-amendment-2-3.md`; recipe of record `RS1-spec.md` Β§11
109
+
110
+ ## License
111
+
112
+ Apache 2.0, inherited from Qwen3-0.6B.