nicholasKluge commited on
Commit
adf1d17
·
verified ·
1 Parent(s): 3a52b1e

Update data_mixtures.md

Browse files
Files changed (1) hide show
  1. data_mixtures.md +58 -58
data_mixtures.md CHANGED
@@ -1,58 +1,58 @@
1
- # Data Mixtures for LilTii-v0.2
2
-
3
- ## Stage 1 (Warmup+Stable) Data Mixture
4
-
5
- For this stage, 40% is Bengali text (~40B tokens), 35% is educational English text (~35B tokens), 14.6% is reasoning-focused English text (~14.6B tokens), and 9.5% is educational math English text (~9.5B tokens). The detailed breakdown is as follows:
6
-
7
- | Dataset Name | Subset | Size (Tokens) | Repetition Factor |
8
- | ------------------------------------------------------------------------------------------------------------ | -------------- | ------------- | ----------------- |
9
- | [Polygl0t/gigakriya-v1](https://huggingface.co/datasets/Polygl0t/gigakriya-v1) | Edu Score of 1 | 5.87B | 2 |
10
- | | Edu Score of 2 | 8.62B | 2 |
11
- | | Edu Score of 3 | 4.25B | 2 |
12
- | | Edu Score of 4 | 1.52B | 2 |
13
- | | Edu Score of 5 | 5.50M | 2 |
14
- | [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) | Edu Score of 3 | 35.00B | 1 |
15
- | [HuggingFaceTB/finemath](https://huggingface.co/datasets/HuggingFaceTB/finemath) | Edu Score of 4 | 8.59B | 1 |
16
- | | Edu Score of 5 | 1.08B | 1 |
17
- | [allenai/big-reasoning-traces](https://huggingface.co/datasets/allenai/big-reasoning-traces) | All | 2.44B | 1 |
18
- | [allenai/math-meta-reasoning-filtered](https://huggingface.co/datasets/allenai/math-meta-reasoning-filtered) | All | 1.24B | 2 |
19
- | [nvidia/OpenScience](https://huggingface.co/datasets/nvidia/OpenScience) | All | 9.87B | 1 |
20
-
21
- During this stage, the learning rate follows a linear warmup for the first 2,000 steps, reaching a peak of 7e-4. It then remains stable at this peak for the next 47,500 steps before transitioning to the next stage.
22
-
23
- ## Stage 2 (Stable) Data Mixture
24
-
25
- For this stage, 40% is Bengali text (~40B tokens), 25% is synthetic English text (~25B tokens), 14% is educational English text (~14B tokens), 14.6% is reasoning-focused English text (~14.6B tokens), and 9.5% is educational math English text (~9.5B tokens).
26
-
27
- | Dataset Name | Subset | Size (Tokens) | Repetition Factor |
28
- | ------------------------------------------------------------------------------------------------------------ | -------------- | ------------- | ----------------- |
29
- | [Polygl0t/gigakriya-v1](https://huggingface.co/datasets/Polygl0t/gigakriya-v1) | Edu Score of 1 | 5.87B | 2 |
30
- | | Edu Score of 2 | 8.62B | 2 |
31
- | | Edu Score of 3 | 4.25B | 2 |
32
- | | Edu Score of 4 | 1.52B | 2 |
33
- | | Edu Score of 5 | 5.50M | 2 |
34
- | [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) | Edu Score of 4 | 14.22B | 1 |
35
- | [HuggingFaceTB/smollm-corpus](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus) (Cosmopedia v2) | All | 25.0B | 1 |
36
- | [HuggingFaceTB/finemath](https://huggingface.co/datasets/HuggingFaceTB/finemath) | Edu Score of 4 | 8.59B | 1 |
37
- | | Edu Score of 5 | 1.08B | 1 |
38
- | [allenai/big-reasoning-traces](https://huggingface.co/datasets/allenai/big-reasoning-traces) | All | 2.44B | 1 |
39
- | [allenai/math-meta-reasoning-filtered](https://huggingface.co/datasets/allenai/math-meta-reasoning-filtered) | All | 1.24B | 2 |
40
- | [nvidia/OpenScience](https://huggingface.co/datasets/nvidia/OpenScience) | All | 9.87B | 1 |
41
-
42
- During this stage, the learning rate remains stable at 7e-4 for the entire duration of 47,500 steps.
43
-
44
- ## Stage 3 (Stable+LinearDecay) Data Mixture
45
-
46
- For this stage, 50% is Bengali text (~15B tokens), 40% is synthetic English text (~12.5B tokens), ~1% is highly educational English text (~0.27B tokens), 8% is reasoning-focused English text (~2.4B tokens), and ~1% is highly-educational math English text (~1B tokens).
47
-
48
- | Dataset Name | Subset | Size (Tokens) | Repetition Factor |
49
- | ---------------------------------------------------------------------------------------------------------- | -------------- | ------------- | ----------------- |
50
- | [Polygl0t/gigakriya-v1](https://huggingface.co/datasets/Polygl0t/gigakriya-v1) | Edu Score of 3 | 4.25B | 3 |
51
- | | Edu Score of 4 | 1.52B | 2 |
52
- | | Edu Score of 5 | 5.50M | 3 |
53
- | [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) | Edu Score of 5 | 0.27B | 4 |
54
- | [HuggingFaceTB/smollm-corpus](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus) (Cosmopedia v2) | Half | 12.5B | 1 |
55
- | [HuggingFaceTB/finemath](https://huggingface.co/datasets/HuggingFaceTB/finemath) | Edu Score of 5 | 1.08B | 1 |
56
- | [allenai/big-reasoning-traces](https://huggingface.co/datasets/allenai/big-reasoning-traces) | All | 2.44B | 1 |
57
-
58
- During this stage, the learning rate starts at 7e-4 and remains stable for the first 3,000 steps. It then linearly decays to 0 over the remaining 12,000 steps. The decay phase covers approximately 25 billion tokens, about 10% of the total training tokens.
 
1
+ # Data Mixtures for LilTii-v0.2
2
+
3
+ ## Stage 1 (Warmup+Stable) Data Mixture
4
+
5
+ For this stage, 40% is Bengali text (40B tokens), 35% is educational English text (35B tokens), 14.6% is reasoning-focused English text (14.6B tokens), and 9.5% is educational math English text (9.5B tokens). The detailed breakdown is as follows:
6
+
7
+ | Dataset Name | Subset | Size (Tokens) | Repetition Factor |
8
+ | ------------------------------------------------------------------------------------------------------------ | -------------- | ------------- | ----------------- |
9
+ | [Polygl0t/gigakriya-v1](https://huggingface.co/datasets/Polygl0t/gigakriya-v1) | Edu Score of 1 | 5.87B | 2 |
10
+ | | Edu Score of 2 | 8.62B | 2 |
11
+ | | Edu Score of 3 | 4.25B | 2 |
12
+ | | Edu Score of 4 | 1.52B | 2 |
13
+ | | Edu Score of 5 | 5.50M | 2 |
14
+ | [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) | Edu Score of 3 | 35.00B | 1 |
15
+ | [HuggingFaceTB/finemath](https://huggingface.co/datasets/HuggingFaceTB/finemath) | Edu Score of 4 | 8.59B | 1 |
16
+ | | Edu Score of 5 | 1.08B | 1 |
17
+ | [allenai/big-reasoning-traces](https://huggingface.co/datasets/allenai/big-reasoning-traces) | All | 2.44B | 1 |
18
+ | [allenai/math-meta-reasoning-filtered](https://huggingface.co/datasets/allenai/math-meta-reasoning-filtered) | All | 1.24B | 2 |
19
+ | [nvidia/OpenScience](https://huggingface.co/datasets/nvidia/OpenScience) | All | 9.87B | 1 |
20
+
21
+ During this stage, the learning rate follows a linear warmup for the first 2,000 steps, reaching a peak of 7e-4. It then remains stable at this peak for the next 47,500 steps before transitioning to the next stage.
22
+
23
+ ## Stage 2 (Stable) Data Mixture
24
+
25
+ For this stage, 40% is Bengali text (40B tokens), 25% is synthetic English text (25B tokens), 14% is educational English text (14B tokens), 14.6% is reasoning-focused English text (14.6B tokens), and 9.5% is educational math English text (9.5B tokens).
26
+
27
+ | Dataset Name | Subset | Size (Tokens) | Repetition Factor |
28
+ | ------------------------------------------------------------------------------------------------------------ | -------------- | ------------- | ----------------- |
29
+ | [Polygl0t/gigakriya-v1](https://huggingface.co/datasets/Polygl0t/gigakriya-v1) | Edu Score of 1 | 5.87B | 2 |
30
+ | | Edu Score of 2 | 8.62B | 2 |
31
+ | | Edu Score of 3 | 4.25B | 2 |
32
+ | | Edu Score of 4 | 1.52B | 2 |
33
+ | | Edu Score of 5 | 5.50M | 2 |
34
+ | [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) | Edu Score of 4 | 14.22B | 1 |
35
+ | [HuggingFaceTB/smollm-corpus](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus) (Cosmopedia v2) | All | 25.0B | 1 |
36
+ | [HuggingFaceTB/finemath](https://huggingface.co/datasets/HuggingFaceTB/finemath) | Edu Score of 4 | 8.59B | 1 |
37
+ | | Edu Score of 5 | 1.08B | 1 |
38
+ | [allenai/big-reasoning-traces](https://huggingface.co/datasets/allenai/big-reasoning-traces) | All | 2.44B | 1 |
39
+ | [allenai/math-meta-reasoning-filtered](https://huggingface.co/datasets/allenai/math-meta-reasoning-filtered) | All | 1.24B | 2 |
40
+ | [nvidia/OpenScience](https://huggingface.co/datasets/nvidia/OpenScience) | All | 9.87B | 1 |
41
+
42
+ During this stage, the learning rate remains stable at 7e-4 for the entire duration of 47,500 steps.
43
+
44
+ ## Stage 3 (Stable+LinearDecay) Data Mixture
45
+
46
+ For this stage, 50% is Bengali text (15B tokens), 40% is synthetic English text (12.5B tokens), 1% is highly educational English text (0.27B tokens), 8% is reasoning-focused English text (2.4B tokens), and 1% is highly-educational math English text (1B tokens).
47
+
48
+ | Dataset Name | Subset | Size (Tokens) | Repetition Factor |
49
+ | ---------------------------------------------------------------------------------------------------------- | -------------- | ------------- | ----------------- |
50
+ | [Polygl0t/gigakriya-v1](https://huggingface.co/datasets/Polygl0t/gigakriya-v1) | Edu Score of 3 | 4.25B | 3 |
51
+ | | Edu Score of 4 | 1.52B | 2 |
52
+ | | Edu Score of 5 | 5.50M | 3 |
53
+ | [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) | Edu Score of 5 | 0.27B | 4 |
54
+ | [HuggingFaceTB/smollm-corpus](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus) (Cosmopedia v2) | Half | 12.5B | 1 |
55
+ | [HuggingFaceTB/finemath](https://huggingface.co/datasets/HuggingFaceTB/finemath) | Edu Score of 5 | 1.08B | 1 |
56
+ | [allenai/big-reasoning-traces](https://huggingface.co/datasets/allenai/big-reasoning-traces) | All | 2.44B | 1 |
57
+
58
+ During this stage, the learning rate starts at 7e-4 and remains stable for the first 3,000 steps. It then linearly decays to 0 over the remaining 12,000 steps. The decay phase covers approximately 25 billion tokens, about 10% of the total training tokens.