nicholasKluge commited on
Commit
316e25c
·
verified ·
1 Parent(s): e8fdeb4

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +30 -4
README.md CHANGED
@@ -229,6 +229,15 @@ tokenizer = AutoTokenizer.from_pretrained(model_id)
229
  model = AutoModelForCausalLM.from_pretrained(model_id, revision=revision)
230
  ```
231
 
 
 
 
 
 
 
 
 
 
232
  ## Intended Uses
233
 
234
  The primary intended use LilTii is to serve as foundations for research and development involving native Bengali language modeling. Checkpoints saved during training are designed to provide a controlled setting for performing comparative experiments, specifically regarding the effects of active pretraining on the performance of currently available benchmarks. You may also fine-tune and adapt LilTii for deployment if your use follows the Apache 2.0 license. If you decide to use LilTii as a basis for your fine-tuned model, please conduct your own risk and bias assessment.
@@ -348,11 +357,28 @@ The table below compares our two versions of LilTii against other base models of
348
 
349
  | | NPM (mean) | ARC Challenge | HellaSwag | MMLU | TruthfulQA MC1 | Bangla MMLU | BoolQ-BN | CommonsenseQA-BN | OpenBookQA-BN | PIQA-BN |
350
  |-----------------|------------|---------------|-----------|----------|----------------|-------------|----------|------------------|---------------|----------|
351
- | LilTii-v0.2 | 9.62724 | 0.261762 | 0.322008 | 0.270631 | 0.254802 | 0.260881 | 0.606481 | 0.324324 | 0.321932 | 0.605005 |
352
- | Qwen3-0.6B-Base | 8.0721 | 0.2284 | 0.286951 | 0.297865 | 0.271447 | 0.325695 | 0.657407 | 0.230139 | 0.305835 | 0.527203 |
353
- | Qwen2.5-0.5B | 6.03669 | 0.229256 | 0.284246 | 0.283144 | 0.28169 | 0.308407 | 0.578704 | 0.22113 | 0.305835 | 0.535909 |
354
- | LilTii-v0.1 | 5.81863 | 0.235244 | 0.303073 | 0.264088 | 0.240717 | 0.239525 | 0.520833 | 0.302211 | 0.327968 | 0.587051 |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
355
 
 
356
 
357
  ## Cite as 🤗
358
 
 
229
  model = AutoModelForCausalLM.from_pretrained(model_id, revision=revision)
230
  ```
231
 
232
+ <details>
233
+ <summary><b>Learning Curves</b></summary>
234
+
235
+ ![Learning Curves](./learning_curves.png)
236
+
237
+ This plot illustrates the evolution of model performance (measured by loss and perplexity) as a function of training time, measured in tokens seen during training.
238
+
239
+ </details>
240
+
241
  ## Intended Uses
242
 
243
  The primary intended use LilTii is to serve as foundations for research and development involving native Bengali language modeling. Checkpoints saved during training are designed to provide a controlled setting for performing comparative experiments, specifically regarding the effects of active pretraining on the performance of currently available benchmarks. You may also fine-tune and adapt LilTii for deployment if your use follows the Apache 2.0 license. If you decide to use LilTii as a basis for your fine-tuned model, please conduct your own risk and bias assessment.
 
357
 
358
  | | NPM (mean) | ARC Challenge | HellaSwag | MMLU | TruthfulQA MC1 | Bangla MMLU | BoolQ-BN | CommonsenseQA-BN | OpenBookQA-BN | PIQA-BN |
359
  |-----------------|------------|---------------|-----------|----------|----------------|-------------|----------|------------------|---------------|----------|
360
+ | LilTii-v0.2 | 9.63 | 26.18 | 32.20 | 27.06 | 25.48 | 26.09 | 60.65 | 32.43 | 32.19 | 60.50 |
361
+ | Qwen3-0.6B-Base | 8.07 | 22.84 | 28.70 | 29.79 | 27.14 | 32.57 | 65.74 | 23.01 | 30.58 | 52.72 |
362
+ | Qwen2.5-0.5B | 6.04 | 22.93 | 28.42 | 28.31 | 28.17 | 30.84 | 57.87 | 22.11 | 30.58 | 53.59 |
363
+ | LilTii-v0.1 | 5.82 | 23.52 | 30.31 | 26.41 | 24.07 | 23.95 | 52.08 | 30.22 | 32.80 | 58.71 |
364
+
365
+ <details>
366
+ <summary><b>Aggregate NPM Across Benchmarks</b></summary>
367
+
368
+ ![NPM vs Compute](./aggregate_npm.png)
369
+
370
+ This plot illustrates the evolution of model performance (measured by NPM mean) as a function of training time, measured in tokens seen during training. NPM scores are compared against two baseline models: Qwen2.5-0.5B and Qwen3-0.6B-Base, which are state-of-the-art multilingual models. The $sp$ (spearman correlation) between token ingestion and NPM mean is displayed, serving as an indicator of the monotonic relationship between tokens seen and performance.
371
+
372
+ </details>
373
+
374
+ <details>
375
+ <summary><b>Performance vs Compute</b></summary>
376
+
377
+ ![Performance vs Compute](./performance_vs_compute.png)
378
+
379
+ This plot compares the compute requirements (measured as $C = 6 * N * D$, where $N$ is the number of parameters and $D$ is the number of tokens processed) against the performance of each model (measured by NPM mean). It highlights the trade-offs between model size, training data, and performance. LilTii models are compared against two baseline models: Qwen2.5-0.5B and Qwen3-0.6B-Base, which are state-of-the-art multilingual models.
380
 
381
+ </details>
382
 
383
  ## Cite as 🤗
384