Vega-1.6-42M-Tetris-Verbose
Note: AI was used in the creation of this project. But that's why you're here, isn't it?
This is the Tetris model for VegaLM1-42M, an experimental SLM trained on various datasets.
The first 262M tokens the model saw came from a 90M selection of fineweb-edu The next 688M was from a 50/25/15/10 split of fineweb-edu, fineweb, wikipedia, and a replay respectively, totaling around 420M tokens The final 2B tokens came from a 60/25/15 split of stack-v3-train, dclm-baseline-1.0, and smollm-corpus, trained for 1 epoch because I'm lazy
A new small finetuning corpus was constructed with data from https://huggingface.co/spaces/DedeProGames/SLM-Tetris-Arena, and this model is intended to dominate. And, oh hell yes, it dominates.
It is HILARIOUSLY good compared to all other models at playing Tetris, and it handily beats every model in the arena, no questions asked.
Benchmarks
Good at Tetris.
Model specs
- Layers: 12
- Hidden size: 512
- Attention heads: 8
- KV heads: 4
- Context length: 2,048 tokens
- Intermediate size: 880
- Vocabulary: 32,000
- Parameters: 42,054,144
- Downloads last month
- 3
Model tree for CNWPlayer/Vega-1.6-42M-Tetris-Verbose
Base model
CNWPlayer/VegaLM1-42M-Base