SozKZ MoE: Mixture of Experts
Collection
Mixture-of-Experts models for Kazakh — upcycled and domain-pretrained MoE architectures • 4 items • Updated
A Kazakh language model trained from scratch on a curated, deduplicated Kazakh text corpus (~1B tokens). Llama architecture with Chinchilla-optimal compute budget.
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("stukenov/kazakh-moe-160M-A50M-domain")
model = AutoModelForCausalLM.from_pretrained("stukenov/kazakh-moe-160M-A50M-domain")
text = "Қазақстан — "
inputs = tokenizer(text, return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=100)
print(tokenizer.decode(output[0], skip_special_tokens=True))
| Parameter | Value |
|---|---|
| Type | mixtral |
| Parameters | ~51M |
| Hidden size | 512 |
| Layers | 8 |
| Attention heads | 8 |
| Vocab size | 50,257 |
| Context length | 1024 |
Trained on stukenov/kazakh-clean-pretrain-v2 — a curated Kazakh corpus processed through a 9-stage cleaning pipeline:
Sources: CC-100, OSCAR, Wikipedia, Leipzig, Kazakh News, Kazakh Books.
See the SLM project for full experiment details.