--- language: - el - en license: apache-2.0 library_name: transformers pipeline_tag: text-generation base_model: glossAPI/apertus-8b-greek-cpt datasets: - glossAPI/apertus-8b-greek-cpt-modern-greek-train - glossAPI/greek-apertus-post-training-data - glossAPI/greek-apertus-sft-data tags: - apertus - greek - glossapi - instruct - sft - chat - v0.5-preview - card-eval-20261001 - public-three-benchmarks model-index: - name: Greek Apertus 8B Instruct results: - task: type: text-generation dataset: name: GreekMMLU (official, no chat template) type: greekmmlu-official-no-chat-template metrics: - type: accuracy name: accuracy (%) value: 70.52 - task: type: text-generation dataset: name: IFEval-el (prompt-strict) type: ifeval-el-prompt-strict metrics: - type: accuracy name: accuracy (%) value: 65.8 - task: type: text-generation dataset: name: GSM8K-el type: gsm8k-el metrics: - type: accuracy name: accuracy (%) value: 59.67 --- # Greek Apertus 8B Instruct v0.5 Greek Apertus is an 8B instruction model based on [Apertus 8B](https://huggingface.co/swiss-ai/Apertus-8B-2509). Greek continued pre-training followed by general instruction, Greek maths and conversation fine-tuning. Greek-extended vocabulary: 148,992 tokens. ## Benchmarks Accuracy (%). | Model | GreekMMLU | Greek IFEval | Greek GSM8K | |---|---:|---:|---:| | [Krikri 1.0](https://huggingface.co/ilsp/Llama-Krikri-8B-Instruct) | 67.94 | 63.77 | 71.72 | | [Krikri 1.5](https://huggingface.co/ilsp/Llama-Krikri-8B-Instruct-v1.5) | 67.69 | 54.90 | **73.01** | | [Apertus Instruct](https://huggingface.co/swiss-ai/Apertus-8B-Instruct-2509) | 66.50 | 54.90 | 56.18 | | [Greek Apertus — general instruction tuning](https://huggingface.co/glossAPI/greek-apertus-8b-instruct/tree/stage1-sft-S1) | **69.95** | **67.28** | 54.21 | | **Greek Apertus Instruct** | **70.52** | **65.80** | 59.67 | ## Training [Greek continued pre-training](https://huggingface.co/glossAPI/apertus-8b-greek-cpt/tree/base-avg30B-50B). | Stage | Checkpoint | Data | |---|---|---| | 1 | [Stage 1](https://huggingface.co/glossAPI/greek-apertus-8b-instruct/tree/stage1-sft-S1) | [Supervised mixtures](https://huggingface.co/datasets/glossAPI/greek-apertus-sft-data/tree/main/stage1-sft-S1) | | 2 | [Stage 2](https://huggingface.co/glossAPI/greek-apertus-8b-instruct/tree/stage2-sft-recipe3) | [Greek maths](https://huggingface.co/datasets/glossAPI/greek-apertus-post-training-data/tree/main/second_stage_sft/stage2-sft-recipe3) | | 3 | [Stage 3](https://huggingface.co/glossAPI/greek-apertus-8b-instruct/tree/stage3-sft-chat-P) | [Greek conversations](https://huggingface.co/datasets/glossAPI/greek-apertus-post-training-data/tree/main/second_stage_sft/stage3-sft-chat-P) | ## Repositories and checkpoints | Resource | Links | |---|---| | Greek base | [Repository](https://huggingface.co/glossAPI/apertus-8b-greek-cpt) | | Pre-training data | [Repository](https://huggingface.co/datasets/glossAPI/apertus-8b-greek-cpt-modern-greek-train) | | SFT data | [Repository](https://huggingface.co/datasets/glossAPI/greek-apertus-sft-data) | | Post-training data | [Repository](https://huggingface.co/datasets/glossAPI/greek-apertus-post-training-data) | | Original Apertus | [Repository](https://huggingface.co/swiss-ai/Apertus-8B-2509) | | Tokenizer | [Repository](https://huggingface.co/fffoivos/apertus-tokenizer-extension) | ## Licence Apache-2.0 for the model weights. ## Acknowledgements This work was implemented thanks to a grant by Swiss AI for compute on CSCS. Collection: [Greek Apertus 8B](https://huggingface.co/collections/glossAPI/greek-apertus-8b-v05-preview-6abecae3e0362dff3324b647).