--- license: apache-2.0 base_model: answerdotai/ModernBERT-base datasets: - ymoslem/TeleMath-router pipeline_tag: text-classification tags: - quality-estimation - llm-routing - cascade - telecom --- # ymoslem/ModernBERT-base-TeleMath-router-qe-classifier-binary-10ep-lr2e-05-e4b-think-5runs Quality estimator for the Stage 2 cascade of CRE-Router on TeleMath. Gemma4-E4B thinking. Gates cluster 3 under the TPOT routing. Given an efficient model's answer, it predicts whether to **accept** it or **escalate** to a stronger model. ## Input format question [SEP] full_output [SEP] num_tokens `full_output` is truncated to its last 1,000 words before tokenisation, and the tokeniser runs at `max_length=4096`. Using 2048 truncates roughly a fifth of these inputs from the right, which removes the final answer and the token count, the two things the classifier most needs. ## Training ModernBERT-base, 10 epochs, lr 2e-5, max_length 4096, on the `train_e4b_think` split of [ymoslem/TeleMath-router](https://huggingface.co/datasets/ymoslem/TeleMath-router). Labels are recomputed from the saved generation text with the current grader, so they carry no dependence on the answer parser that was in use when the generations were produced. ## Intended use Deployed only where Stage 1 gates to the corresponding tier. It is trained on that tier's own outputs and does not transfer to another model's outputs.