Darwin-180B-RSI / .eval_results /lexam_hard.yaml
SeaWolf-AI's picture
eval: LEXam-hard 45.72
9eb2e46 verified
Raw History Blame Contribute Delete
596 Bytes
- dataset:
id: joelniklaus/LEXam-hard
task_id: lexam_hard
value: 45.72
date: '2026-09-30'
source:
url: https://huggingface.co/FINAL-Bench/Darwin-180B-RSI
name: Model Card
notes: "LEXam-hard, all 518 open questions; single sample (temperature=1.0, top_p=0.95, top_k=20); thinking budget 32,768 tokens, responses truncated at 32K regenerated with a 120K budget (60 items); judged by DeepSeek-R1-0528 with the LEXam paper judge prompt (eval.yaml); score = mean of German and English mean grades x 100 (de 44.37, en 47.07); 1 item without a parsable grade counted as 0; bf16"