Duplicated Qwen3.6-27B row with contradictory scores, and 20 of 34 boards have an unseparated #1
Two things from leaderboard.parquet: a duplicated model row with contradictory scores, and 20 of 34 boards whose #1 isn't separated from #2.
First, credit where it's due, because it's load-bearing here. You publish scores at full float precision, so every denominator can be recovered exactly — Fraction(v).limit_denominator() on the published values returns 83/198 for niilc, 6511 for BigBenchHard, 7097 for JMMLU, 14042 for MMLU-en, and so on for all 37 *_exact_match columns. I've looked at several other leaderboards this week that publish one decimal, and none of them can be checked from outside at all. This one can, entirely because of that choice.
1. Qwen/Qwen3.6-27B appears twice, and the two rows disagree
leaderboard.parquet has 21 rows but only 20 distinct (model, revision) pairs. Qwen/Qwen3.6-27B at revision main is present twice, and the two rows disagree on 138 of the 175 numeric columns:
| column | row A | row B |
|---|---|---|
mmlu_en_exact_match |
0.000000 | 0.856003 |
jmmlu_exact_match |
0.101029 | 0.806115 |
gpqa_main_ja_exact_match |
0.002237 | 0.483221 |
drop_exact_match |
0.600000 | 0.653382 |
One row carries 26 exact zeros across the numeric columns and the other 42. A score of exactly 0.0 on MMLU-English next to 0.856 for the same model and revision looks like one of the two is a failed or partial run rather than a real result. Whichever row the front end happens to take is the one users see.
2. Rank separation
With the denominators recovered above, SE = sqrt(p(1-p)/n) and 20,000 simulated draws per board:
| board | n | #1 | #2 | gap in leader's SE | P(measured #1 isn't best) |
|---|---|---|---|---|---|
| niilc | 198 | 41.92% | 41.92% | 0.00σ | 67.2% |
| drop | 9535 | 65.39% | 65.34% | 0.11σ | 62.9% |
| gpqa_main_ja | 447 | 48.77% | 48.32% | 0.19σ | 47.1% |
| gpqa_extended_ja | 542 | 47.05% | 46.49% | 0.26σ | 54.3% |
| jemhopqa | 120 | 54.17% | 52.50% | 0.37σ | 73.2% |
| jsick | 4927 | 84.64% | 84.43% | 0.40σ | 43.4% |
| … | |||||
| jmmlu | 7097 | 84.35% | 81.15% | 7.42σ | 0.0% |
| mmlu_prox_en | 11759 | 71.79% | 64.76% | 16.95σ | 0.0% |
| mmlu_prox_ja | 11759 | 67.78% | 59.10% | 20.15σ | 0.0% |
20 of 34 boards have a #1 within 1.35σ of #2. niilc is an exact tie — Gemma-2-Llama-Swallow-27b-it and llm-jp-3-8x13b-instruct3 are both exactly 83/198, which the full precision makes verifiable rather than a guess.
The pattern worth flagging: several of the unseparated pairs are successive versions of the same model. On drop and gpqa_main_ja the top two are Qwen3.5-27B and Qwen3.6-27B; on jamc-qa they are llm-jp-3-8x13b-instruct3 and llm-jp-3.1-8x13b-instruct4. Those are exactly the comparisons a reader is most likely to draw a conclusion from — "the newer one is better" — and on these boards the data doesn't support an ordering either way.
The large boards behave completely differently: MMLU-ProX, JMMLU, MMMLU and TriviaQA all separate cleanly at 7–20σ. So this isn't "the leaderboard is noise" — it's that the small boards can't resolve what the table displays as a strict order.
Reproduce
import pandas as pd, math
from fractions import Fraction
df = pd.read_parquet("hf://datasets/llm-jp/leaderboard-contents-v2/leaderboard.parquet")
print("rows", len(df), "distinct (model,revision)", df[["model","revision"]].drop_duplicates().shape[0])
dup = df[df.duplicated(["model","revision"], keep=False)]
print(dup[["model","mmlu_en_exact_match","gpqa_main_ja_exact_match"]].to_string(index=False))
s = df[["model","niilc_exact_match"]].dropna().sort_values("niilc_exact_match", ascending=False)
top = s["niilc_exact_match"].iloc[0]
print("niilc top value as a fraction:", Fraction(float(top)).limit_denominator(20000),
"| tied models:", (s["niilc_exact_match"] == top).sum())
Limits
The published score is treated as truth and noise as independent across models — the standard framing, and a model rather than a fact. The denominators are recovered as the LCM of the reduced fractions per column, so a column where every model happened to land on a coarser fraction could in principle be a multiple of the true size; that would make the SEs conservative, not anti-conservative.
Suggestion
The duplicate row looks like a straightforward data fix. For the second point, significance tiers or a CI column would stop the table asserting an order the small boards can't support — no score would change. Happy to be told I've misread something about how the runs are aggregated, in which case I'll say so here.
Hello @Ipezygj
I hope you are doing well,
Thank you for your observations and great questions!
Yes, we are using full float precision, and we believe that transparency is key for good evaluation of OSS LLMs. 😊
Please keep in mind that the leaderboard has many evaluation options. For example, "Qwen/Qwen3.6-27B", the "Apply Chat Template" setting differs from RL-tuned to multimodal (or even the vLLM version in the backend). It may be necessary to find a way to properly reflect the evaluation settings in the "Run Name". At the very least, as you pointed out, the issue of "different evaluation scores despite duplicate evaluations of the same model" is not occurring.
This leaderboard is not intended to provide a pure ranking with winners and losers, but to study strengths and weaknesses of various open source LLMs. Furthermore, NIILC evaluation data is originally limited from 2003 (https://mynlp.is.s.u-tokyo.ac.jp/niilc-qa/). This is something that could happen by chance. In the near future, llm-jp-eval (evaluation framework behind the Open Japanese LLM Leaderboard) is already preparing to add evaluation data such as Tool Calling, that will clearly differentiate models, so I believe this is a problem that will naturally be resolved over time.
You can find more information about datasets, evaluation style, and philosophy directly on the llm-jp-eval github here:
https://github.com/llm-jp/llm-jp-eval/blob/dev/README_en.md
Thank you again, please feel free to launch evaluations on interesting LLMs! 🤗
Best,
Akim