Can an API-served model (OpenAI-compatible endpoint) be evaluated on v2?

#5
by dave-at-apmic - opened

Hi leaderboard team,

I originally sent this question to LLM-jp via their contact form, and the LLM-jp Secretariat asked me to post it here instead, since the leaderboard is hosted with Hugging Face's support.

Background

I'm Dave Sung, CTO and co-founder of APMIC, a Taiwan-based company that fine-tunes language models for enterprise customers. We run llm-jp-eval internally, so we're familiar with the suite.

Question

Is there any path for an API-served model (not open weights) to be evaluated on the Open Japanese LLM Leaderboard v2?

Our model ACE-3 is fine-tuned from google/gemma-4-26B-A4B-it, which is already on the leaderboard. However, ACE-3 is served through a commercial API rather than published as open weights. Since v2 deploys each submission to a Hugging Face inference endpoint, I suspect the answer is no — but I'd rather ask than assume.

What we can offer

If a path does exist, we would:

  • Provide a dedicated OpenAI-compatible endpoint, free and unmetered for the duration of the evaluation run
  • Accept whatever labelling or flagging you consider appropriate for a result that is not reproducible from open weights (e.g. a separate "API / closed-weights" category, or a clear marker on the row)

If there's no such path today, I'd also be interested to know whether it's something you'd consider for the future, and whether there's anything we could contribute to make it feasible.

Thanks for your time and for maintaining the leaderboard.

Dave Sung
CTO & Co-founder, APMIC
[email protected]

Hello @dave-at-apmic

Thank you for the message!

Researchers of the fine-tuning and evaluation team have been talking about your proposal for the OpenAI style API. Just to let you know, the philosophy of "Open Japanese LLM Leaderboard" is to evaluate open-source models. It might be a unfair advantage to evaluate to closed models via API calls. Plus, the computing power for the evaluations are given through academic grants.

Please be patient, I will come back to you as soon as possible.

Best,

Akim Mousterou

Hi Akim,

Thank you for discussing this with the team.

APMIC can provide a free, unmetered OpenAI-compatible endpoint, so the evaluation would not consume LLM-jp’s academic GPU resources.

We can also provide a preliminary self-evaluation report, including per-task scores, raw results, configurations, model version, and decoding parameters.

We understand the fairness concerns and are happy for the result to appear in a separate, clearly labelled “API / closed-weights” category rather than the main ranking.

Please let us know if this approach would be helpful.

Hello @dave-at-apmic ,

Thank you for the message and your generosity. I talked to the "Fine tuned and Evaluation" team again about your idea and APMIC. The topic will be debated for the next assembly by the various stakeholders of LLM-jp.

The Nejumi team ( @kyamamoto-nv , @olachinkei , etc. ) did an excellent job with the "Nejumi LLM Leaderboard 4". The leaderboard has a great evaluation taxonomy for corporate AI labs. You should probably try to evaluate your models with Nejumi.

Japanese version:
https://nejumi.ai/

English version:
https://wandb.ai/llm-leaderboard/nejumi-leaderboard4/reports/Nejumi-LLM-Leaderboaed-4--VmlldzoxNDQyMzkxMA 

I will keep your update anyway! Thank you again for your patience.

Best,

Akim

Hi Akim,

Thank you so much for raising this with the Fine-tuning and Evaluation team, and for putting it on the agenda for the next assembly. We really appreciate the time you've spent on it.

Thanks also for the pointer to Nejumi LLM Leaderboard 4. We'll reach out to the Nejumi team about evaluating our models there.

Looking forward to hearing the outcome of the assembly. Thanks again for your patience and support.

Best regards,
Dave

Sign up or log in to comment