Can an API-served model (OpenAI-compatible endpoint) be evaluated on v2?
4
#5 opened 20 days ago
by
dave-at-apmic
Duplicated Qwen3.6-27B row with contradictory scores, and 20 of 34 boards have an unseparated #1
1
#4 opened about 2 months ago
by
Ipezygj
Stuck in Pending Evaluation Queue
👍 1
2
#3 opened 3 months ago
by
flux-chinmay
Cost + latency + hallucination evaluation for Japanese LLM selection
3
#2 opened 4 months ago
by
vigneshwar234