2026-06-17
- Modele nou testate: GLM 5.2, Kimi K2.7 Code, Claude Fable 5, Nemotron 3 Ultra, Qwen3.7 Plus, MiniMax M3, Step 3.7 Flash, Claude Opus 4.8 Added benchmark coverage for newly released models missing from the changelog: Z.ai GLM 5.2, MoonshotAI Kimi K2.7 Code, Anthropic Claude Fable 5, NVIDIA Nemotron 3 Ultra 550B A55B, Qwen 3.7 Plus, MiniMax M3, StepFun Step 3.7 Flash, and Anthropic Claude Opus 4.8.
- Funcționalitate nouă: Updated scoring to use per-category bias adjustments, so category-level differences are normalized before they roll into leaderboard results.
- Remediere de bug: Adjusted missing-test handling so models are not scored as if unavailable tests were valid wrong answers.
- UX: Leaderboard search now supports comma-separated model queries, so searches like "deepseek, glm" show matches for either model family.