2026-09-02 新测试模型: Claude Fable 5.1, Mercury 2.5 Preview, Granite 4.2 8B, Gemini 3.8 Flash, Muse Spark 1.3
2026-08-15 新功能: dots-studio/dots-3-note-preview:free 已添加免费的 Dots3-Note Preview,用于可用性和推理模式的快速测试。 新功能: 本地运行的成本和价值比较现在使用 RTX 3090 的估算电费。
2026-08-14 新测试模型: Gemini 3.7 Flash, Qwen3.8 27B 新功能: 新增通过 Ollama 执行本地基准测试,并记录提供商来源、正确处理未定价的本地算力,同时显示清晰的本地运行标签。
2026-08-12 新测试模型: Seed 2.1 Turbo, Qwen3.8 2.4T A95B, Seed-2.0-Code, Grok 4.6 新功能: 新增了带来源依据的参数量、MoE 激活参数、模型开放状态和架构信息,并在模型页和对比页中展示,同时加入了新的排行榜筛选器。
2026-08-02 新功能: Model pages now show benchmark-suite coverage so incomplete results are clear at a glance.
2026-07-28 新测试模型: Trinity Large Thinking, Qwen3.7 Flash 新功能: Added site-wide Spotlight Search with keyboard navigation and a shortcut for finding pages, models, and comparisons.
2026-07-26 新功能: Added editorial model recommendations for affordable, intelligent, and fast choices. UX: Comparison pages now use draggable model summary cards with clearer model selection and ordering.
2026-07-17 新增测试: Added an executable JavaScript tool-calling benchmark with reference-case validation.
2026-07-16 新测试模型: Muse Spark 1.1, Kimi K3 新功能: The benchmark runner now supports executable tools with validated tool-call schemas and sandboxed execution.
2026-07-03 新功能: Launched AI World Cup, where benchmarked models generate football strategies and compete in simulated matches.
2026-06-17 新测试模型: GLM 5.2 Bug 修复: Adjusted missing-test handling so models are not scored as if unavailable tests were valid wrong answers. UX: Leaderboard search now supports comma-separated model queries, so searches like "deepseek, glm" show matches for either model family.
2026-06-12 新测试模型: Kimi K2.7 Code 新功能: Updated scoring to use per-category bias adjustments, so category-level differences are normalized before they roll into leaderboard results.
2026-06-05 新功能: Added model-generated visual showcases to model and comparison pages. 新功能: Added multi-model comparison pages with model recommendations and category-based ranking.
2026-05-27 新功能: Added current-price cost calculations while preserving original tested-at pricing for auditability.
2026-05-21 新测试模型: Grok Build 0.1 新增测试: Added a new benchmark test and improved answer judging. Bug 修复: 在提供商验证要求启用 reasoning 后,已移除不受支持的 xAI Grok Build 0.1 无 reasoning 变体。
2026-05-08 新增测试: Added a new benchmark test to expand suite coverage. Bug 修复: Reasoning chips and compare labels now recognize the minimal reasoning variant instead of falling back to auto. UX: Model pages now order sibling reasoning-variant chips from highest effort to lowest.
2026-04-28 新测试模型: Qwen3.5 Plus 2026-04-20, Qwen3.6 27B, Qwen3.6 35B A3B, Qwen3.6 Flash, Qwen3.6 Max Preview
2026-04-26 UX: 改进了移动端对比下拉菜单的位置,压缩了模型页面布局,并将运行历史拆分为按模型划分的分片,以减少页面加载的历史数据。 Bug 修复: 运行历史现在会合并同一测试套件中近似重复的重测,并在模型页面以直接对比表显示所有公开运行。
2026-04-24 新测试模型: DeepSeek V4 Flash 0423, DeepSeek V4 Pro, GPT-5.5 Bug 修复: 更新日志中的模型链接现在会解析到规范的在线模型页面,模型页面之间也会互相链接到不同推理变体。
2026-04-23 新测试模型: Ling-2.6-1T, Hy3 preview 新功能: 运行历史 - 模型页面现在会显示历史公开运行记录以及并排运行对比表。 (示例模型页面) UX: 排行榜现在支持基于 URL 的分页、筛选,以及从排名列表直接发起对比操作。 Bug 修复: 首页搜索、筛选计数和分页状态现在会在整个数据集范围内保持一致。 重新测试: GLM 5.1 已重新运行完整基准测试套件,并清理了该模型的公开运行历史快照。 Bug 修复: 已阻止未实际重新测试的无关模型获得新的 tested_at 时间戳。
2026-03-18 新测试模型: MiniMax M2.7 新功能: Added input and output pricing metrics to model and comparison pages.
2026-03-12 新测试模型: Nemotron 3 Super, Hunter Alpha, Grok 4.20, Grok 4.20 Beta, Grok 4.20 Multi Agent Beta 新功能: Added archived-model support across benchmark runs and public model displays. 新功能: Added generated Open Graph cards for richer shared links.
2026-03-06 新功能: Added the Methodology section and expanded charts with latency, cost, token, and model-switching views.
2026-03-04 新测试模型: GPT-5.2 Chat, GPT 5.3 Chat 新功能: Added interactive comparison charts for direct model analysis. 新增测试: Added a VAT compliance micro-audit benchmark with structured tool requirements.
2026-03-03 新测试模型: Trinity Large Preview, DeepSeek V3.2, Gemini 2.5 Flash, Gemini 3.1 Flash Lite, Gemini 3.1 Flash Lite Preview, GPT-5 Mini
2026-02-27 新测试模型: Seed-2.0-Mini 新功能: Added reasoning-quality and consistency charts to model comparisons.
2026-02-26 新测试模型: LFM2-24B-A2B, Qwen3.5-122B-A10B, Qwen3.5-27B, Qwen3.5-35B-A3B, Qwen3.5-Flash 新功能: Added responsive live leaderboard search by model, company, and reasoning mode.
2026-02-18 新测试模型: Claude Opus 4.6, Gemini 3 Flash Preview, Gemini 3 PRO Preview, GPT-5.2, MiMo-V2-Flash
2026-02-16 新测试模型: Kimi K2.5, GPT-5 Nano, gpt-oss-120b, Qwen3 Coder Next, Qwen3.5 Plus 2026-02-15, Step 3.5 Flash, GLM 4.7 Flash, GLM 5