AI BENCHY
Advertise here

更新日志

一个按日期分组的产品与基准更新简明记录。我们用它记录新测试的模型、重新测试、基准变更以及已经发布的 UX/产品工作。

2026-09-05

  • 新功能: 可将 SVG 展示作品复制为图片,或下载带水印的 PNG。预览现在显示模型排名。
  • Bug 修复: 恢复了有效的 SVG 展示作品,并明确了超时错误信息。

2026-08-20

2026-08-15

  • 新功能: dots-studio/dots-3-note-preview:free 已添加免费的 Dots3-Note Preview,用于可用性和推理模式的快速测试。
  • 新功能: 本地运行的成本和价值比较现在使用 RTX 3090 的估算电费。

2026-08-14

  • 新测试模型: Gemini 3.7 Flash, Qwen3.8 27B
  • 新功能: 新增通过 Ollama 执行本地基准测试,并记录提供商来源、正确处理未定价的本地算力,同时显示清晰的本地运行标签。

2026-08-12

  • 新测试模型: Seed 2.1 Turbo, Qwen3.8 2.4T A95B, Seed-2.0-Code, Grok 4.6
  • 新功能: 新增了带来源依据的参数量、MoE 激活参数、模型开放状态和架构信息,并在模型页和对比页中展示,同时加入了新的排行榜筛选器。

2026-08-02

  • 新功能: Model pages now show benchmark-suite coverage so incomplete results are clear at a glance.

2026-07-28

  • 新测试模型: Trinity Large Thinking, Qwen3.7 Flash
  • 新功能: Added site-wide Spotlight Search with keyboard navigation and a shortcut for finding pages, models, and comparisons.

2026-07-26

  • 新功能: Added editorial model recommendations for affordable, intelligent, and fast choices.
  • UX: Comparison pages now use draggable model summary cards with clearer model selection and ordering.

2026-07-18

2026-07-17

  • 新增测试: Added an executable JavaScript tool-calling benchmark with reference-case validation.

2026-07-16

  • 新测试模型: Muse Spark 1.1, Kimi K3
  • 新功能: The benchmark runner now supports executable tools with validated tool-call schemas and sandboxed execution.

2026-07-03

  • 新功能: Launched AI World Cup, where benchmarked models generate football strategies and compete in simulated matches.

2026-06-17

  • 新测试模型: GLM 5.2
  • Bug 修复: Adjusted missing-test handling so models are not scored as if unavailable tests were valid wrong answers.
  • UX: Leaderboard search now supports comma-separated model queries, so searches like "deepseek, glm" show matches for either model family.

2026-06-16

  • 新功能: Added cost sorting and filtering across leaderboard and category views.

2026-06-12

  • 新测试模型: Kimi K2.7 Code
  • 新功能: Updated scoring to use per-category bias adjustments, so category-level differences are normalized before they roll into leaderboard results.

2026-06-06

  • 新功能: Model showcases now support shareable lightbox views, filtering, and score details.

2026-06-05

  • 新功能: Added model-generated visual showcases to model and comparison pages.
  • 新功能: Added multi-model comparison pages with model recommendations and category-based ranking.

2026-06-04

  • 新测试模型: Nemotron 3 Ultra
  • 新增测试: Added a coding benchmark with executable reference cases.

2026-05-27

  • 新功能: Added current-price cost calculations while preserving original tested-at pricing for auditability.

2026-05-22

  • 新测试模型: Qwen3.7 Max
  • 新增测试: 新增一个 Coding 测试类别,专注于在 C++ 解决方案中查找错误。

2026-05-21

  • 新测试模型: Grok Build 0.1
  • 新增测试: Added a new benchmark test and improved answer judging.
  • Bug 修复: 在提供商验证要求启用 reasoning 后,已移除不受支持的 xAI Grok Build 0.1 无 reasoning 变体。

2026-05-08

  • 新增测试: Added a new benchmark test to expand suite coverage.
  • Bug 修复: Reasoning chips and compare labels now recognize the minimal reasoning variant instead of falling back to auto.
  • UX: Model pages now order sibling reasoning-variant chips from highest effort to lowest.

2026-05-06

2026-04-26

  • UX: 改进了移动端对比下拉菜单的位置,压缩了模型页面布局,并将运行历史拆分为按模型划分的分片,以减少页面加载的历史数据。
  • Bug 修复: 运行历史现在会合并同一测试套件中近似重复的重测,并在模型页面以直接对比表显示所有公开运行。

2026-04-25

  • 新功能: 新增可靠性分数遥测,将目标 API 和速率限制失败与错误答案分开跟踪。

2026-04-24

2026-04-23

  • 新测试模型: Ling-2.6-1T, Hy3 preview
  • 新功能: 运行历史 - 模型页面现在会显示历史公开运行记录以及并排运行对比表。 (示例模型页面)
  • UX: 排行榜现在支持基于 URL 的分页、筛选,以及从排名列表直接发起对比操作。
  • Bug 修复: 首页搜索、筛选计数和分页状态现在会在整个数据集范围内保持一致。
  • 重新测试: GLM 5.1 已重新运行完整基准测试套件,并清理了该模型的公开运行历史快照。
  • Bug 修复: 已阻止未实际重新测试的无关模型获得新的 tested_at 时间戳。

2026-04-11

2026-03-18

  • 新测试模型: MiniMax M2.7
  • 新功能: Added input and output pricing metrics to model and comparison pages.

2026-03-06

  • 新功能: Added the Methodology section and expanded charts with latency, cost, token, and model-switching views.

2026-03-04

  • 新测试模型: GPT-5.2 Chat, GPT 5.3 Chat
  • 新功能: Added interactive comparison charts for direct model analysis.
  • 新增测试: Added a VAT compliance micro-audit benchmark with structured tool requirements.

2026-03-02

  • 新功能: Added share actions with copyable links across model and comparison pages.

2026-02-27

  • 新测试模型: Seed-2.0-Mini
  • 新功能: Added reasoning-quality and consistency charts to model comparisons.

更新日志页面已创建

这个更新日志是在上线后才开始记录的,所以部分更早的更新没有列出。