AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Changelog

A simple log of product and benchmark updates, grouped by date. We use it to note newly tested models, re-tests, benchmark changes, and shipped UX/product work.

2026-09-05

  • New Feature: Copy SVG showcases as images or download watermarked PNGs. Previews now show model rank.
  • Bug Fix: Restored valid SVG showcases and clarified timeout errors.

2026-08-20

2026-08-15

  • New Feature: dots-studio/dots-3-note-preview:free Added Dots3-Note Preview (free) for availability and reasoning-mode smoke tests.
  • New Feature: Local runs now use estimated RTX 3090 electricity costs in cost and value comparisons.

2026-08-14

  • New Models Tested: Gemini 3.7 Flash, Qwen3.8 27B
  • New Feature: Added local Ollama benchmark execution with provider provenance, unpriced local-compute handling, and clear Local run labels.

2026-08-12

2026-08-02

  • New Feature: Model pages now show benchmark-suite coverage so incomplete results are clear at a glance.

2026-07-28

  • New Models Tested: Trinity Large Thinking, Qwen3.7 Flash
  • New Feature: Added site-wide Spotlight Search with keyboard navigation and a shortcut for finding pages, models, and comparisons.

2026-07-26

  • New Feature: Added editorial model recommendations for affordable, intelligent, and fast choices.
  • UX: Comparison pages now use draggable model summary cards with clearer model selection and ordering.

2026-07-18

2026-07-17

  • New Tests Added: Added an executable JavaScript tool-calling benchmark with reference-case validation.

2026-07-16

  • New Models Tested: Muse Spark 1.1, Kimi K3
  • New Feature: The benchmark runner now supports executable tools with validated tool-call schemas and sandboxed execution.

2026-07-03

  • New Feature: Launched AI World Cup, where benchmarked models generate football strategies and compete in simulated matches.

2026-06-17

  • New Models Tested: GLM 5.2
  • Bug Fix: Adjusted missing-test handling so models are not scored as if unavailable tests were valid wrong answers.
  • UX: Leaderboard search now supports comma-separated model queries, so searches like "deepseek, glm" show matches for either model family.

2026-06-16

  • New Feature: Added cost sorting and filtering across leaderboard and category views.

2026-06-12

  • New Models Tested: Kimi K2.7 Code
  • New Feature: Updated scoring to use per-category bias adjustments, so category-level differences are normalized before they roll into leaderboard results.

2026-06-06

  • New Feature: Model showcases now support shareable lightbox views, filtering, and score details.

2026-06-05

  • New Feature: Added model-generated visual showcases to model and comparison pages.
  • New Feature: Added multi-model comparison pages with model recommendations and category-based ranking.

2026-06-04

  • New Models Tested: Nemotron 3 Ultra
  • New Tests Added: Added a coding benchmark with executable reference cases.

2026-05-27

  • New Feature: Added current-price cost calculations while preserving original tested-at pricing for auditability.

2026-05-22

  • New Models Tested: Qwen3.7 Max
  • New Tests Added: Added a new Coding test category focused on bug-finding in C++ solutions.

2026-05-21

  • New Models Tested: Grok Build 0.1
  • New Tests Added: Added a new benchmark test and improved answer judging.
  • Bug Fix: Removed the unsupported xAI Grok Build 0.1 no-reasoning variant after provider validation required reasoning.

2026-05-08

  • New Tests Added: Added a new benchmark test to expand suite coverage.
  • Bug Fix: Reasoning chips and compare labels now recognize the minimal reasoning variant instead of falling back to auto.
  • UX: Model pages now order sibling reasoning-variant chips from highest effort to lowest.

2026-05-06

2026-04-26

  • UX: Improved mobile compare dropdown placement, tightened model page layout, and split run history into per-model shards so pages load less historical data.
  • Bug Fix: Run history now groups near-duplicate same-suite retests and shows all public runs in a direct comparison table on model pages.

2026-04-25

  • New Feature: Added Reliability score telemetry so target API and rate-limit failures are tracked separately from wrong answers.

2026-04-23

  • New Models Tested: Ling-2.6-1T, Hy3 preview
  • New Feature: Run history - Model pages now show historical public runs and a side-by-side run comparison table. (Example model page)
  • UX: The leaderboard now supports URL-backed pagination, filters, and direct compare actions from the ranking list.
  • Bug Fix: Homepage search, filter counts, and pagination state now stay consistent across the full dataset.
  • Re-test: GLM 5.1 Reran the full benchmark suite and cleaned up the public run-history snapshot for this model.
  • Bug Fix: Stopped unrelated models from receiving a fresh tested_at timestamp when they were not actually retested.

2026-04-11

2026-03-18

  • New Models Tested: MiniMax M2.7
  • New Feature: Added input and output pricing metrics to model and comparison pages.

2026-03-06

  • New Feature: Added the Methodology section and expanded charts with latency, cost, token, and model-switching views.

2026-03-04

  • New Models Tested: GPT-5.2 Chat, GPT 5.3 Chat
  • New Feature: Added interactive comparison charts for direct model analysis.
  • New Tests Added: Added a VAT compliance micro-audit benchmark with structured tool requirements.

2026-03-02

  • New Feature: Added share actions with copyable links across model and comparison pages.

2026-02-27

  • New Models Tested: Seed-2.0-Mini
  • New Feature: Added reasoning-quality and consistency charts to model comparisons.

Changelog page created

We started this changelog after launch, so some older updates are missing.