AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

変更履歴

日付ごとにまとめた、製品とベンチマーク更新のシンプルな記録です。新たにテストしたモデル、再テスト、ベンチマーク変更、公開済みの UX/製品改善をここに記録します。

2026-09-08

  • 新機能: ライトモードは紙のような配色、ダークモードはグラファイト調に。コンパクトな比較ヘッダー、わかりやすいモデル選択、モバイルで省スペースのおすすめ表示を備えています。
  • バグ修正: メニュー項目の文字の重なりと、ドロップダウンが切れる問題を修正しました。

2026-09-07

  • 新機能: サイト全体を、読みやすい文字、統一された操作部品、落ち着いた配色のライト・ダークテーマに刷新しました。モデル検索、比較の操作、モバイルの指標カードは引き続き利用できます。
  • 新しくテストしたモデル: Qwen3.8 Max (0902)

2026-09-05

  • 新機能: SVG作品を画像としてコピーしたり、透かし付きPNGとしてダウンロードできるようになりました。プレビューにはモデル順位が表示されます。
  • バグ修正: 有効なSVG作品の表示を復旧し、タイムアウトのエラー表示を明確にしました。

2026-09-04

2026-08-29

2026-08-20

  • 新しくテストしたモデル: GLM 5.3

2026-08-15

  • 新機能: dots-studio/dots-3-note-preview:free 可用性と推論モードの簡易テスト用に Dots3-Note Preview(無料)を追加しました。
  • 新機能: ローカル実行のコストと価値の比較に RTX 3090 の推定電気料金を使用するようになりました。

2026-08-14

  • 新しくテストしたモデル: Gemini 3.7 Flash, Qwen3.8 27B
  • 新機能: Ollama によるローカルベンチマーク実行、プロバイダーの来歴、価格未設定のローカル計算資源の適切な処理、明確なローカル実行タグを追加しました。

2026-08-12

  • 新しくテストしたモデル: Seed 2.1 Turbo, Qwen3.8 2.4T A95B, Seed-2.0-Code, Grok 4.6
  • 新機能: 出典付きのパラメータ数、MoE のアクティブパラメータ、公開状況、アーキテクチャをモデルページと比較ページに追加し、ランキングに新しいフィルターを追加しました。

2026-08-04

2026-08-02

  • 新機能: Model pages now show benchmark-suite coverage so incomplete results are clear at a glance.

2026-07-28

  • 新しくテストしたモデル: Trinity Large Thinking, Qwen3.7 Flash
  • 新機能: Added site-wide Spotlight Search with keyboard navigation and a shortcut for finding pages, models, and comparisons.

2026-07-26

  • 新機能: Added editorial model recommendations for affordable, intelligent, and fast choices.
  • UX: Comparison pages now use draggable model summary cards with clearer model selection and ordering.

2026-07-22

2026-07-20

2026-07-18

  • 新しくテストしたモデル: Inkling

2026-07-17

  • 追加された新テスト: Added an executable JavaScript tool-calling benchmark with reference-case validation.

2026-07-16

  • 新しくテストしたモデル: Muse Spark 1.1, Kimi K3
  • 新機能: The benchmark runner now supports executable tools with validated tool-call schemas and sandboxed execution.

2026-07-03

  • 新機能: Launched AI World Cup, where benchmarked models generate football strategies and compete in simulated matches.

2026-06-17

  • 新しくテストしたモデル: GLM 5.2
  • バグ修正: Adjusted missing-test handling so models are not scored as if unavailable tests were valid wrong answers.
  • UX: Leaderboard search now supports comma-separated model queries, so searches like "deepseek, glm" show matches for either model family.

2026-06-16

  • 新機能: Added cost sorting and filtering across leaderboard and category views.

2026-06-12

  • 新しくテストしたモデル: Kimi K2.7 Code
  • 新機能: Updated scoring to use per-category bias adjustments, so category-level differences are normalized before they roll into leaderboard results.

2026-06-06

  • 新機能: Model showcases now support shareable lightbox views, filtering, and score details.

2026-06-05

  • 新機能: Added model-generated visual showcases to model and comparison pages.
  • 新機能: Added multi-model comparison pages with model recommendations and category-based ranking.

2026-06-04

  • 新しくテストしたモデル: Nemotron 3 Ultra
  • 追加された新テスト: Added a coding benchmark with executable reference cases.

2026-06-03

2026-06-01

2026-05-27

  • 新機能: Added current-price cost calculations while preserving original tested-at pricing for auditability.

2026-05-22

  • 新しくテストしたモデル: Qwen3.7 Max
  • 追加された新テスト: C++ ソリューションのバグ発見に焦点を当てた新しい Coding テストカテゴリを追加しました。

2026-05-21

  • 新しくテストしたモデル: Grok Build 0.1
  • 追加された新テスト: Added a new benchmark test and improved answer judging.
  • バグ修正: プロバイダー検証で reasoning が必須とされたため、サポートされていない xAI Grok Build 0.1 の no-reasoning バリアントを削除しました。

2026-05-10

2026-05-08

  • 追加された新テスト: Added a new benchmark test to expand suite coverage.
  • バグ修正: Reasoning chips and compare labels now recognize the minimal reasoning variant instead of falling back to auto.
  • UX: Model pages now order sibling reasoning-variant chips from highest effort to lowest.

2026-05-06

  • 新しくテストしたモデル: Cobuddy

2026-04-30

  • 新しくテストしたモデル: Owl Alpha

2026-04-26

  • UX: モバイルの比較ドロップダウン位置を改善し、モデルページのレイアウトを引き締め、実行履歴をモデル別シャードに分割してページが読み込む履歴データを減らしました。
  • バグ修正: 実行履歴では同じテストスイートのほぼ重複する再テストをまとめ、モデルページで全公開実行を直接比較テーブルとして表示するようになりました。

2026-04-25

  • 新機能: 信頼性スコアのテレメトリを追加し、対象APIとレート制限の失敗を誤答とは別に追跡するようにしました。

2026-04-24

  • 新しくテストしたモデル: DeepSeek V4 Flash 0423, DeepSeek V4 Pro, GPT-5.5
  • バグ修正: 変更履歴のモデルリンクは正規の公開モデルページに解決されるようになり、モデルページ間でも推論バリアントを相互に移動できるようになりました。

2026-04-23

  • 新しくテストしたモデル: Ling-2.6-1T, Hy3 preview
  • 新機能: 実行履歴 - モデルページで過去の公開実行履歴と実行同士の並列比較テーブルを表示するようになりました。 (モデルページの例)
  • UX: リーダーボードは URL ベースのページネーション、フィルター、ランキング一覧からの直接比較操作に対応しました。
  • バグ修正: トップページの検索、フィルター件数、ページネーション状態がデータセット全体で一貫して保たれるようになりました。
  • 再テスト: GLM 5.1 完全なベンチマークスイートを再実行し、このモデルの公開実行履歴スナップショットを整理しました。
  • バグ修正: 実際には再テストしていない無関係なモデルに新しい tested_at タイムスタンプが付かないようにしました。

2026-04-20

  • 新しくテストしたモデル: Kimi K2.6

2026-04-11

  • 新しくテストしたモデル: GLM 5.1

2026-03-21

2026-03-20

2026-03-18

  • 新しくテストしたモデル: MiniMax M2.7
  • 新機能: Added input and output pricing metrics to model and comparison pages.

2026-03-15

2026-03-06

  • 新機能: Added the Methodology section and expanded charts with latency, cost, token, and model-switching views.

2026-03-04

  • 新しくテストしたモデル: GPT-5.2 Chat, GPT 5.3 Chat
  • 新機能: Added interactive comparison charts for direct model analysis.
  • 追加された新テスト: Added a VAT compliance micro-audit benchmark with structured tool requirements.

2026-03-02

  • 新機能: Added share actions with copyable links across model and comparison pages.

2026-02-27

  • 新しくテストしたモデル: Seed-2.0-Mini
  • 新機能: Added reasoning-quality and consistency charts to model comparisons.

変更ログページを作成しました

この変更ログは公開後に開始したため、古い更新の一部はここに載っていません。