不正解 の失敗
どのAIモデルで 不正解 が起きやすいかを確認し、選ぶ前に信頼性のリスクを見極められます。
316/316
モデルを絞り込む
現在の検索条件とフィルターに一致するモデルはありません。
| 順位 | モデル | 企業 | 不正解 件数 | スコア | 合計コスト | 正解テスト | 応答時間(平均) |
|---|---|---|---|---|---|---|---|
| #135 | GPT-5.6 Sol none | OpenAI | 9 | 7.0 | $0.201 | 12/22 | 2.17s |
| #143 | KAT-Coder-Pro V2.5 medium | Kwaipilot | 9 | 6.9 | $0.478 | 11/22 | 24.9s |
| #153 | DeepSeek V4 Pro none | DeepSeek | 9 | 6.8 | $0.210 | 9/22 | 11.7s |
| #192 | gpt-oss-120b medium | OpenAI | 9 | 6.1 | $0.020 | 9/22 | 20.8s |
| #211 | Dots 3 Note Preview medium | Dots Studio | 9 | 5.9 | $0.000 | 8/22 | 67.1s |
| #219 | North Mini Code medium | Cohere | 9 | 5.7 | $0.000 | 8/22 | 142.0s |
| #222 | Inkling Small low | Thinkingmachines | 9 | 5.7 | $0.055 | 9/22 | 2.07s |
| #228 | KAT-Coder-Air V2.5 high | Kwaipilot | 9 | 5.6 | $0.077 | 7/22 | 15.9s |
| #231 | Gemma 4 26B A4B none | 9 | 5.6 | $0.015 | 9/22 | 7.58s | |
| #257 | Granite 4.2 8B medium | IBM Granite | 9 | 5.2 | $0.009 | 7/22 | 46.0s |
| #274 | Ling-2.6-flash none | Inclusionai | 9 | 4.9 | $0.002 | 6/22 | 10.7s |
| #276 | Hy4 preview none | Tencent | 9 | 4.8 | $0.041 | 3/22 | 39.2s |
| #288 | Cobuddy medium | Baidu | 9 | 4.7 | $0.000 | 7/21 | 39.9s |
| #296 | Elephant Alpha none | Openrouter | 9 | 4.3 | $0.000 | 5/21 | 1.22s |
| #297 | GLM 4.7 Flash medium | Z.ai | 9 | 4.3 | $0.168 | 4/22 | 134.3s |