AI BENCHY Compare
MoonshotAI: Kimi K2.5 vs OpenAI: GPT-5 Mini
Vergelijken:
Benchmarks gegenereerd uit AI BENCHY-testsuites op: 2026-03-04
| Metriek | MoonshotAI: Kimi K2.5 none Releasedatum: 2026-01-27 | OpenAI: GPT-5 Mini medium Releasedatum: 2025-08-07 |
|---|---|---|
| Rang | #45 | #30 |
| Gem. score | 3.87 | 6.05 |
| Consistentie | 9.00 | 8.87 |
| Kosten per resultaat | 0.261 | 1.122 |
| Totale kosten | $0.011 | $0.090 |
| Correcte tests | ||
| Slaagpercentage per poging | 33.3% | 60.0% |
| Instabiele tests | 2 | 2 |
| Uitvoer-tokens | 1,968 | 4,974 |
| Redeneer-tokens | 0 | 37,568 |
Score vs totale kosten
Categorie-uitsplitsing
| Anti-AI-trucs | Score | Consistentie | Slaagpercentage per poging | Instabiele tests | Correcte tests | Uitvoer-tokens | Redeneer-tokens |
|---|---|---|---|---|---|---|---|
| MoonshotAI: Kimi K2.5 | 2.67 | 7.86 | 11.1% | 1 | 363 | 0 | |
| OpenAI: GPT-5 Mini | 7.00 | 9.62 | 66.7% | 0 | 1,645 | 5,824 |
| Gecombineerd | Score | Consistentie | Slaagpercentage per poging | Instabiele tests | Correcte tests | Uitvoer-tokens | Redeneer-tokens |
|---|---|---|---|---|---|---|---|
| MoonshotAI: Kimi K2.5 | 1.00 | 10.00 | 0.0% | 0 | 53 | 0 | |
| OpenAI: GPT-5 Mini | 10.00 | 10.00 | 100.0% | 0 | 251 | 2,176 |
| Gegevensparsering en extractie | Score | Consistentie | Slaagpercentage per poging | Instabiele tests | Correcte tests | Uitvoer-tokens | Redeneer-tokens |
|---|---|---|---|---|---|---|---|
| MoonshotAI: Kimi K2.5 | 5.50 | 5.81 | 83.3% | 1 | 995 | 0 | |
| OpenAI: GPT-5 Mini | 9.88 | 10.00 | 100.0% | 0 | 453 | 3,200 |
| Domeinspecifiek | Score | Consistentie | Slaagpercentage per poging | Instabiele tests | Correcte tests | Uitvoer-tokens | Redeneer-tokens |
|---|---|---|---|---|---|---|---|
| MoonshotAI: Kimi K2.5 | 4.00 | 10.00 | 33.3% | 0 | 29 | 0 | |
| OpenAI: GPT-5 Mini | 1.00 | 7.21 | 22.2% | 1 | 293 | 14,016 |
| Instructies opvolgen | Score | Consistentie | Slaagpercentage per poging | Instabiele tests | Correcte tests | Uitvoer-tokens | Redeneer-tokens |
|---|---|---|---|---|---|---|---|
| MoonshotAI: Kimi K2.5 | 5.00 | 9.99 | 50.0% | 0 | 61 | 0 | |
| OpenAI: GPT-5 Mini | 7.00 | 6.64 | 66.7% | 1 | 318 | 4,992 |
| Puzzle Solving | Score | Consistentie | Slaagpercentage per poging | Instabiele tests | Correcte tests | Uitvoer-tokens | Redeneer-tokens |
|---|---|---|---|---|---|---|---|
| MoonshotAI: Kimi K2.5 | 2.00 | 9.92 | 0.0% | 0 | 247 | 0 | |
| OpenAI: GPT-5 Mini | 4.33 | 9.78 | 33.3% | 0 | 1,527 | 5,760 |
| Toolaanroepen | Score | Consistentie | Slaagpercentage per poging | Instabiele tests | Correcte tests | Uitvoer-tokens | Redeneer-tokens |
|---|---|---|---|---|---|---|---|
| MoonshotAI: Kimi K2.5 | 10.00 | 10.00 | 100.0% | 0 | 220 | 0 | |
| OpenAI: GPT-5 Mini | 10.00 | 10.00 | 100.0% | 0 | 487 | 1,600 |
Snelle vergelijking
Vergelijkingspaar wisselen
Kimi K2.5nonevsGLM 4.7 FlashmediumGPT-5 MinimediumvsQwen3.5 Plus 2026-02-15noneGPT-5 MinimediumvsGLM 5noneClaude Sonnet 4.6nonevsGPT-5 MinimediumKimi K2.5nonevsQwen3 Coder NextmediumGemini 2.5 FlashnonevsGPT-5 MinimediumGPT-5 MinimediumvsQwen3.5-35B-A3BnoneDeepSeek V3.2nonevsGPT-5 MinimediumGPT-5 MinimediumvsQwen3.5-122B-A10BnoneGemini 3 Flash PreviewnonevsGPT-5 MinimediumGemini 3.1 Flash Lite PreviewnonevsGPT-5 MinimediumGemini 3.1 Flash Lite PreviewlowvsGPT-5 Minimedium