AI BENCHY Compare
MoonshotAI: Kimi K2.5 vs xAI: Grok 4.20
Last updated at: 2026-04-02
| Metric | Kimi K2.5 Kimi K2.5 medium | Grok 4.20 Grok 4.20 medium |
|---|---|---|
| Score | 7.2 | 7.1 |
| Rank | #39 | #40 |
| Consistency | 7.2 | 8.2 |
| Tests Correct | ||
| Attempt pass rate | 72.6% | 66.7% |
| Flaky tests | 6 | 4 |
| Total Runs | 51 | 51 |
| Cost per result | 2.232 | 7.358 |
| Total Cost | $0.201 | $0.663 |
| Input Price | $0.383 / 1M | $2.000 / 1M |
| Output Price | $1.909 / 1M | $6.000 / 1M |
| Output Tokens | 40,907 | 1,494 |
| Reasoning Tokens | 75,121 | 97,078 |
| Response Time (avg) | 64.59s | 9.50s |
| Response Time (max) | 137.29s | 29.87s |
| Response Time (total) | 645.93s | 161.54s |
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Category Breakdown
Quick Compare
Switch Comparison Pair
Claude Sonnet 4.6nonevsKimi K2.5mediumClaude Sonnet 4.6nonevsGrok 4.20mediumQwen3.5 Plus 2026-02-15nonevsGrok 4.20mediumKimi K2.5mediumvsGPT-5.3 ChatnoneGemma 4 31BnonevsGrok 4.20mediumGrok 4.20mediumvsGLM 5noneKimi K2.5mediumvsQwen3.5 Plus 2026-02-15noneGPT-5.3 ChatnonevsGrok 4.20mediumGemma 4 31BnonevsKimi K2.5mediumKimi K2.5mediumvsGLM 5noneKimi K2.5mediumvsGPT-5.2 ChatnoneGemini 3.1 Flash Lite PreviewnonevsKimi K2.5medium