Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Compared models

Kimi K2.6 (medium) vs Kimi K2.5 (medium) vs GLM 5 (medium) vs Claude Opus 4.7 (medium) benchmark comparison: Claude Opus 4.7 (medium) leads on Score with 8.7. Kimi K2.5 (medium) leads on Reliability with 10.0. GLM 5 (medium) has the lowest Total Cost at $0.307. Claude Opus 4.7 (medium) is fastest at 7.61s.

Last updated at: 2026-07-28

Rank
#78
Total Output Tokens
391,540
Response Time (avg)
109.98s
Total Cost
$0.831
Rank
#87
Total Output Tokens
227,367
Response Time (avg)
99.00s
Total Cost
$0.600
Rank
#50
Total Output Tokens
124,566
Response Time (avg)
33.54s
Total Cost
$0.307
Rank
#20
Total Output Tokens
29,990
Response Time (avg)
7.61s
Total Cost
$1.477
Recommended model Claude Opus 4.7 (medium)

It has the best score here (8.7), while responding about 10.6x faster than the other models in this comparison.

Detailed comparison

Metric Kimi K2.6 Kimi K2.6 medium Release: 2026-04-20 Kimi K2.5 Kimi K2.5 medium Release: 2026-01-27 GLM 5 GLM 5 medium Release: 2026-02-12 Claude Opus 4.7 Claude Opus 4.7 medium Release: 2026-04-16
Score 7.2 7.0 7.7 8.7
Rank #78 #87 #50 #20
Reliability 9.4 10.0 10.0 10.0
Consistency 8.3 7.0 8.1 9.6
Tests Correct
Attempt pass rate 63.6% 65.2% 78.8% 83.3%
Flaky tests 4 8 4 1
Total Runs 66 66 63 66
Cost per result 9.821 4.789 1.668 8.201
Total Cost $0.831 $0.600 $0.307 $1.477
Input Price $0.646 / 1M $0.571 / 1M $0.950 / 1M $5.000 / 1M
Output Price $2.720 / 1M $2.850 / 1M $2.551 / 1M $25.000 / 1M
Total Input Tokens 68,902 118,448 35,224 145,252
Output Tokens 111,680 62,124 21,570 24,948
Reasoning Tokens 279,860 165,243 102,996 5,042
Response Time (avg) 109.98s 99.00s 33.54s 7.61s
Response Time (max) 876.20s 281.00s 99.85s 65.40s
Response Time (total) 2309.56s 1485.04s 435.99s 159.91s

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#78 MoonshotAI: Kimi K2.6

medium
Cost
$0.013
Time
103.4s
Tokens
3,620 tok

#87 MoonshotAI: Kimi K2.5

medium
Cost
$0.030
Time
58.6s
Tokens
8,683 tok

#50 GLM 5

medium
Cost
$0.005
Time
20.7s
Tokens
2,068 tok

#20 Claude Opus 4.7

medium
Cost
$0.059
Time
26.8s
Tokens
2,475 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Kimi K2.6 5.7 8.6 33.3% 0 214.42s 2,925 9,970 77,189
Kimi K2.5 6.1 4.6 66.7% 2 217.49s 6,935 5,705 74,693
GLM 5 10.0 10.0 100.0% 0 74.30s 7,254 2,997 52,930
Claude Opus 4.7 7.6 7.2 77.8% 1 12.96s 10,635 7,629 1,114

Quick Compare

Switch Comparison Pair