Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Compared models

MiniMax M2.7 (medium) vs Kimi K2.5 (medium) vs GLM 5 (medium) vs Gemini 3.1 Flash Lite Preview (medium) benchmark comparison: GLM 5 (medium) leads on Score with 7.7. MiniMax M2.7 (medium) leads on Reliability with 10.0. Gemini 3.1 Flash Lite Preview (medium) has the lowest Total Cost at $0.115. Gemini 3.1 Flash Lite Preview (medium) is fastest at 4.61s.

Last updated at: 2026-08-07

Rank
#205
Total Output Tokens
137,594
Response Time (avg)
41.28s
Total Cost
$0.176
Rank
#97
Total Output Tokens
227,367
Response Time (avg)
99.00s
Total Cost
$0.600
Rank
#59
Total Output Tokens
124,566
Response Time (avg)
33.54s
Total Cost
$0.307
Rank
#82
Total Output Tokens
56,983
Response Time (avg)
4.61s
Total Cost
$0.115
Recommended model Gemini 3.1 Flash Lite Preview (medium)

Its score stays close to the best score here (7.3 vs 7.7), while costing about 3.1x less than the other models in this comparison.

Detailed comparison

Metric MiniMax M2.7 MiniMax M2.7 medium Release: 2026-03-18 Kimi K2.5 Kimi K2.5 medium Release: 2026-01-27 GLM 5 GLM 5 medium Release: 2026-02-12 Gemini 3.1 Flash Lite Preview Gemini 3.1 Flash Lite Preview medium Release: 2026-03-03
Score 5.0 7.0 7.7 7.3
Rank #205 #97 #59 #82
Reliability 10.0 10.0 10.0 10.0
Consistency 6.6 7.0 8.1 9.9
Benchmark coverage 22/22 tests · 66/66 attempts 22/22 tests · 66/66 attempts 21/22 tests · 63/66 attempts 22/22 tests · 66/66 attempts
Tests Correct
Attempt pass rate 45.5% 65.2% 78.8% 59.1%
Flaky tests 9 8 4 0
Total Runs 66 66 63 66
Cost per result 3.906 4.789 1.668 0.884
Total Cost $0.176 $0.600 $0.307 $0.115
Input Price $0.270 / 1M $0.571 / 1M $0.950 / 1M $0.250 / 1M
Output Price $1.080 / 1M $2.850 / 1M $2.551 / 1M $1.500 / 1M
Total Input Tokens 114,518 118,448 35,224 117,480
Output Tokens 18,558 62,124 21,570 10,589
Reasoning Tokens 119,036 165,243 102,996 46,394
Response Time (avg) 41.28s 99.00s 33.54s 4.61s
Response Time (max) 196.21s 281.00s 99.85s 18.34s
Response Time (total) 866.81s 1485.04s 435.99s 101.39s

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#205 MiniMax M2.7

medium
Cost
$0.022
Time
22.8s
Tokens
9,250 tok

#97 MoonshotAI: Kimi K2.5

medium
Cost
$0.030
Time
58.6s
Tokens
8,683 tok

#59 GLM 5

medium
Cost
$0.005
Time
20.7s
Tokens
2,068 tok

#82 Gemini 3.1 Flash Lite Preview

medium
Cost
$0.003
Time
5.2s
Tokens
1,944 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
MiniMax M2.7 5.7 9.1 33.3% 0 101.89s 2,961 1,231 38,841
Kimi K2.5 6.1 4.6 66.7% 2 217.49s 6,935 5,705 74,693
GLM 5 10.0 10.0 100.0% 0 74.30s 7,254 2,997 52,930
Gemini 3.1 Flash Lite Preview 5.5 10.0 33.3% 0 4.09s 8,126 461 8,597

Quick Compare

Switch Comparison Pair