Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

MiniMax: MiniMax M2.5 vs xAI: Grok 4.20

Summary

MiniMax M2.5 vs Grok 4.20 benchmark comparison: MiniMax M2.5 leads on average score with 4.7 vs 4.4. Grok 4.20 has the lower benchmark cost at $0.057 vs $0.164. Grok 4.20 is faster at 1.11s vs 65.37s, with pass rates of 46.0% vs 28.6%.

Recommended model: Grok 4.20 - Its score stays close to the best score here (4.4 vs 4.7), while costing about 2.9x less than MiniMax M2.5.

Last updated at: 2026-07-02

Metric MiniMax M2.5 MiniMax M2.5 medium Release: 2026-02-12 Grok 4.20 Grok 4.20 none Release: 2026-03-31
Score 4.7 4.4
Rank #151 #160
Reliability 10.0 N/A
Consistency 6.5 8.5
Tests Correct
Attempt pass rate 46.0% 28.6%
Flaky tests 9 0
Total Runs 63 54
Cost per result 7.900 1.570
Total Cost $0.164 $0.057
Input Price $0.120 / 1M $1.250 / 1M
Output Price $0.480 / 1M $2.500 / 1M
Total Input Tokens 43,706 41,313
Output Tokens 109,495 1,923
Reasoning Tokens 330,814 0
Response Time (avg) 65.37s 1.11s
Response Time (max) 251.36s 6.04s
Response Time (total) 849.76s 19.96s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#151 MiniMax M2.5

medium
Invalid SVG
Cost
$0.000
Time
300.0s
Tokens
0 tok

#160 xAI: Grok 4.20

none
Cost
$0.004
Time
6.5s
Tokens
1,367 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
MiniMax M2.5 7.9 6.3 83.3% 2 20.82s 612 286 45,344
Grok 4.20 4.8 10.0 25.0% 0 501ms 1,986 267 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
MiniMax M2.5 3.4 9.1 0.0% 0 188.58s 6,076 357 106,177
Grok 4.20 1.1 3.1 0.0% 0 1.22s 1,074 312 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
MiniMax M2.5 4.5 2.1 66.7% 1 60.39s 21,104 740 9,713
Grok 4.20 3.0 10.0 0.0% 0 6.04s 17,673 282 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
MiniMax M2.5 4.6 1.7 66.7% 2 7.48s 6,584 266 3,835
Grok 4.20 10.0 10.0 100.0% 0 522ms 7,749 207 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
MiniMax M2.5 2.9 4.4 22.2% 2 237.27s 308 105,047 133,487
Grok 4.20 3.0 10.0 0.0% 0 687ms 1,746 325 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
MiniMax M2.5 3.8 2.5 33.3% 1 6.63s 492 25 1,686
Grok 4.20 4.8 10.0 0.0% 0 659ms 819 83 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
MiniMax M2.5 7.5 10.0 50.0% 0 621ms 699 156 1,495
Grok 4.20 6.3 10.0 50.0% 0 445ms 1,350 60 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
MiniMax M2.5 5.3 7.2 44.4% 1 11.21s 495 1,069 9,605
Grok 4.20 5.3 10.0 33.3% 0 473ms 1,671 198 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
MiniMax M2.5 10.0 10.0 100.0% 0 15.35s 7,123 269 937
Grok 4.20 10.0 10.0 100.0% 0 4.63s 7,245 189 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
MiniMax M2.5 3.0 10.0 0.0% 0 80.79s 213 1,280 18,535
Grok 4.20 0.0 0.0 0.0% 0 0ms 0 0 0

Quick Compare

Switch Comparison Pair