Navigate
AI BENCHY
Advertise here

MiniMax M2.7 (medium) vs Mistral Small 4

Mistral Small 4 leads on average score with 5.1 vs 5.0. Mistral Small 4 has the lower benchmark cost at $0.022 vs $0.163. Mistral Small 4 is faster at 1.20s vs 41.28s, with pass rates of 45.5% vs 25.8%.

Last updated at: 2026-08-02

Rank
#200
Total Output Tokens
137,594
Response Time (avg)
41.28s
Total Cost
$0.163
Rank
#193
Total Output Tokens
9,812
Response Time (avg)
1.20s
Total Cost
$0.022
Recommended model Mistral Small 4

It has the best score here (5.1), while costing about 7.5x less than MiniMax M2.7 (medium).

Detailed comparison

Metric MiniMax M2.7 MiniMax M2.7 medium Release: 2026-03-18 Mistral Small 4 Mistral Small 4 none Release: 2026-03-16
Score 5.0 5.1
Rank #200 #193
Reliability 10.0 10.0
Consistency 6.6 9.6
Benchmark coverage 22/22 tests · 66/66 attempts 22/22 tests · 66/66 attempts
Tests Correct
Attempt pass rate 45.5% 25.8%
Flaky tests 9 1
Total Runs 66 66
Cost per result 3.906 0.432
Total Cost $0.163 $0.022
Input Price $0.250 / 1M $0.150 / 1M
Output Price $1.000 / 1M $0.600 / 1M
Total Input Tokens 114,518 104,708
Output Tokens 18,558 9,812
Reasoning Tokens 119,036 0
Response Time (avg) 41.28s 1.20s
Response Time (max) 196.21s 13.16s
Response Time (total) 866.81s 26.38s

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#200 MiniMax M2.7

medium
Cost
$0.022
Time
22.8s
Tokens
9,250 tok

#193 Mistral Small 4

none
Cost
$0.002
Time
10.4s
Tokens
2,370 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
MiniMax M2.7 5.7 9.1 33.3% 0 101.89s 2,961 1,231 38,841
Mistral Small 4 3.7 9.7 0.0% 0 901ms 7,636 619 0

Quick Compare

Switch Comparison Pair