Navigate
AI BENCHY
Advertise here

Mercury 2.5 Preview (low) vs MiMo-V2.5

MiMo-V2.5 leads on average score with 5.1 vs 5.0. Mercury 2.5 Preview (low) has the lower benchmark cost at $0.011 vs $0.025. Mercury 2.5 Preview (low) is faster at 1.31s vs 4.68s, with pass rates of 47.0% vs 27.3%.

Last updated at: 2026-09-02

Rank
#254
Total Output Tokens
31,975
Response Time (avg)
1.31s
Total Cost
$0.011
Rank
#249
Total Output Tokens
16,464
Response Time (avg)
4.68s
Total Cost
$0.025
Recommended model Mercury 2.5 Preview (low)

Its score stays close to the best score here (5.0 vs 5.1), while costing about 2.4x less than MiMo-V2.5.

Detailed comparison

Metric Mercury 2.5 Preview Mercury 2.5 Preview low Release: 2026-09-02 MiMo-V2.5 MiMo-V2.5 none Release: 2026-04-22
Score 5.0 5.1
Rank #254 #249
Reliability 10.0 10.0
Consistency 7.0 9.2
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 47.0% 27.3%
Flaky tests 8 2
Total Runs 66 66
Cost per result 0.169 0.769
Total Cost $0.011 $0.025
Input Price $0.040 / 1M $0.140 / 1M
Output Price $0.150 / 1M $0.280 / 1M
Total Input Tokens 133,525 141,052
Output Tokens 7,259 16,464
Reasoning Tokens 24,716 0
Response Time (avg) 1.31s 4.68s
Response Time (max) 7.29s 55.36s
Response Time (total) 28.77s 103.02s
Parameters ~100B 310B total (15B active)
Availability Closed Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#254 Mercury 2.5 Preview

low
Cost
$0.001
Time
1.5s
Tokens
1,169 tok

#249 MiMo-V2.5

none
Cost
$0.007
Time
267.4s
Tokens
25,283 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2.5 Preview 5.5 10.0 33.3% 0 972ms 7,794 695 2,263
MiMo-V2.5 5.5 10.0 33.3% 0 3.24s 7,440 696 0

Quick Compare

Switch Comparison Pair