Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

Inception: Mercury 2 vs MiniMax: MiniMax M2.7

Last updated at: 2026-06-03

Metric Mercury 2 Mercury 2 none Release: 2026-02-24 MiniMax M2.7 MiniMax M2.7 medium Release: 2026-03-18
Score 4.6 5.4
Rank #153 #128
Reliability 10.0 10.0
Consistency 9.1 6.8
Tests Correct
Attempt pass rate 25.0% 48.3%
Flaky tests 2 8
Total Runs 60 60
Cost per result 0.216 2.076
Total Cost $0.009 $0.104
Input Price $0.250 / 1M $0.279 / 1M
Output Price $0.750 / 1M $1.200 / 1M
Total Input Tokens 25,515 33,493
Output Tokens 3,001 8,224
Reasoning Tokens 0 73,373
Response Time (avg) 614ms 29.86s
Response Time (max) 1.27s 117.04s
Response Time (total) 12.28s 567.39s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2 3.0 10.0 0.0% 0 483ms 631 286 0
MiniMax M2.7 7.9 6.3 83.3% 2 40.32s 654 3,010 17,716
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2 3.5 9.4 0.0% 0 831ms 4,631 1,650 0
MiniMax M2.7 6.7 9.7 50.0% 0 54.73s 2,083 474 22,402
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2 3.0 10.0 0.0% 0 606ms 4,821 131 0
MiniMax M2.7 4.7 1.6 66.7% 1 41.03s 14,233 369 4,480
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2 7.3 5.9 83.3% 1 667ms 6,362 180 0
MiniMax M2.7 6.3 5.8 66.7% 1 21.95s 7,152 187 5,882
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2 5.3 7.2 44.4% 1 534ms 784 46 0
MiniMax M2.7 3.0 10.0 0.0% 0 19.00s 245 8 2,796
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2 4.8 10.0 0.0% 0 628ms 495 159 0
MiniMax M2.7 3.9 2.5 33.3% 1 38.70s 486 92 5,204
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2 6.5 10.0 50.0% 0 551ms 691 82 0
MiniMax M2.7 3.8 5.8 33.3% 1 12.80s 687 350 2,600
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2 3.1 10.0 0.0% 0 535ms 694 251 0
MiniMax M2.7 5.9 7.2 55.6% 1 24.87s 675 362 7,840
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2 10.0 10.0 100.0% 0 1.27s 6,193 197 0
MiniMax M2.7 4.7 1.6 66.7% 1 12.05s 7,067 304 1,001
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2 3.0 10.0 0.0% 0 548ms 213 19 0
MiniMax M2.7 3.0 10.0 0.0% 0 22.77s 211 3,068 3,452

Quick Compare

Switch Comparison Pair