AI BENCHY Compare

Inception: Mercury 2 vs MiniMax: MiniMax M3

Summary

Mercury 2 vs MiniMax M3 benchmark comparison: MiniMax M3 leads on average score with 7.6 vs 7.5. Mercury 2 has the lower benchmark cost at $0.058 vs $0.131. Mercury 2 is faster at 2.24s vs 68.17s, with pass rates of 54.0% vs 65.1%.

Recommended model: Mercury 2 - Its score stays close to the best score here (7.5 vs 7.6), while costing about 2.3x less than MiniMax M3.

Last updated at: 2026-06-18

Metric	Mercury 2 Mercury 2 medium Release: 2026-02-24	MiniMax M3 MiniMax M3 medium Release: 2026-06-01

Metric	Mercury 2 Mercury 2 medium Release: 2026-02-24	MiniMax M3 MiniMax M3 medium Release: 2026-06-01
Score	7.5	7.6
Rank	#44	#40
Reliability	10.0	9.6
Consistency	8.8	7.9
Tests Correct
Attempt pass rate	54.0%	65.1%
Flaky tests	3	5
Total Runs	63	63
Cost per result	0.578	1.187
Total Cost	$0.058	$0.131
Input Price	$0.250 / 1M	$0.300 / 1M
Output Price	$0.750 / 1M	$1.200 / 1M
Total Input Tokens	35,116	46,546
Output Tokens	4,048	49,036
Reasoning Tokens	61,219	92,543
Response Time (avg)	2.24s	68.17s
Response Time (max)	14.63s	431.03s
Response Time (total)	44.72s	1363.38s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#44 Mercury 2

medium

Cost: $0.002
Time: 2.1s
Tokens: 1,702 tok

#40 MiniMax M3

medium

Cost: $0.012
Time: 154.4s
Tokens: 10,018 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
Mercury 2	6.9	9.9	50.0%	0		1.12s	554	2,546	2,609
MiniMax M3	5.5	3.7	66.7%	3		14.95s	2,526	874	3,414

Coding	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
Mercury 2	8.2	7.7	77.8%	1		2.04s	7,065	296	11,328
MiniMax M3	6.1	6.5	55.6%	1		144.74s	5,804	6,223	32,667

Combined	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
Mercury 2	10.0	10.0	100.0%	0		3.28s	12,909	268	4,887
MiniMax M3	10.0	10.0	100.0%	0		65.30s	14,760	1,306	6,253

Data parsing and extraction	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
Mercury 2	7.3	5.9	83.3%	1		1.11s	6,234	183	1,656
MiniMax M3	10.0	10.0	100.0%	0		14.92s	8,088	514	3,164

Domain specific	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
Mercury 2	2.9	7.2	11.1%	1		6.48s	695	41	30,754
MiniMax M3	5.5	9.3	33.3%	0		233.13s	869	16,254	19,070

General Intelligence	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
Mercury 2	4.8	10.0	0.0%	0		821ms	456	137	542
MiniMax M3	5.1	3.4	33.3%	1		33.25s	954	2,487	2,523

Instructions following	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
Mercury 2	10.0	10.0	100.0%	0		1.07s	340	14	958
MiniMax M3	9.8	10.0	100.0%	0		6.14s	1,623	103	920

Puzzle Solving	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
Mercury 2	5.4	10.0	33.3%	0		949ms	601	361	2,781
MiniMax M3	7.9	9.9	66.7%	0		49.91s	2,079	11,946	13,761

Tool Calling	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
Mercury 2	10.0	10.0	100.0%	0		1.89s	6,080	180	1,956
MiniMax M3	10.0	10.0	100.0%	0		11.91s	9,168	281	555

Trivia	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
Mercury 2	3.0	10.0	0.0%	0		2.58s	182	22	3,748
MiniMax M3	3.0	10.0	0.0%	0		100.80s	675	9,048	10,216

Quick Compare

Switch Comparison Pair

DeepSeek V4 ProhighvsMiniMax M3medium Mercury 2mediumvsGPT-5.3 Chatnone DeepSeek V4 ProhighvsMercury 2medium MiniMax M3mediumvsStep 3.7 Flashlow MiniMax M3mediumvsGPT-5.3 Chatnone Mercury 2mediumvsStep 3.7 Flashlow Gemini 3 Flash PreviewlowvsMercury 2medium Claude Sonnet 4.6nonevsMercury 2medium Gemini 3 Flash PreviewlowvsMiniMax M3medium Claude Sonnet 4.6nonevsMiniMax M3medium Claude Opus 4.8nonevsMercury 2medium DeepSeek V4 PrononevsMercury 2medium