Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

MiniMax: MiniMax M2.5 vs OpenAI: GPT-5.4 Nano

Last updated at: 2026-03-17

Metric MiniMax M2.5 MiniMax M2.5 medium Release: 2026-02-12 Free Available GPT-5.4 Nano GPT-5.4 Nano none Release: 2026-03-17
Rank #50 #73
Score 5.9 4.3
Consistency 5.4 7.3
Cost per result 4.987 0.404
Total Cost $0.250 $0.009
Tests Correct
Attempt pass rate 60.8% 29.4%
Flaky tests 10 6
Total Runs 51 51
Output Tokens 107,044 2,185
Reasoning Tokens 206,422 0
Response Time (avg) 39.65s 1.39s
Response Time (max) 237.27s 3.84s
Response Time (total) 396.47s 23.70s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
MiniMax M2.5 7.9 6.3 83.3% 2 20.82s 286 45,344
GPT-5.4 Nano 3.5 8.0 16.7% 1 1.18s 800 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
MiniMax M2.5 4.5 2.1 66.7% 1 60.39s 740 9,713
GPT-5.4 Nano 3.0 10.0 0.0% 0 3.84s 280 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
MiniMax M2.5 4.6 1.7 66.7% 2 7.48s 266 3,835
GPT-5.4 Nano 6.5 10.0 50.0% 0 1.11s 219 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
MiniMax M2.5 2.9 4.4 22.2% 2 237.27s 105,047 133,487
GPT-5.4 Nano 2.9 4.4 22.2% 2 926ms 52 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
MiniMax M2.5 3.8 2.5 33.3% 1 6.63s 25 1,686
GPT-5.4 Nano 3.8 2.5 33.3% 1 1.31s 180 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
MiniMax M2.5 8.1 6.8 83.3% 1 4.64s 252 1,873
GPT-5.4 Nano 5.0 6.8 33.3% 1 787ms 84 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
MiniMax M2.5 5.3 7.2 44.4% 1 11.54s 159 9,547
GPT-5.4 Nano 3.7 7.3 22.2% 1 1.29s 348 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
MiniMax M2.5 10.0 10.0 100.0% 0 15.35s 269 937
GPT-5.4 Nano 10.0 10.0 100.0% 0 3.40s 222 0

Quick Compare

Switch Comparison Pair