Navigate
AI BENCHY
Compare Charts
❤️ Made by XCS
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

MiniMax: MiniMax M2.5 vs OpenAI: GPT-5.4

Compare:

Last updated at: 2026-03-05

Metric MiniMax: MiniMax M2.5 medium Release: 2026-02-12 OpenAI: GPT-5.4 none Release: 2026-03-05
Rank #42 #44
Avg Score 4.8 4.6
Tests Correct
Consistency 5.8 8.9
Cost per result 4.937 1.496
Total Cost $0.247 $0.090
Attempt pass rate 62.2% 44.4%
Flaky tests 8 2
common.totalAttempts 45 (15 x 3) 45 (15 x 3)
Output Tokens 107,019 1,635
Reasoning Tokens 204,504 0
Response Time (avg) 47.58s 1.46s
Response Time (max) 237.27s 2.89s
Response Time (total) 380.62s 21.86s

Top Models by Score

Response Time (avg)

Score vs Total Cost

Avg Score vs Response Time (avg)

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
MiniMax: MiniMax M2.5 9.3 7.9 88.9% 1 32.42s 286 45,112
OpenAI: GPT-5.4 10.0 7.3 11.1% 1 1.41s 388 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
MiniMax: MiniMax M2.5 10.0 2.1 66.7% 1 60.39s 740 9,713
OpenAI: GPT-5.4 10.0 10.0 0.0% 0 2.89s 291 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
MiniMax: MiniMax M2.5 10.0 1.7 66.7% 2 7.48s 266 3,835
OpenAI: GPT-5.4 9.9 10.0 100.0% 0 1.04s 222 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
MiniMax: MiniMax M2.5 10.0 4.4 22.2% 2 237.27s 105,047 133,487
OpenAI: GPT-5.4 4.0 7.2 44.4% 1 1.07s 50 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
MiniMax: MiniMax M2.5 8.0 6.8 83.3% 1 4.64s 252 1,873
OpenAI: GPT-5.4 5.5 10.0 50.0% 0 1.07s 81 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
MiniMax: MiniMax M2.5 4.0 7.2 44.4% 1 11.54s 159 9,547
OpenAI: GPT-5.4 4.0 9.8 33.3% 0 1.52s 357 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
MiniMax: MiniMax M2.5 10.0 10.0 100.0% 0 15.35s 269 937
OpenAI: GPT-5.4 10.0 10.0 100.0% 0 2.75s 246 0

Quick Compare

Switch Comparison Pair