Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

OpenAI: gpt-oss-120b vs Xiaomi: MiMo-V2.5-Pro

Last updated at: 2026-05-22

Metric gpt-oss-120b gpt-oss-120b none Release: 2025-08-05 Free Available MiMo-V2.5-Pro MiMo-V2.5-Pro none Release: 2026-04-22
Score 5.2 5.6
Rank #129 #115
Reliability 10.0 10.0
Consistency 8.7 8.5
Tests Correct
Attempt pass rate 36.8% 41.7%
Flaky tests 3 4
Total Runs 57 60
Cost per result 0.201 0.637
Total Cost $0.011 $0.039
Input Price $0.000 / 1M $1.000 / 1M
Output Price $0.000 / 1M $3.000 / 1M
Output Tokens 51,505 3,067
Reasoning Tokens 0 0
Response Time (avg) 21.86s 1.84s
Response Time (max) 113.71s 8.32s
Response Time (total) 349.78s 36.84s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 6.5 10.0 50.0% 0 32.84s 8,676 0
MiMo-V2.5-Pro 3.3 8.1 8.3% 1 2.67s 994 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 4.3 1.1 66.7% 1 9.57s 3,232 0
MiMo-V2.5-Pro 5.0 6.7 33.3% 1 1.80s 479 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 3.0 10.0 0.0% 0 0ms 0 0
MiMo-V2.5-Pro 3.0 10.0 0.0% 0 3.54s 596 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 6.5 10.0 50.0% 0 7.12s 598 0
MiMo-V2.5-Pro 10.0 10.0 100.0% 0 1.32s 249 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 3.0 10.0 0.0% 0 34.98s 29,483 0
MiMo-V2.5-Pro 5.3 10.0 33.3% 0 877ms 27 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 4.8 10.0 0.0% 0 10.79s 615 0
MiMo-V2.5-Pro 4.0 10.0 0.0% 0 2.58s 87 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 9.8 10.0 100.0% 0 5.10s 1,982 0
MiMo-V2.5-Pro 6.4 10.0 50.0% 0 1.03s 66 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 4.4 4.5 44.5% 2 9.51s 3,781 0
MiMo-V2.5-Pro 6.7 4.7 77.8% 2 1.32s 297 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 3.0 10.0 0.0% 0 0ms 0 0
MiMo-V2.5-Pro 10.0 10.0 100.0% 0 3.30s 258 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 3.0 10.0 0.0% 0 47.29s 3,138 0
MiMo-V2.5-Pro 3.0 10.0 0.0% 0 1.89s 14 0

Quick Compare

Switch Comparison Pair