Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

OpenAI: GPT-5.3 Chat vs Xiaomi: MiMo-V2-Flash

Last updated at: 2026-06-04

Metric GPT-5.3 Chat GPT-5.3 Chat none Release: 2026-03-03 MiMo-V2-Flash MiMo-V2-Flash medium Release: 2025-12-16
Score 7.2 7.2
Rank #63 #64
Reliability 10.0 10.0
Consistency 8.1 8.8
Tests Correct
Attempt pass rate 66.7% 65.1%
Flaky tests 5 3
Total Runs 63 63
Cost per result 3.605 0.343
Total Cost $0.433 $0.043
Input Price $1.750 / 1M $0.100 / 1M
Output Price $14.000 / 1M $0.300 / 1M
Total Input Tokens 34,209 40,111
Output Tokens 26,617 12,476
Reasoning Tokens 0 125,039
Response Time (avg) 6.34s 20.11s
Response Time (max) 18.33s 96.01s
Response Time (total) 133.13s 301.59s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.3 Chat 6.7 8.1 58.3% 1 3.86s 606 3,167 0
MiMo-V2-Flash 8.1 7.9 83.3% 1 15.85s 621 1,674 23,559
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.3 Chat 5.6 4.7 55.6% 2 10.52s 7,302 6,632 0
MiMo-V2-Flash 6.0 7.2 55.6% 1 10.71s 7,177 474 13,505
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.3 Chat 10.0 10.0 100.0% 0 11.96s 11,019 2,614 0
MiMo-V2-Flash 9.8 10.0 100.0% 0 75.68s 18,676 442 26,859
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.3 Chat 10.0 10.0 100.0% 0 2.21s 7,140 942 0
MiMo-V2-Flash 6.5 10.0 50.0% 0 0ms 2,622 153 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.3 Chat 3.5 4.4 33.3% 2 13.01s 723 8,264 0
MiMo-V2-Flash 5.9 7.2 55.6% 1 96.01s 739 8,374 42,461
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.3 Chat 4.6 10.0 0.0% 0 1.99s 477 319 0
MiMo-V2-Flash 4.0 10.0 0.0% 0 4.20s 492 87 488
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.3 Chat 9.8 10.0 100.0% 0 3.51s 660 1,491 0
MiMo-V2-Flash 10.0 10.0 100.0% 0 4.28s 678 75 3,504
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.3 Chat 10.0 10.0 100.0% 0 2.99s 642 1,758 0
MiMo-V2-Flash 7.7 10.0 66.7% 0 3.87s 670 864 1,948
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.3 Chat 10.0 10.0 100.0% 0 8.36s 5,445 861 0
MiMo-V2-Flash 10.0 10.0 100.0% 0 27.78s 8,220 321 12,715
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.3 Chat 3.0 10.0 0.0% 0 4.38s 195 569 0
MiMo-V2-Flash 3.0 10.0 0.0% 0 1.96s 216 12 0

Quick Compare

Switch Comparison Pair