Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Qwen: Qwen3.6 Max Preview vs Xiaomi: MiMo-V2.5-Pro

Last updated at: 2026-05-08

Metric Qwen3.6 Max Preview Qwen3.6 Max Preview none Release: 2026-04-20 MiMo-V2.5-Pro MiMo-V2.5-Pro medium Release: 2026-04-22
Score 7.2 8.1
Rank #54 #18
Reliability 10.0 10.0
Consistency 9.1 9.2
Tests Correct
Attempt pass rate 64.9% 74.1%
Flaky tests 2 2
Total Runs 57 54
Cost per result 0.755 1.661
Total Cost $0.083 $0.200
Input Price $1.040 / 1M $1.000 / 1M
Output Price $6.240 / 1M $3.000 / 1M
Output Tokens 4,751 2,790
Reasoning Tokens 0 52,001
Response Time (avg) 3.31s 16.23s
Response Time (max) 20.51s 84.22s
Response Time (total) 62.80s 292.10s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.6 Max Preview 5.2 7.9 41.7% 1 2.63s 513 0
MiMo-V2.5-Pro 10.0 10.0 100.0% 0 3.26s 323 1,179
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.6 Max Preview 5.0 2.0 66.7% 1 3.45s 426 0
MiMo-V2.5-Pro 10.0 10.0 100.0% 0 32.58s 543 7,485
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.6 Max Preview 3.0 10.0 0.0% 0 20.51s 2,842 0
MiMo-V2.5-Pro 10.0 10.0 100.0% 0 53.36s 348 11,870
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.6 Max Preview 10.0 10.0 100.0% 0 2.87s 243 0
MiMo-V2.5-Pro 7.3 5.8 83.3% 1 18.81s 260 8,383
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.6 Max Preview 7.7 10.0 66.7% 0 1.22s 18 0
MiMo-V2.5-Pro 5.3 10.0 33.3% 0 37.87s 275 17,023
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.6 Max Preview 4.3 10.0 0.0% 0 1.62s 76 0
MiMo-V2.5-Pro 5.5 10.0 0.0% 0 4.02s 155 163
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.6 Max Preview 9.8 10.0 100.0% 0 1.45s 69 0
MiMo-V2.5-Pro 9.9 10.0 100.0% 0 2.77s 82 803
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.6 Max Preview 10.0 10.0 100.0% 0 2.38s 323 0
MiMo-V2.5-Pro 6.7 7.9 55.6% 1 5.16s 493 2,187
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.6 Max Preview 10.0 10.0 100.0% 0 5.27s 222 0
MiMo-V2.5-Pro 10.0 10.0 100.0% 0 16.87s 311 2,908
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.6 Max Preview 3.0 10.0 0.0% 0 1.97s 19 0
MiMo-V2.5-Pro - - - - - - - -

Quick Compare

Switch Comparison Pair