Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

Qwen: Qwen3.5-35B-A3B vs Xiaomi: MiMo-V2.5

Last updated at: 2026-05-28

Metric Qwen3.5-35B-A3B Qwen3.5-35B-A3B medium Release: 2026-02-24 MiMo-V2.5 MiMo-V2.5 medium Release: 2026-04-22
Score 7.3 7.4
Rank #65 #57
Reliability 10.0 10.0
Consistency 7.5 8.4
Tests Correct
Attempt pass rate 73.3% 70.0%
Flaky tests 6 4
Total Runs 60 60
Cost per result 4.865 2.876
Total Cost $0.368 $0.052
Input Price $0.139 / 1M $0.140 / 1M
Output Price $1.000 / 1M $0.280 / 1M
Output Tokens 31,242 2,806
Reasoning Tokens 330,546 161,888
Response Time (avg) 69.66s 20.35s
Response Time (max) 409.98s 97.49s
Response Time (total) 1393.17s 406.94s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-35B-A3B 10.0 10.0 100.0% 0 21.13s 798 42,652
MiMo-V2.5 10.0 10.0 100.0% 0 4.14s 281 1,739
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-35B-A3B 6.5 10.0 50.0% 0 244.54s 14,456 88,431
MiMo-V2.5 6.9 6.2 66.7% 1 64.48s 536 44,967
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-35B-A3B 4.7 1.6 66.7% 1 75.34s 775 12,485
MiMo-V2.5 10.0 10.0 100.0% 0 16.86s 363 7,609
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-35B-A3B 7.3 5.9 83.3% 1 59.33s 235 19,493
MiMo-V2.5 2.7 5.7 16.7% 1 6.33s 306 5,714
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-35B-A3B 4.1 4.4 44.5% 2 88.34s 41 46,368
MiMo-V2.5 5.3 10.0 33.3% 0 34.53s 507 49,478
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-35B-A3B 2.8 1.6 33.3% 1 30.30s 20 3,753
MiMo-V2.5 5.4 2.5 66.7% 1 5.37s 121 418
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-35B-A3B 10.0 10.0 100.0% 0 24.45s 97 17,361
MiMo-V2.5 9.9 10.0 100.0% 0 1.80s 88 801
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-35B-A3B 8.2 7.2 88.9% 1 33.13s 3,592 26,585
MiMo-V2.5 8.2 7.2 88.9% 1 20.25s 279 33,254
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-35B-A3B 10.0 10.0 100.0% 0 4.65s 309 1,365
MiMo-V2.5 10.0 10.0 100.0% 0 7.29s 303 2,424
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-35B-A3B 3.0 10.0 0.0% 0 177.35s 10,919 72,053
MiMo-V2.5 3.0 10.0 0.0% 0 51.29s 22 15,484

Quick Compare

Switch Comparison Pair