AI BENCHY Compare

Qwen: Qwen3.5 Plus 2026-02-15 vs Xiaomi: MiMo-V2-Flash

ベンチマークは AI BENCHY テストスイートから次の日時に生成: 2026-05-29

指標	Qwen3.5 Plus 2026-02-15 Qwen3.5 Plus 2026-02-15 none リリース: 2026-02-15	MiMo-V2-Flash MiMo-V2-Flash medium リリース: 2025-12-16

指標	Qwen3.5 Plus 2026-02-15 Qwen3.5 Plus 2026-02-15 none リリース: 2026-02-15	MiMo-V2-Flash MiMo-V2-Flash medium リリース: 2025-12-16
スコア	6.4	7.1
順位	#94	#77
信頼性	10.0	10.0
一貫性	9.3	8.7
正解テスト
試行ごとの合格率	48.3%	63.3%
不安定なテスト	2	3
総実行回数	60	60
結果あたりのコスト	0.195	0.345
合計コスト	$0.018	$0.038
入力価格	$0.260 / 1M	$0.100 / 1M
出力価格	$1.560 / 1M	$0.300 / 1M
出力トークン	2,474	12,458
推論トークン	0	115,182
応答時間（平均）	2.40s	20.28s
応答時間（最大）	6.65s	96.01s
応答時間（合計）	33.56s	283.87s

スコア上位モデル

スコア vs 総コスト

応答時間（平均）

スコア vs 応答時間（平均）

合計出力トークン

スコア vs 合計出力トークン

カテゴリ内訳

反AIトリック	スコア	一貫性	試行ごとの合格率	不安定なテスト	正解テスト	応答時間（平均）	出力トークン	推論トークン
Qwen3.5 Plus 2026-02-15	4.8	10.0	25.0%	0		1.91s	517	0
MiMo-V2-Flash	8.1	7.9	83.3%	1		15.85s	1,674	23,559

コーディング	スコア	一貫性	試行ごとの合格率	不安定なテスト	正解テスト	応答時間（平均）	出力トークン	推論トークン
Qwen3.5 Plus 2026-02-15	4.9	6.9	16.7%	1		2.54s	467	0
MiMo-V2-Flash	4.1	5.8	33.3%	1		7.20s	456	3,648

複合	スコア	一貫性	試行ごとの合格率	不安定なテスト	正解テスト	応答時間（平均）	出力トークン	推論トークン
Qwen3.5 Plus 2026-02-15	3.0	10.0	0.0%	0		6.65s	314	0
MiMo-V2-Flash	9.8	10.0	100.0%	0		75.68s	442	26,859

データ解析と抽出	スコア	一貫性	試行ごとの合格率	不安定なテスト	正解テスト	応答時間（平均）	出力トークン	推論トークン
Qwen3.5 Plus 2026-02-15	10.0	10.0	100.0%	0		1.89s	243	0
MiMo-V2-Flash	6.5	10.0	50.0%	0		0ms	153	0

ドメイン特化	スコア	一貫性	試行ごとの合格率	不安定なテスト	正解テスト	応答時間（平均）	出力トークン	推論トークン
Qwen3.5 Plus 2026-02-15	5.3	10.0	33.3%	0		1.17s	17	0
MiMo-V2-Flash	5.9	7.2	55.6%	1		96.01s	8,374	42,461

汎用知能	スコア	一貫性	試行ごとの合格率	不安定なテスト	正解テスト	応答時間（平均）	出力トークン	推論トークン
Qwen3.5 Plus 2026-02-15	4.4	3.0	33.3%	1		2.26s	117	0
MiMo-V2-Flash	4.0	10.0	0.0%	0		4.20s	87	488

指示追従	スコア	一貫性	試行ごとの合格率	不安定なテスト	正解テスト	応答時間（平均）	出力トークン	推論トークン
Qwen3.5 Plus 2026-02-15	10.0	10.0	100.0%	0		1.67s	72	0
MiMo-V2-Flash	10.0	10.0	100.0%	0		4.28s	75	3,504

パズル解決	スコア	一貫性	試行ごとの合格率	不安定なテスト	正解テスト	応答時間（平均）	出力トークン	推論トークン
Qwen3.5 Plus 2026-02-15	7.7	10.0	66.7%	0		2.71s	494	0
MiMo-V2-Flash	7.7	10.0	66.7%	0		3.87s	864	1,948

ツール呼び出し	スコア	一貫性	試行ごとの合格率	不安定なテスト	正解テスト	応答時間（平均）	出力トークン	推論トークン
Qwen3.5 Plus 2026-02-15	10.0	10.0	100.0%	0		3.33s	222	0
MiMo-V2-Flash	10.0	10.0	100.0%	0		27.78s	321	12,715

雑学	スコア	一貫性	試行ごとの合格率	不安定なテスト	正解テスト	応答時間（平均）	出力トークン	推論トークン
Qwen3.5 Plus 2026-02-15	3.0	10.0	0.0%	0		1.11s	11	0
MiMo-V2-Flash	3.0	10.0	0.0%	0		1.96s	12	0

クイック比較

比較ペアを切り替え

Claude Sonnet 4.6nonevsMiMo-V2-Flashmedium Qwen3.6 Max PreviewnonevsMiMo-V2-Flashmedium DeepSeek V4 ProhighvsMiMo-V2-Flashmedium Step 3.7 FlashhighvsMiMo-V2-Flashmedium Mercury 2mediumvsQwen3.5 Plus 2026-02-15none Ring-2.6-1TnonevsMiMo-V2-Flashmedium Claude Opus 4.8nonevsMiMo-V2-Flashmedium GPT-5 NanomediumvsQwen3.5 Plus 2026-02-15none Kimi K2.5mediumvsQwen3.5 Plus 2026-02-15none Gemini 3.1 Flash LiteminimalvsQwen3.5 Plus 2026-02-15none Step 3.7 FlashlowvsMiMo-V2-Flashmedium GPT-5.3 ChatnonevsMiMo-V2-Flashmedium