Hy4 preview (low) leads on average score with 5.6 vs 5.4. MiMo-V2.6-Flash has the lower benchmark cost at $0.037 vs $1.745. MiMo-V2.6-Flash is faster at 3.64s vs 214.25s, with pass rates of 19.7% vs 31.8%.
Last updated at:
2026-09-21
Compared models
Rank
#253
Total Output Tokens
682,538
Response Time (avg)
214.25s
Total Cost
$1.745
Rank
#258
Total Output Tokens
74,105
Response Time (avg)
3.64s
Total Cost
$0.037
Recommended modelMiMo-V2.6-Flash
Its score stays close to the best score here (5.4 vs 5.6), while costing about 47.3x less than Hy4 preview (low).
A test is fully passed only if every run passed for that test.Wrong answer: 14Did not follow instructions: 1Invalid tool call: 1Response Time (avg)3.64sResponse Time (max)55.14sResponse Time (total)80.06sA test is fully passed only if every run passed for that test.…
Attempt pass rate
19.7%Attempt pass rate = passed attempts / total attempts across runs.…
31.8%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
3Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
2Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
Total Runs
66Total Runs…
66Total Runs…
Cost per result
58.157
0.615
Total Cost
$1.745
$0.037
Input Price
$0.834 / 1MInput Price…
$0.140 / 1MInput Price…
Output Price
$2.501 / 1MOutput Price…
$0.280 / 1MOutput Price…
Total Input Tokens
45,154Total Input Tokens…
115,134Total Input Tokens…
Output Tokens
4,022Output Tokens…
74,105Output Tokens…
Reasoning Tokens
678,516Reasoning Tokens…
0Reasoning Tokens…
Response Time (avg)
214.25sResponse Time (avg)…
3.64sResponse Time (avg)…
Response Time (max)
597.45sResponse Time (max)…
55.14sResponse Time (max)…
Response Time (total)
4713.56sResponse Time (total)…
80.06sResponse Time (total)…
Parameters
770B total (49B active)
309B total (15B active)
Availability
Open source
Open source
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
6.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
8.9Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 4Response Time (avg)129.60sResponse Time (max)237.97sResponse Time (total)518.41sA test is fully passed only if every run passed for that test.…
4.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
25.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)816msResponse Time (max)1.57sResponse Time (total)3.26sA test is fully passed only if every run passed for that test.…
5.2Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
8.6Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 3Response Time (avg)219.15sResponse Time (max)259.22sResponse Time (total)657.45sA test is fully passed only if every run passed for that test.…
3.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.6Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
11.1%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Did not follow instructions: 1Response Time (avg)744msResponse Time (max)1.03sResponse Time (total)2.23sA test is fully passed only if every run passed for that test.…
2.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.8Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
16.7%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 2Response Time (avg)139.43sResponse Time (max)171.85sResponse Time (total)278.85sA test is fully passed only if every run passed for that test.…
3.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
9.1Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Invalid tool call: 1Wrong answer: 1Response Time (avg)4.78sResponse Time (max)7.06sResponse Time (total)9.56sA test is fully passed only if every run passed for that test.…
8.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)88.53sResponse Time (max)120.56sResponse Time (total)177.06sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)715msResponse Time (max)756msResponse Time (total)1.43sA test is fully passed only if every run passed for that test.…
4.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2API error: 1Response Time (avg)470.24sResponse Time (max)547.52sResponse Time (total)1410.71sA test is fully passed only if every run passed for that test.…
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)722msResponse Time (max)878msResponse Time (total)2.17sA test is fully passed only if every run passed for that test.…
7.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)110.43sResponse Time (max)110.43sResponse Time (total)110.43sA test is fully passed only if every run passed for that test.…
5.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)779msResponse Time (max)779msResponse Time (total)779msA test is fully passed only if every run passed for that test.…
9.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)77.72sResponse Time (max)91.04sResponse Time (total)155.43sA test is fully passed only if every run passed for that test.…
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)810msResponse Time (max)964msResponse Time (total)1.62sA test is fully passed only if every run passed for that test.…
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
2Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Did not follow instructions: 1Wrong answer: 1Response Time (avg)238.03sResponse Time (max)373.63sResponse Time (total)714.10sA test is fully passed only if every run passed for that test.…
3.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
22.2%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)739msResponse Time (max)822msResponse Time (total)2.22sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)93.67sResponse Time (max)93.67sResponse Time (total)93.67sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)1.65sResponse Time (max)1.65sResponse Time (total)1.65sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Response Time (avg)597.45sResponse Time (max)597.45sResponse Time (total)597.45sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)55.14sResponse Time (max)55.14sResponse Time (total)55.14sA test is fully passed only if every run passed for that test.…