Navigate
AI BENCHY
Advertise here

Mistral Small 4 (medium) vs Qwen3.5-9B

The average score is effectively tied at 5.1 vs 5.1. Qwen3.5-9B has the lower benchmark cost at $0.021 vs $0.096. Mistral Small 4 (medium) is faster at 10.77s vs 19.17s, with pass rates of 42.4% vs 19.7%.

Last updated at: 2026-07-28

Rank
#187
Total Output Tokens
131,824
Response Time (avg)
10.77s
Total Cost
$0.096
Rank
#189
Total Output Tokens
37,484
Response Time (avg)
19.17s
Total Cost
$0.021
Recommended model Mistral Small 4 (medium)

It has the best score here (5.1), while responding about 1.8x faster than Qwen3.5-9B.

Detailed comparison

Metric Mistral Small 4 Mistral Small 4 medium Release: 2026-03-16 Qwen3.5-9B Qwen3.5-9B none Release: 2026-03-02
Score 5.1 5.1
Rank #187 #189
Reliability 10.0 10.0
Consistency 7.0 9.7
Tests Correct
Attempt pass rate 42.4% 19.7%
Flaky tests 8 1
Total Runs 66 66
Cost per result 1.913 0.490
Total Cost $0.096 $0.021
Input Price $0.150 / 1M $0.100 / 1M
Output Price $0.600 / 1M $0.150 / 1M
Total Input Tokens 140,494 144,407
Output Tokens 39,462 37,484
Reasoning Tokens 92,362 0
Response Time (avg) 10.77s 19.17s
Response Time (max) 59.15s 382.06s
Response Time (total) 236.94s 421.74s

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#187 Mistral Small 4

medium
Cost
$0.006
Time
47.9s
Tokens
9,857 tok

#189 Qwen3.5-9B

none
Invalid SVG
Cost
$0.000
Time
300.0s
Tokens
0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mistral Small 4 4.4 5.1 33.3% 2 39.98s 7,636 11,635 54,715
Qwen3.5-9B 3.9 7.8 11.1% 1 5.60s 7,913 1,042 0

Quick Compare

Switch Comparison Pair