Navigate
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Mistral Small 4 (medium) vs Qwen3.5-9B

Mistral Small 4 (medium) leads on average score with 5.1 vs 5.1. Qwen3.5-9B has the lower benchmark cost at $0.021 vs $0.097. Mistral Small 4 (medium) is faster at 10.80s vs 19.19s, with pass rates of 43.9% vs 19.7%.

Last updated at: 2026-09-10

Compared models

Rank
#267
Total Output Tokens
133,597
Response Time (avg)
10.80s
Total Cost
$0.097
Rank
#272
Total Output Tokens
37,484
Response Time (avg)
19.19s
Total Cost
$0.021
Recommended model Mistral Small 4 (medium)

It has the best score here (5.1), while responding about 1.8x faster than Qwen3.5-9B.

Detailed comparison

Metric Mistral Small 4 Mistral Small 4 medium Release: 2026-03-16 Qwen3.5-9B Qwen3.5-9B none Release: 2026-03-02
Score 5.1 5.1
Rank #267 #272
Reliability 10.0 10.0
Consistency 7.0 9.7
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 43.9% 19.7%
Flaky tests 8 1
Total Runs 66 66
Cost per result 1.923 0.490
Total Cost $0.097 $0.021
Input Price $0.150 / 1M $0.100 / 1M
Output Price $0.600 / 1M $0.150 / 1M
Total Input Tokens 140,559 144,416
Output Tokens 40,268 37,484
Reasoning Tokens 93,329 0
Response Time (avg) 10.80s 19.19s
Response Time (max) 59.15s 382.06s
Response Time (total) 237.60s 422.25s
Parameters 119B total (6B active) 9B
Availability Open source Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#267 Mistral Small 4

medium
Cost
$0.006
Time
47.9s
Tokens
9,857 tok

#272 Qwen3.5-9B

none
Reached the allocated time limit (300 seconds) without receiving showcase output.
Cost
$0.000
Time
300.0s
Tokens
0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mistral Small 4 4.4 5.1 33.3% 2 39.98s 7,636 11,635 54,715
Qwen3.5-9B 3.9 7.8 11.1% 1 5.60s 7,913 1,042 0

Quick Compare

Switch Comparison Pair