Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Mistral Small 4 vs Qwen3 Coder Next

The average score is effectively tied at 5.1 vs 5.1. Mistral Small 4 has the lower benchmark cost at $0.022 vs $0.025. Mistral Small 4 is faster at 1.20s vs 9.12s, with pass rates of 25.8% vs 25.8%.

Last updated at: 2026-07-28

Rank
#185
Total Output Tokens
9,812
Response Time (avg)
1.20s
Total Cost
$0.022
Rank
#186
Total Output Tokens
11,808
Response Time (avg)
9.12s
Total Cost
$0.025
Recommended model Mistral Small 4

It has the best score here (5.1), while responding about 7.6x faster than Qwen3 Coder Next.

Detailed comparison

Metric Mistral Small 4 Mistral Small 4 none Release: 2026-03-16 Qwen3 Coder Next Qwen3 Coder Next none Release: 2026-02-03
Score 5.1 5.1
Rank #185 #186
Reliability 10.0 10.0
Consistency 9.6 9.7
Tests Correct
Attempt pass rate 25.8% 25.8%
Flaky tests 1 1
Total Runs 66 66
Cost per result 0.432 0.488
Total Cost $0.022 $0.025
Input Price $0.150 / 1M $0.110 / 1M
Output Price $0.600 / 1M $0.800 / 1M
Total Input Tokens 104,708 134,218
Output Tokens 9,812 11,808
Reasoning Tokens 0 0
Response Time (avg) 1.20s 9.12s
Response Time (max) 13.16s 45.14s
Response Time (total) 26.38s 145.94s

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#185 Mistral Small 4

none
Cost
$0.002
Time
10.4s
Tokens
2,370 tok

#186 Qwen3 Coder Next

none
Invalid SVG
Cost
$0.058
Time
246.3s
Tokens
64,126 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mistral Small 4 3.7 9.7 0.0% 0 901ms 7,636 619 0
Qwen3 Coder Next 4.6 7.9 22.2% 1 2.22s 7,442 621 0

Quick Compare

Switch Comparison Pair