Navigate
Advertise here

Mistral Small 4 vs Qwen3 Coder Next

The average score is effectively tied at 5.1 vs 5.1. Mistral Small 4 has the lower benchmark cost at $0.022 vs $0.026. Mistral Small 4 is faster at 1.20s vs 8.63s, with pass rates of 25.8% vs 25.8%.

Last updated at: 2026-09-10

Compared models

Rank
#266
Total Output Tokens
9,812
Response Time (avg)
1.20s
Total Cost
$0.022
Rank
#269
Total Output Tokens
11,808
Response Time (avg)
8.63s
Total Cost
$0.026
Recommended model Mistral Small 4

It has the best score here (5.1), while responding about 7.2x faster than Qwen3 Coder Next.

Detailed comparison

Metric Mistral Small 4 Mistral Small 4 none Release: 2026-03-16 Qwen3 Coder Next Qwen3 Coder Next none Release: 2026-02-03
Score 5.1 5.1
Rank #266 #269
Reliability 10.0 10.0
Consistency 9.6 9.7
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 25.8% 25.8%
Flaky tests 1 1
Total Runs 66 66
Cost per result 0.432 0.488
Total Cost $0.022 $0.026
Input Price $0.150 / 1M $0.120 / 1M
Output Price $0.600 / 1M $0.800 / 1M
Total Input Tokens 104,717 134,227
Output Tokens 9,812 11,808
Reasoning Tokens 0 0
Response Time (avg) 1.20s 8.63s
Response Time (max) 13.16s 45.14s
Response Time (total) 26.41s 146.72s
Parameters 119B total (6B active) 80B total (3B active)
Availability Open source Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#266 Mistral Small 4

none
Cost
$0.002
Time
10.4s
Tokens
2,370 tok

#269 Qwen3 Coder Next

none
Invalid SVG
Cost
$0.058
Time
246.3s
Tokens
64,126 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mistral Small 4 3.7 9.7 0.0% 0 901ms 7,636 619 0
Qwen3 Coder Next 4.6 7.9 22.2% 1 2.22s 7,442 621 0

Quick Compare

Switch Comparison Pair