Navigate
AI BENCHY
Advertise here

Mistral Small 4 (medium) vs Qwen3 Coder Next

The average score is effectively tied at 5.1 vs 5.1. Qwen3 Coder Next has the lower benchmark cost at $0.026 vs $0.096. Qwen3 Coder Next is faster at 9.12s vs 10.77s, with pass rates of 42.4% vs 25.8%.

Last updated at: 2026-07-28

Rank
#187
Total Output Tokens
131,824
Response Time (avg)
10.77s
Total Cost
$0.096
Rank
#186
Total Output Tokens
11,808
Response Time (avg)
9.12s
Total Cost
$0.026
Recommended model Qwen3 Coder Next

It has the best score here (5.1), while costing about 3.7x less than Mistral Small 4 (medium).

Detailed comparison

Metric Mistral Small 4 Mistral Small 4 medium Release: 2026-03-16 Qwen3 Coder Next Qwen3 Coder Next none Release: 2026-02-03
Score 5.1 5.1
Rank #187 #186
Reliability 10.0 10.0
Consistency 7.0 9.7
Tests Correct
Attempt pass rate 42.4% 25.8%
Flaky tests 8 1
Total Runs 66 66
Cost per result 1.913 0.488
Total Cost $0.096 $0.026
Input Price $0.150 / 1M $0.120 / 1M
Output Price $0.600 / 1M $0.800 / 1M
Total Input Tokens 140,494 134,218
Output Tokens 39,462 11,808
Reasoning Tokens 92,362 0
Response Time (avg) 10.77s 9.12s
Response Time (max) 59.15s 45.14s
Response Time (total) 236.94s 145.94s

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#187 Mistral Small 4

medium
Cost
$0.006
Time
47.9s
Tokens
9,857 tok

#186 Qwen3 Coder Next

none
Invalid SVG
Cost
$0.058
Time
246.3s
Tokens
64,126 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mistral Small 4 4.4 5.1 33.3% 2 39.98s 7,636 11,635 54,715
Qwen3 Coder Next 4.6 7.9 22.2% 1 2.22s 7,442 621 0

Quick Compare

Switch Comparison Pair