Navigate
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Mistral Large 4 (medium) vs Qwen3.5 Plus 2026-04-20

Mistral Large 4 (medium) leads on average score with 6.2 vs 6.1. Qwen3.5 Plus 2026-04-20 has the lower benchmark cost at $0.216 vs $1.456. Qwen3.5 Plus 2026-04-20 is faster at 18.49s vs 265.03s, with pass rates of 55.1% vs 44.9%.

Last updated at: 2026-10-06

Compared models

Rank
#240
Total Output Tokens
627,233
Response Time (avg)
265.03s
Total Cost
$1.456
Rank
#245
Total Output Tokens
69,664
Response Time (avg)
18.49s
Total Cost
$0.216
Recommended model Qwen3.5 Plus 2026-04-20

Its score stays close to the best score here (6.1 vs 6.2), while costing about 6.8x less than Mistral Large 4 (medium).

Detailed comparison

Metric Mistral Large 4 Mistral Large 4 medium Release: 2026-10-06 Qwen3.5 Plus 2026-04-20 Qwen3.5 Plus 2026-04-20 none Release: 2026-04-20
Score 6.2 6.1
Rank #240 #245
Reliability 9.2 9.8
Consistency 7.6 8.4
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 55.1% 44.9%
Flaky tests 7 5
Total Runs 69 69
Cost per result 16.168 2.816
Total Cost $1.456 $0.216
Input Price $0.680 / 1M $0.300 / 1M
Output Price $2.090 / 1M $1.800 / 1M
Cache Read Price $0.070 / 1M N/A
Cache Write Price N/A $0.375 / 1M
Total Input Tokens 212,024 299,984
Output Tokens 166,422 69,664
Reasoning Tokens 562,095 0
Response Time (avg) 265.03s 18.49s
Response Time (max) 1760.20s 206.05s
Response Time (total) 6095.69s 425.33s
Parameters 1.05T total (49B active) ~397B total (~17B active)
Availability Closed Closed

Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#240 Mistral Large 4

medium
Cost
$0.011
Time
63.4s
Tokens
5,291 tok

#245 Qwen3.5 Plus 2026-04-20

none
Cost
$0.008
Time
77.0s
Tokens
4,369 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mistral Large 4 3.4 7.2 22.2% 1 1201.87s 7,189 33,530 303,140
Qwen3.5 Plus 2026-04-20 3.9 7.8 11.1% 1 1.69s 7,913 480 0

Quick Compare

Switch Comparison Pair