Navigate
AI BENCHY
Advertise here

Mercury 2.5 Preview (low) vs Qwen3.5-9B

Qwen3.5-9B leads on average score with 5.1 vs 5.0. Mercury 2.5 Preview (low) has the lower benchmark cost at $0.011 vs $0.021. Mercury 2.5 Preview (low) is faster at 1.31s vs 19.19s, with pass rates of 47.0% vs 19.7%.

Last updated at: 2026-09-02

Rank
#254
Total Output Tokens
31,975
Response Time (avg)
1.31s
Total Cost
$0.011
Rank
#250
Total Output Tokens
37,484
Response Time (avg)
19.19s
Total Cost
$0.021
Recommended model Mercury 2.5 Preview (low)

Its score stays close to the best score here (5.0 vs 5.1), while costing about 2.0x less than Qwen3.5-9B.

Detailed comparison

Metric Mercury 2.5 Preview Mercury 2.5 Preview low Release: 2026-09-02 Qwen3.5-9B Qwen3.5-9B none Release: 2026-03-02
Score 5.0 5.1
Rank #254 #250
Reliability 10.0 10.0
Consistency 7.0 9.7
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 47.0% 19.7%
Flaky tests 8 1
Total Runs 66 66
Cost per result 0.169 0.490
Total Cost $0.011 $0.021
Input Price $0.040 / 1M $0.100 / 1M
Output Price $0.150 / 1M $0.150 / 1M
Total Input Tokens 133,525 144,416
Output Tokens 7,259 37,484
Reasoning Tokens 24,716 0
Response Time (avg) 1.31s 19.19s
Response Time (max) 7.29s 382.06s
Response Time (total) 28.77s 422.25s
Parameters ~100B 9B
Availability Closed Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#254 Mercury 2.5 Preview

low
Cost
$0.001
Time
1.5s
Tokens
1,169 tok

#250 Qwen3.5-9B

none
Invalid SVG
Cost
$0.000
Time
300.0s
Tokens
0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2.5 Preview 5.5 10.0 33.3% 0 972ms 7,794 695 2,263
Qwen3.5-9B 3.9 7.8 11.1% 1 5.60s 7,913 1,042 0

Quick Compare

Switch Comparison Pair