Navigate
Advertise here

Mercury 2.5 (low) vs Qwen3 Coder Next

The average score is effectively tied at 5.1 vs 5.1. Mercury 2.5 (low) has the lower benchmark cost at $0.011 vs $0.026. Mercury 2.5 (low) is faster at 1.35s vs 8.63s, with pass rates of 39.4% vs 25.8%.

Last updated at: 2026-09-08

Compared models

Rank
#265
Total Output Tokens
33,719
Response Time (avg)
1.35s
Total Cost
$0.011
Rank
#264
Total Output Tokens
11,808
Response Time (avg)
8.63s
Total Cost
$0.026
Recommended model Mercury 2.5 (low)

It has the best score here (5.1), while costing about 2.4x less than Qwen3 Coder Next.

Detailed comparison

Metric Mercury 2.5 Mercury 2.5 low Release: 2026-09-08 Qwen3 Coder Next Qwen3 Coder Next none Release: 2026-02-03
Score 5.1 5.1
Rank #265 #264
Reliability 9.8 10.0
Consistency 8.1 9.7
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 39.4% 25.8%
Flaky tests 5 1
Total Runs 66 66
Cost per result 0.177 0.488
Total Cost $0.011 $0.026
Input Price $0.040 / 1M $0.120 / 1M
Output Price $0.150 / 1M $0.800 / 1M
Total Input Tokens 138,020 134,227
Output Tokens 6,083 11,808
Reasoning Tokens 27,636 0
Response Time (avg) 1.35s 8.63s
Response Time (max) 7.48s 45.14s
Response Time (total) 29.74s 146.72s
Parameters ~100B 80B total (3B active)
Availability Closed Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#265 Mercury 2.5

low
Cost
$0.001
Time
2.4s
Tokens
1,236 tok

#264 Qwen3 Coder Next

none
Invalid SVG
Cost
$0.058
Time
246.3s
Tokens
64,126 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2.5 5.5 10.0 33.3% 0 972ms 7,909 521 2,269
Qwen3 Coder Next 4.6 7.9 22.2% 1 2.22s 7,442 621 0

Quick Compare

Switch Comparison Pair