Navigate
AI BENCHY
Advertise here

DeepSeek V3.2 vs Mercury 2.5 Preview (low)

The average score is effectively tied at 5.0 vs 5.0. Mercury 2.5 Preview (low) has the lower benchmark cost at $0.011 vs $0.054. Mercury 2.5 Preview (low) is faster at 1.31s vs 17.89s, with pass rates of 36.4% vs 47.0%.

Last updated at: 2026-09-02

Rank
#256
Total Output Tokens
42,099
Response Time (avg)
17.89s
Total Cost
$0.054
Rank
#254
Total Output Tokens
31,975
Response Time (avg)
1.31s
Total Cost
$0.011
Recommended model Mercury 2.5 Preview (low)

It has the best score here (5.0), while costing about 5.3x less than DeepSeek V3.2.

Detailed comparison

Metric DeepSeek V3.2 DeepSeek V3.2 none Release: 2025-12-01 Mercury 2.5 Preview Mercury 2.5 Preview low Release: 2026-09-02
Score 5.0 5.0
Rank #256 #254
Reliability 10.0 10.0
Consistency 8.1 7.0
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 36.4% 47.0%
Flaky tests 5 8
Total Runs 66 66
Cost per result 0.870 0.169
Total Cost $0.054 $0.011
Input Price $0.269 / 1M $0.040 / 1M
Output Price $0.400 / 1M $0.150 / 1M
Total Input Tokens 135,831 133,525
Output Tokens 42,099 7,259
Reasoning Tokens 0 24,716
Response Time (avg) 17.89s 1.31s
Response Time (max) 115.89s 7.29s
Response Time (total) 393.53s 28.77s
Parameters 671B total (37B active) ~100B
Availability Open source Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#256 DeepSeek V3.2

none
Cost
$0.002
Time
7.0s
Tokens
1,046 tok

#254 Mercury 2.5 Preview

low
Cost
$0.001
Time
1.5s
Tokens
1,169 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V3.2 3.1 6.9 11.1% 1 14.54s 7,279 4,528 0
Mercury 2.5 Preview 5.5 10.0 33.3% 0 972ms 7,794 695 2,263

Quick Compare

Switch Comparison Pair