Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

DeepSeek: DeepSeek V3.2 vs Inception: Mercury 2

Last updated at: 2026-06-03

Metric DeepSeek V3.2 DeepSeek V3.2 medium Release: 2025-12-01 Mercury 2 Mercury 2 medium Release: 2026-02-24
Score 6.8 6.5
Rank #77 #89
Reliability 10.0 10.0
Consistency 7.5 8.8
Tests Correct
Attempt pass rate 63.3% 51.7%
Flaky tests 6 3
Total Runs 60 60
Cost per result 0.368 0.611
Total Cost $0.033 $0.055
Input Price $0.229 / 1M $0.250 / 1M
Output Price $0.344 / 1M $0.750 / 1M
Total Input Tokens 35,744 32,570
Output Tokens 7,177 4,022
Reasoning Tokens 68,297 58,405
Response Time (avg) 53.34s 2.27s
Response Time (max) 189.03s 14.63s
Response Time (total) 1066.71s 43.20s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V3.2 8.2 7.9 83.3% 1 24.23s 448 3,247 6,953
Mercury 2 6.9 9.9 50.0% 0 1.12s 554 2,546 2,609
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V3.2 3.9 5.8 33.3% 1 184.97s 3,128 640 21,230
Mercury 2 7.2 6.5 66.7% 1 2.29s 4,519 270 8,514
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V3.2 10.0 10.0 100.0% 0 93.11s 14,283 571 6,296
Mercury 2 10.0 10.0 100.0% 0 3.28s 12,909 268 4,887
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V3.2 10.0 10.0 100.0% 0 36.09s 7,388 207 7,693
Mercury 2 7.3 5.9 83.3% 1 1.11s 6,234 183 1,656
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V3.2 2.9 4.4 22.2% 2 24.27s 472 21 6,838
Mercury 2 2.9 7.2 11.1% 1 6.48s 695 41 30,754
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V3.2 3.4 2.5 33.3% 1 58.29s 314 49 2,189
Mercury 2 4.8 10.0 0.0% 0 821ms 456 137 542
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V3.2 10.0 10.0 100.0% 0 35.78s 627 1,397 2,845
Mercury 2 10.0 10.0 100.0% 0 1.07s 340 14 958
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V3.2 7.0 7.2 55.6% 1 37.69s 594 518 6,375
Mercury 2 5.4 10.0 33.3% 0 949ms 601 361 2,781
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V3.2 10.0 10.0 100.0% 0 34.81s 8,307 507 859
Mercury 2 10.0 10.0 100.0% 0 1.89s 6,080 180 1,956
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V3.2 3.0 10.0 0.0% 0 83.99s 183 20 7,019
Mercury 2 3.0 10.0 0.0% 0 2.58s 182 22 3,748

Quick Compare

Switch Comparison Pair