Navigate
AI BENCHY
Advertise here

DeepSeek V4 Pro (high) vs Qwen3.8 27B (medium)

Qwen3.8 27B (medium) leads on average score with 7.8 vs 7.7. Local hardware cost was not measured, so cost comparisons are unavailable. Qwen3.8 27B (medium) is faster at 33.05s vs 92.50s, with pass rates of 63.6% vs 66.7%.

Last updated at: 2026-08-15

Rank
#75
Total Output Tokens
209,444
Response Time (avg)
92.50s
Total Cost
$0.555
Rank
#68
Total Output Tokens
136,162
Response Time (avg)
33.05s
Total Cost
N/A
Recommended model DeepSeek V4 Pro (high)

It has the best overall balance of score, reliability, cost, and response time in this comparison.

Detailed comparison

Metric DeepSeek V4 Pro DeepSeek V4 Pro high Release: 2026-04-24 Qwen3.8 27B Qwen3.8 27B medium Release: 2026-08-14
Score 7.7 7.8
Rank #75 #68
Reliability 10.0 9.6
Consistency 7.7 9.7
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 63.6% 66.7%
Flaky tests 6 1
Total Runs 66 66
Cost per result 2.251 N/A
Total Cost $0.555 N/A
Input Price $1.169 / 1M N/A
Output Price $2.337 / 1M N/A
Total Input Tokens 90,757 98,266
Output Tokens 22,928 1,861
Reasoning Tokens 186,516 134,301
Response Time (avg) 92.50s 33.05s
Response Time (max) 416.76s 214.63s
Response Time (total) 2035.01s 694.04s
Parameters 1.6T total (49B active) 27.3B
Availability Open source Weights available

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#75 DeepSeek V4 Pro

high
Cost
$0.023
Time
257.6s
Tokens
14,870 tok

#68 Qwen3.8 27B

medium
Cost
N/A
Time
43.4s
Tokens
2,956 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 6.3 8.7 33.3% 0 243.00s 5,090 383 84,580
Qwen3.8 27B 10.0 10.0 100.0% 0 40.50s 7,893 585 26,808

Quick Compare

Switch Comparison Pair