Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

DeepSeek: DeepSeek V4 Pro vs Qwen: Qwen3.6 35B A3B

Summary

DeepSeek V4 Pro vs Qwen3.6 35B A3B benchmark comparison: DeepSeek V4 Pro leads on average score with 7.2 vs 6.7. DeepSeek V4 Pro has the lower benchmark cost at $0.034 vs $0.146. DeepSeek V4 Pro is faster at 6.41s vs 18.08s, with pass rates of 52.4% vs 63.5%.

Recommended model: DeepSeek V4 Pro - It has the best score here (7.2), while costing about 4.4x less than Qwen3.6 35B A3B.

Last updated at: 2026-06-18

Metric DeepSeek V4 Pro DeepSeek V4 Pro none Release: 2026-04-24 Qwen3.6 35B A3B Qwen3.6 35B A3B medium Release: 2026-04-20
Score 7.2 6.7
Rank #58 #75
Reliability 9.9 10.0
Consistency 8.8 9.6
Tests Correct
Attempt pass rate 52.4% 63.5%
Flaky tests 3 1
Total Runs 63 63
Cost per result 0.333 1.094
Total Cost $0.034 $0.146
Input Price $0.435 / 1M $0.140 / 1M
Output Price $0.870 / 1M $1.000 / 1M
Total Input Tokens 53,558 16,385
Output Tokens 11,424 19,632
Reasoning Tokens 0 130,219
Response Time (avg) 6.41s 18.08s
Response Time (max) 30.09s 86.11s
Response Time (total) 134.66s 343.61s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#58 DeepSeek V4 Pro

none
Invalid SVG
Cost
$0.000
Time
300.0s
Tokens
0 tok

#75 Qwen3.6 35B A3B

medium
Invalid SVG
Cost
$0.000
Time
300.0s
Tokens
0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 3.2 6.1 16.7% 2 4.02s 540 1,168 0
Qwen3.6 35B A3B 10.0 10.0 100.0% 0 6.02s 672 1,154 12,385
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 5.6 10.0 33.3% 0 13.38s 7,275 5,500 0
Qwen3.6 35B A3B 7.7 10.0 66.7% 0 50.55s 5,051 7,929 37,223
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 9.5 10.0 100.0% 0 23.74s 27,529 2,235 0
Qwen3.6 35B A3B 3.0 10.0 0.0% 0 0ms 0 0 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 10.0 10.0 100.0% 0 4.61s 7,568 200 0
Qwen3.6 35B A3B 10.0 10.0 100.0% 0 12.99s 7,776 2,591 9,968
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 5.3 10.0 33.3% 0 3.72s 666 24 0
Qwen3.6 35B A3B 5.3 7.2 44.4% 1 22.50s 771 6,193 39,116
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 5.0 10.0 0.0% 0 2.05s 471 126 0
Qwen3.6 35B A3B 4.4 9.9 0.0% 0 8.66s 516 129 4,569
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 6.3 5.8 66.7% 1 4.12s 627 713 0
Qwen3.6 35B A3B 10.0 10.0 100.0% 0 7.50s 699 219 7,404
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 10.0 10.0 100.0% 0 3.61s 594 442 0
Qwen3.6 35B A3B 8.0 10.0 66.7% 0 5.95s 696 655 9,228
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 10.0 10.0 100.0% 0 7.40s 8,105 328 0
Qwen3.6 35B A3B 3.0 10.0 0.0% 0 0ms 0 0 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 3.0 10.0 0.0% 0 5.76s 183 688 0
Qwen3.6 35B A3B 3.0 10.0 0.0% 0 32.90s 204 762 10,326

Quick Compare

Switch Comparison Pair