Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Anthropic: Claude Sonnet 5 vs DeepSeek: DeepSeek V4 Flash

Summary

Claude Sonnet 5 vs DeepSeek V4 Flash benchmark comparison: DeepSeek V4 Flash leads on average score with 8.3 vs 7.9. DeepSeek V4 Flash has the lower benchmark cost at $0.029 vs $0.550. Claude Sonnet 5 is faster at 9.94s vs 45.85s, with pass rates of 79.4% vs 74.6%.

Recommended model: DeepSeek V4 Flash - It has the best score here (8.3), while costing about 19.3x less than Claude Sonnet 5.

Last updated at: 2026-06-30

Metric Claude Sonnet 5 Claude Sonnet 5 medium Release: 2026-06-30 DeepSeek V4 Flash DeepSeek V4 Flash high Release: 2026-04-24
Score 7.9 8.3
Rank #30 #23
Reliability 10.0 10.0
Consistency 9.0 8.5
Tests Correct
Attempt pass rate 79.4% 74.6%
Flaky tests 3 4
Total Runs 63 63
Cost per result 3.662 0.299
Total Cost $0.550 $0.029
Input Price $2.000 / 1M $0.098 / 1M
Output Price $10.000 / 1M $0.196 / 1M
Total Input Tokens 67,416 39,745
Output Tokens 34,012 10,310
Reasoning Tokens 7,673 123,501
Response Time (avg) 9.94s 45.85s
Response Time (max) 56.94s 218.13s
Response Time (total) 208.71s 962.79s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#30 Claude Sonnet 5

medium
Cost
$0.007
Time
6.4s
Tokens
832 tok

#23 DeepSeek V4 Flash

high
Cost
$0.003
Time
93.1s
Tokens
7,926 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 10.0 10.0 100.0% 0 3.80s 834 1,220 446
DeepSeek V4 Flash 8.3 10.0 75.0% 0 28.51s 540 140 7,770
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 9.0 7.9 88.9% 1 17.28s 10,590 13,153 2,379
DeepSeek V4 Flash 7.8 10.0 66.7% 0 50.60s 7,279 395 34,862
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 4.5 2.1 66.7% 1 37.01s 29,394 4,848 2,170
DeepSeek V4 Flash 10.0 10.0 100.0% 0 76.57s 14,016 465 7,347
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 10.0 10.0 100.0% 0 3.16s 10,503 312 0
DeepSeek V4 Flash 10.0 10.0 100.0% 0 28.03s 7,290 201 1,179
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 7.7 10.0 66.7% 0 20.38s 975 12,140 1,994
DeepSeek V4 Flash 4.1 4.4 44.5% 2 100.31s 666 27 59,249
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 4.8 3.2 33.3% 1 4.32s 708 264 0
DeepSeek V4 Flash 6.1 3.1 66.7% 1 25.15s 471 79 632
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 9.9 10.0 100.0% 0 3.10s 909 318 269
DeepSeek V4 Flash 10.0 10.0 100.0% 0 15.36s 627 63 1,622
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 7.7 10.0 66.7% 0 2.98s 894 407 121
DeepSeek V4 Flash 8.2 7.2 88.9% 1 26.11s 594 196 1,767
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 10.0 10.0 100.0% 0 10.70s 12,351 433 90
DeepSeek V4 Flash 10.0 10.0 100.0% 0 74.73s 8,079 228 542
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 3.0 10.0 0.0% 0 7.06s 258 917 204
DeepSeek V4 Flash 3.0 10.0 0.0% 0 54.46s 183 8,516 8,531

Quick Compare

Switch Comparison Pair