Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

Anthropic: Claude Opus 4.8 vs DeepSeek: DeepSeek V4 Flash

Last updated at: 2026-05-28

Metric Claude Opus 4.8 Claude Opus 4.8 medium Release: 2026-05-28 DeepSeek V4 Flash DeepSeek V4 Flash high Release: 2026-04-24 Free Available
Score 8.7 7.6
Rank #12 #45
Reliability 10.0 10.0
Consistency 9.6 8.4
Tests Correct
Attempt pass rate 83.3% 73.3%
Flaky tests 1 4
Total Runs 60 60
Cost per result 6.285 0.309
Total Cost $1.006 $0.028
Input Price $5.000 / 1M $0.100 / 1M
Output Price $25.000 / 1M $0.200 / 1M
Output Tokens 23,201 10,302
Reasoning Tokens 5,901 115,740
Response Time (avg) 9.34s 46.36s
Response Time (max) 38.03s 218.13s
Response Time (total) 186.84s 927.27s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 3.95s 1,179 478
DeepSeek V4 Flash 8.3 10.0 75.0% 0 28.51s 140 7,770
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 14.97s 6,651 1,381
DeepSeek V4 Flash 6.8 10.0 50.0% 0 58.13s 387 27,101
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 9.8 10.0 100.0% 0 38.03s 5,260 1,588
DeepSeek V4 Flash 10.0 10.0 100.0% 0 76.57s 465 7,347
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 7.1 5.6 83.3% 1 12.29s 481 312
DeepSeek V4 Flash 10.0 10.0 100.0% 0 28.03s 201 1,179
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 5.3 10.0 33.3% 0 14.15s 7,477 900
DeepSeek V4 Flash 4.1 4.4 44.5% 2 100.31s 27 59,249
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 2.46s 237 0
DeepSeek V4 Flash 6.1 3.1 66.7% 1 25.15s 79 632
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 3.32s 373 320
DeepSeek V4 Flash 10.0 10.0 100.0% 0 15.36s 63 1,622
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 3.95s 791 483
DeepSeek V4 Flash 8.2 7.2 88.9% 1 26.11s 196 1,767
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 8.96s 301 225
DeepSeek V4 Flash 10.0 10.0 100.0% 0 74.73s 228 542
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 3.0 10.0 0.0% 0 6.14s 451 214
DeepSeek V4 Flash 3.0 10.0 0.0% 0 54.46s 8,516 8,531

Quick Compare

Switch Comparison Pair