Navigate
AI BENCHY
Your ad here

AI BENCHY Compare

DeepSeek: DeepSeek V4 Flash vs OpenAI: GPT-5.4 Mini

Last updated at: 2026-05-01

Metric DeepSeek V4 Flash DeepSeek V4 Flash none Release: 2026-04-24 GPT-5.4 Mini GPT-5.4 Mini none Release: 2026-03-17
Score 5.3 5.1
Rank #109 #117
Reliability N/A N/A
Consistency 9.1 8.6
Tests Correct
Attempt pass rate 33.3% 35.2%
Flaky tests 2 3
Total Runs 54 54
Cost per result 0.147 0.630
Total Cost $0.008 $0.032
Input Price $0.140 / 1M $0.750 / 1M
Output Price $0.280 / 1M $4.500 / 1M
Output Tokens 4,444 2,418
Reasoning Tokens 0 0
Response Time (avg) 29.39s 1.17s
Response Time (max) 111.96s 2.52s
Response Time (total) 529.10s 21.01s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 3.0 10.0 0.0% 0 20.18s 174 0
GPT-5.4 Mini 3.1 8.1 8.3% 1 929ms 654 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 6.3 10.0 0.0% 0 24.04s 471 0
GPT-5.4 Mini 10.0 10.0 100.0% 0 1.19s 333 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 4.5 2.1 66.7% 1 111.96s 2,664 0
GPT-5.4 Mini 3.0 10.0 0.0% 0 2.52s 298 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 23.79s 195 0
GPT-5.4 Mini 10.0 10.0 100.0% 0 1.30s 222 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 5.3 10.0 33.3% 0 19.73s 18 0
GPT-5.4 Mini 3.5 4.4 33.3% 2 937ms 88 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 4.2 9.9 0.0% 0 23.74s 67 0
GPT-5.4 Mini 4.8 10.0 0.0% 0 1.82s 174 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 6.5 10.0 50.0% 0 17.54s 321 0
GPT-5.4 Mini 6.3 10.0 50.0% 0 728ms 101 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 3.1 7.3 11.1% 1 22.96s 207 0
GPT-5.4 Mini 5.4 10.0 33.3% 0 860ms 293 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 77.93s 327 0
GPT-5.4 Mini 3.0 10.0 0.0% 0 2.32s 255 0

Quick Compare

Switch Comparison Pair