Navigate
Advertise here

DeepSeek V4.1 Flash (high) vs GPT-6 Sol (low)

DeepSeek V4.1 Flash (high) leads on average score with 8.7 vs 8.7. GPT-6 Sol (low) has the lower benchmark cost at $0.477 vs $0.617. GPT-6 Sol (low) is faster at 8.72s vs 33.27s, with pass rates of 75.4% vs 87.0%.

Last updated at: 2026-10-07

Compared models

Rank
#52
Total Output Tokens
464,156
Response Time (avg)
33.27s
Total Cost
$0.617
Rank
#59
Total Output Tokens
13,211
Response Time (avg)
8.72s
Total Cost
$0.477
Recommended model GPT-6 Sol (low)

Its score stays close to the best score here (8.7 vs 8.7), while responding about 3.8x faster than DeepSeek V4.1 Flash (high).

Detailed comparison

Metric DeepSeek V4.1 Flash DeepSeek V4.1 Flash high Release: 2026-09-10 GPT-6 Sol GPT-6 Sol low Release: 2026-09-23
Score 8.7 8.7
Rank #52 #59
Reliability 9.7 10.0
Consistency 9.0 8.6
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 75.4% 87.0%
Flaky tests 3 4
Total Runs 69 69
Cost per result 1.843 2.650
Total Cost $0.617 $0.477
Input Price $0.013 / 1M $2.000 / 1M
Output Price $1.320 / 1M $10.000 / 1M
Cache Read Price $0.013 / 1M $0.200 / 1M
Cache Write Price N/A $2.500 / 1M
Total Input Tokens 282,490 172,359
Output Tokens 8,621 5,026
Reasoning Tokens 455,535 8,185
Response Time (avg) 33.27s 8.72s
Response Time (max) 205.53s 48.43s
Response Time (total) 765.12s 200.61s
Parameters 748B total (16B active) ~2T total (~150B active)
Availability Open source Closed

Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#52 DeepSeek V4.1 Flash

high
Cost
$0.035
Time
111.0s
Tokens
29,201 tok

#59 GPT-6 Sol

low
Cost
$0.018
Time
17.4s
Tokens
1,850 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4.1 Flash 10.0 10.0 100.0% 0 49.19s 7,509 376 108,285
GPT-6 Sol 6.2 4.8 66.7% 2 8.74s 7,302 405 1,296

Switch Comparison Pair