Navigate
Advertise here

DeepSeek V4.1 Flash (medium) vs GPT-6 Sol (low)

DeepSeek V4.1 Flash (medium) leads on average score with 8.7 vs 8.7. GPT-6 Sol (low) has the lower benchmark cost at $0.477 vs $0.591. GPT-6 Sol (low) is faster at 8.72s vs 32.88s, with pass rates of 73.9% vs 87.0%.

Last updated at: 2026-10-07

Compared models

Rank
#53
Total Output Tokens
444,665
Response Time (avg)
32.88s
Total Cost
$0.591
Rank
#59
Total Output Tokens
13,211
Response Time (avg)
8.72s
Total Cost
$0.477
Recommended model GPT-6 Sol (low)

Its score stays close to the best score here (8.7 vs 8.7), while responding about 3.8x faster than DeepSeek V4.1 Flash (medium).

Detailed comparison

Metric DeepSeek V4.1 Flash DeepSeek V4.1 Flash medium Release: 2026-09-10 GPT-6 Sol GPT-6 Sol low Release: 2026-09-23
Score 8.7 8.7
Rank #53 #59
Reliability 9.6 10.0
Consistency 9.3 8.6
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 73.9% 87.0%
Flaky tests 2 4
Total Runs 69 69
Cost per result 1.757 2.650
Total Cost $0.591 $0.477
Input Price $0.013 / 1M $2.000 / 1M
Output Price $1.320 / 1M $10.000 / 1M
Cache Read Price $0.013 / 1M $0.200 / 1M
Cache Write Price N/A $2.500 / 1M
Total Input Tokens 280,643 172,359
Output Tokens 7,342 5,026
Reasoning Tokens 437,323 8,185
Response Time (avg) 32.88s 8.72s
Response Time (max) 276.31s 48.43s
Response Time (total) 756.34s 200.61s
Parameters 748B total (16B active) ~2T total (~150B active)
Availability Open source Closed

Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#53 DeepSeek V4.1 Flash

medium
Cost
$0.014
Time
34.7s
Tokens
11,541 tok

#59 GPT-6 Sol

low
Cost
$0.018
Time
17.4s
Tokens
1,850 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4.1 Flash 10.0 10.0 100.0% 0 34.45s 7,509 396 71,345
GPT-6 Sol 6.2 4.8 66.7% 2 8.74s 7,302 405 1,296

Switch Comparison Pair