Navigate
Advertise here

DeepSeek V4.1 Flash vs gpt-oss-120b (medium)

gpt-oss-120b (medium) leads on average score with 5.9 vs 5.8. gpt-oss-120b (medium) has the lower benchmark cost at $0.030 vs $0.391. DeepSeek V4.1 Flash is faster at 16.63s vs 31.58s, with pass rates of 36.2% vs 47.8%.

Last updated at: 2026-10-07

Compared models

Rank
#266
Total Output Tokens
292,779
Response Time (avg)
16.63s
Total Cost
$0.391
Rank
#258
Total Output Tokens
108,704
Response Time (avg)
31.58s
Total Cost
$0.030
Recommended model gpt-oss-120b (medium)

It has the best score here (5.9), while costing about 13.5x less than DeepSeek V4.1 Flash.

Detailed comparison

Metric DeepSeek V4.1 Flash DeepSeek V4.1 Flash none Release: 2026-09-10 gpt-oss-120b gpt-oss-120b medium Release: 2025-08-05
Score 5.8 5.9
Rank #266 #258
Reliability 9.7 10.0
Consistency 9.1 8.1
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 36.2% 47.8%
Flaky tests 3 5
Total Runs 69 69
Cost per result 2.806 0.335
Total Cost $0.391 $0.030
Input Price $0.013 / 1M $0.037 / 1M
Output Price $1.320 / 1M $0.170 / 1M
Cache Read Price $0.013 / 1M N/A
Cache Write Price N/A N/A
Total Input Tokens 337,756 285,293
Output Tokens 292,779 29,565
Reasoning Tokens 0 89,266
Response Time (avg) 16.63s 31.58s
Response Time (max) 247.66s 203.90s
Response Time (total) 382.53s 536.78s
Parameters 748B total (16B active) 117B total (5.1B active)
Availability Open source Open source

Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#266 DeepSeek V4.1 Flash

none
Cost
$0.021
Time
55.4s
Tokens
17,181 tok

#258 gpt-oss-120b

medium
Cost
$0.001
Time
26.7s
Tokens
555 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4.1 Flash 5.5 10.0 33.3% 0 83.66s 7,284 263,438 0
gpt-oss-120b 5.9 7.0 55.6% 1 38.37s 7,782 3,365 11,973

Switch Comparison Pair