Navigate
Advertise here

DeepSeek V4.1 Flash (low) vs Qwen3.8 Flash Next (low)

DeepSeek V4.1 Flash (low) leads on average score with 8.0 vs 8.0. Qwen3.8 Flash Next (low) has the lower benchmark cost at ~$0.045 vs $0.526. Qwen3.8 Flash Next (low) is faster at 23.66s vs 28.04s, with pass rates of 73.9% vs 79.7%.

Last updated at: 2026-10-05

Compared models

Rank
#91
Total Output Tokens
371,547
Response Time (avg)
28.04s
Total Cost
$0.526
Rank
#96
Total Output Tokens
135,295
Response Time (avg)
23.66s
Total Cost
~$0.045
Recommended model Qwen3.8 Flash Next (low)

Its score stays close to the best score here (8.0 vs 8.0), while costing about 11.7x less than DeepSeek V4.1 Flash (low).

Detailed comparison

Metric DeepSeek V4.1 Flash DeepSeek V4.1 Flash low Release: 2026-09-10 Qwen3.8 Flash Next Qwen3.8 Flash Next low Release: 2026-10-04
Score 8.0 8.0
Rank #91 #96
Reliability 9.6 10.0
Consistency 8.7 8.0
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 73.9% 79.7%
Flaky tests 4 6
Total Runs 69 69
Cost per result 1.576 ~0.299
Total Cost $0.526 ~$0.045
Input Price $0.003 / 1M N/A
Output Price $2.400 / 1M N/A
Cache Read Price $0.003 / 1M N/A
Cache Write Price N/A N/A
Total Input Tokens 266,849 342,163
Output Tokens 8,284 1,518
Reasoning Tokens 363,263 133,777
Response Time (avg) 28.04s 23.66s
Response Time (max) 149.84s 166.72s
Response Time (total) 645.03s 544.16s
Parameters 748B total (16B active) 180B total (6B active)
Availability Open source Weights available

Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#91 DeepSeek V4.1 Flash

low
Cost
$0.017
Time
43.6s
Tokens
14,069 tok

#96 Qwen3.8 Flash Next

low
Cost
~$0.002
Time
47.2s
Tokens
5,270 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4.1 Flash 10.0 10.0 100.0% 0 40.66s 7,509 378 81,978
Qwen3.8 Flash Next 10.0 10.0 100.0% 0 29.46s 8,127 397 24,147

Quick Compare

Switch Comparison Pair

Gemini 3.8 FlashmediumvsQwen3.8 Flash NextlowGPT-5.4 NanomediumvsQwen3.8 Flash NextlowClaude Opus 4.7mediumvsDeepSeek V4.1 FlashlowDeepSeek V4.1 FlashlowvsGPT-5.6 LunahighQwen3.8 Flash NextlowvsGLM 5.3 FlashXmaxClaude Fable 5.1highvsQwen3.8 Flash NextlowNemotron 3 UltramediumFree AvailablevsQwen3.8 Flash NextlowDeepSeek V4.1 FlashlowvsNemotron 3 UltramediumFree AvailableGPT-5.4 MinimediumvsQwen3.8 Flash NextlowGPT-5.6 LunahighvsQwen3.8 Flash NextlowDeepSeek V4.1 FlashlowvsGPT-5.4 NanomediumDeepSeek V4.1 FlashlowvsGemini 3.8 Flashmedium