Advertise here
#119

Qwen3.8 Flash Next

Qwen Release: 2026-10-04 Tested on: 2026-10-05 01:35 qwen/qwen3.8-flash-next::medium
180B total (6B active)MoEWeights available
(xhigh) (medium) (low) (none)

Summary

Qwen3.8 Flash Next scores 7.6 on AI BENCHY and ranks #119. It has 10.0 reliability, a 75.4% pass rate, ~$0.047 total cost, and 24.79s average response time.

What makes Qwen3.8 Flash Next unique: It stands out most in Coding, where it ranks #1, while Agentic is its weakest area at #17.

Model facts

Researched on 2026-10-04

Reported
Parameters
180B total (6B active)
Architecture
MoE
Availability
Weights available
License
Qwen Community License 1.0

Qwen reports a 125B MoE language core with 6B activated, plus 51B n-gram embeddings and 4B MTP. The 180B total sums those components; 6B is the activated language-core count, not all lookup/draft work. Public weights use the custom Qwen Community License 1.0. This benchmark uses the ISTA-DASLab GSQ-RCO IQ3_S GGUF through Strata 0.1.39 on a local RTX 3090, with vision disabled and a 65,536-token runtime context. The compression is a quantized version of the exact public checkpoint, not the managed Qwen3.8-Flash API model.

Score

7.6

Consistency

7.9

Total Cost (Current Price)

~$0.047

Total Output Tokens

143,965

Total Input Tokens

278,792

Estimated benchmark electricity cost

~$0.047

Tests Correct

Wrong Tests: 8

Attempt pass rate: 75.4%

Flaky tests

6

Flaky tests had mixed outcomes across runs (at least one pass and one fail).

Response Time (avg)

24.79s

Response Time (max): 159.62s

Response Time (total): 570.26s

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#119 Qwen3.8 Flash Next

medium
Cost
~$0.002
Time
40.7s
Tokens
4,059 tok

Charts

Choose the first model, then click a second model to open a side-by-side page.

Total Output Tokens

Score vs Total Output Tokens

Quick Compare

Category Breakdown

Category Score Consistency Tests Correct
Agentic 4.7 3.1
Anti-AI Tricks 10.0 10.0
Coding 10.0 10.0
Combined 6.4 5.8
Data parsing and extraction 10.0 10.0
Domain specific 2.9 4.4
General Intelligence 4.8 3.2
Instructions following 9.8 10.0
Puzzle Solving 8.2 7.2
Tool Calling 10.0 10.0
Trivia 3.0 10.0

Compared models