AI BENCHY
Advertise here
#46

Qwen3.8 2.4T A95B

Qwen Release: 2026-08-12 Tested on: 2026-08-12 23:44 qwen/qwen3.8-2.4t-a95b::low
2.4T total (95B active)MoEWeights available
(high) (low)

Summary

Qwen3.8 2.4T A95B scores 8.2 on AI BENCHY and ranks #46. It has 9.1 reliability, a 78.8% pass rate, $3.113 total cost, and 132.39s average response time.

What makes Qwen3.8 2.4T A95B unique: It stands out most in Puzzle Solving, where it ranks #1, while Coding is its weakest area at #13. It uses unusually many reasoning tokens, which can help explain its slower or more expensive runs.

Model facts

Researched on 2026-08-12

Reported
Parameters
2.4T total (95B active)
Architecture
MoE
Availability
Weights available
License
Apache-2.0

Exact open-weight sparse checkpoint; public weights use the Qwen Apache-2.0 release license.

Score

8.2

Consistency

8.9

Total Output Tokens

554,485

Total Input Tokens

113,340

Input Price

$2.000 / 1M

Output Price

$6.000 / 1M

Tests Correct

Wrong Tests: 6

Attempt pass rate: 78.8%

Flaky tests

3

Flaky tests had mixed outcomes across runs (at least one pass and one fail).

Response Time (avg)

132.39s

Response Time (max): 534.21s

Response Time (total): 2912.56s

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#46 Qwen3.8 2.4T A95B

low
Cost
$0.100
Time
193.2s
Tokens
16,767 tok

Price History

Historical pricing data for this model from OpenRouter.

Date Input Price Output Price
2026-08-12 23:12 $2.000 / 1M $6.000 / 1M

Charts

Choose the first model, then click a second model to open a side-by-side page.

Total Output Tokens

Score vs Total Output Tokens

Quick Compare

Category Breakdown

Category Score Consistency Tests Correct
Anti-AI Tricks 10.0 10.0
Coding 7.6 7.2
Combined 8.2 6.9
Data parsing and extraction 10.0 10.0
Domain specific 5.5 9.3
General Intelligence 6.1 3.1
Instructions following 9.8 10.0
Puzzle Solving 10.0 10.0
Tool Calling 10.0 10.0
Trivia 3.0 10.0

Compared models