Navigate
AI BENCHY
Advertise here

Gemini 3.1 Flash Lite (low) vs Qwen3.5 Plus 2026-02-15

The average score is effectively tied at 6.5 vs 6.4. Qwen3.5 Plus 2026-02-15 has the lower benchmark cost at $0.073 vs $0.621. Qwen3.5 Plus 2026-02-15 is faster at 9.85s vs 16.26s, with pass rates of 59.1% vs 48.5%.

Last updated at: 2026-07-28

Rank
#118
Total Output Tokens
397,885
Response Time (avg)
16.26s
Total Cost
$0.621
Rank
#120
Total Output Tokens
29,370
Response Time (avg)
9.85s
Total Cost
$0.073
Recommended model Qwen3.5 Plus 2026-02-15

It has the best score here (6.4), while costing about 8.6x less than Gemini 3.1 Flash Lite (low).

Detailed comparison

Metric Gemini 3.1 Flash Lite Gemini 3.1 Flash Lite low Release: 2026-05-08 Qwen3.5 Plus 2026-02-15 Qwen3.5 Plus 2026-02-15 none Release: 2026-02-15
Score 6.5 6.4
Rank #118 #120
Reliability 10.0 10.0
Consistency 9.2 9.4
Tests Correct
Attempt pass rate 59.1% 48.5%
Flaky tests 2 2
Total Runs 66 66
Cost per result 5.170 0.751
Total Cost $0.621 $0.073
Input Price $0.250 / 1M $0.260 / 1M
Output Price $1.500 / 1M $1.560 / 1M
Total Input Tokens 94,224 102,646
Output Tokens 7,759 29,370
Reasoning Tokens 390,126 0
Response Time (avg) 16.26s 9.85s
Response Time (max) 318.02s 123.00s
Response Time (total) 357.64s 157.63s

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#118 Gemini 3.1 Flash Lite

low
Cost
$0.003
Time
4.0s
Tokens
1,479 tok

#120 Qwen3.5 Plus 2026-02-15

none
Cost
$0.012
Time
153.2s
Tokens
7,787 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 5.5 10.0 33.3% 0 1.53s 8,132 471 1,072
Qwen3.5 Plus 2026-02-15 4.3 7.9 11.1% 1 2.05s 7,913 473 0

Quick Compare

Switch Comparison Pair