Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Google: Gemma 4 31B vs Qwen: Qwen3.5-Flash

Summary

Gemma 4 31B vs Qwen3.5-Flash benchmark comparison: Qwen3.5-Flash leads on average score with 6.8 vs 6.1. Gemma 4 31B has the lower benchmark cost at $0.004 vs $0.080. Gemma 4 31B is faster at 4.05s vs 63.29s, with pass rates of 47.6% vs 71.4%.

Recommended model: Gemma 4 31B - Its score stays close to the best score here (6.1 vs 6.8), while costing about 26.5x less than Qwen3.5-Flash.

Last updated at: 2026-07-02

Metric Gemma 4 31B Gemma 4 31B none Release: 2026-04-02 Free Available Qwen3.5-Flash Qwen3.5-Flash medium Release: 2026-02-24
Score 6.1 6.8
Rank #101 #73
Reliability 10.0 10.0
Consistency 10.0 8.1
Tests Correct
Attempt pass rate 47.6% 71.4%
Flaky tests 0 5
Total Runs 63 63
Cost per result 0.034 0.871
Total Cost $0.004 $0.080
Input Price $0.120 / 1M $0.065 / 1M
Output Price $0.350 / 1M $0.260 / 1M
Total Input Tokens 20,911 38,926
Output Tokens 1,407 2,088
Reasoning Tokens 0 294,598
Response Time (avg) 4.05s 63.29s
Response Time (max) 26.13s 234.29s
Response Time (total) 76.87s 1265.85s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#101 Gemma 4 31B

none
Cost
$0.001
Time
12.8s
Tokens
795 tok

#73 Qwen3.5-Flash

medium
Cost
$0.002
Time
25.8s
Tokens
4,294 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 6.5 10.0 50.0% 0 1.85s 852 45 0
Qwen3.5-Flash 10.0 10.0 100.0% 0 59.11s 672 383 32,992
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 5.5 10.0 33.3% 0 11.19s 8,381 735 0
Qwen3.5-Flash 3.7 7.2 22.2% 1 58.87s 6,685 302 90,081
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 3.0 10.0 0.0% 0 0ms 0 0 0
Qwen3.5-Flash 10.0 10.0 100.0% 0 17.78s 14,934 483 8,270
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 10.0 10.0 100.0% 0 2.25s 8,352 285 0
Qwen3.5-Flash 7.3 5.9 83.3% 1 56.99s 6,061 235 16,237
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 7.7 10.0 66.7% 0 3.22s 903 27 0
Qwen3.5-Flash 5.3 7.2 44.4% 1 146.50s 581 58 43,615
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 10.0 10.0 100.0% 0 2.09s 576 117 0
Qwen3.5-Flash 6.1 3.1 66.7% 1 40.05s 516 99 38,486
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 6.5 10.0 50.0% 0 2.84s 795 78 0
Qwen3.5-Flash 10.0 10.0 100.0% 0 63.49s 699 98 14,139
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 6.5 10.0 33.3% 0 4.23s 828 108 0
Qwen3.5-Flash 8.2 7.2 88.9% 1 27.61s 381 89 12,457
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 3.0 10.0 0.0% 0 0ms 0 0 0
Qwen3.5-Flash 10.0 10.0 100.0% 0 10.33s 8,193 309 1,284
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 3.0 10.0 0.0% 0 1.25s 224 12 0
Qwen3.5-Flash 3.0 10.0 0.0% 0 48.98s 204 32 37,037

Quick Compare

Switch Comparison Pair