Navigate
AI BENCHY
Your ad here

AI BENCHY Compare

Qwen: Qwen3.5-9B vs Z.ai: GLM 4.7 Flash

Last updated at: 2026-03-12

Metric Qwen3.5-9B Qwen3.5-9B medium Release: 2026-03-02 GLM 4.7 Flash GLM 4.7 Flash medium Release: 2026-01-19
Rank #66 #62
Avg Score 2.6 3.1
Consistency 7.4 6.4
Cost per result 0.779 1.040
Total Cost $0.024 $0.042
Tests Correct
Attempt pass rate 35.4% 41.7%
Flaky tests 5 7
Total Runs 48 48
Output Tokens 17,930 38,682
Reasoning Tokens 139,706 64,952
Response Time (avg) 71.44s 36.84s
Response Time (max) 226.38s 174.55s
Response Time (total) 928.77s 331.58s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Avg Score vs Response Time (avg)

Total Output Tokens

Avg Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-9B 4.0 7.2 55.6% 1 31.54s 2,410 10,913
GLM 4.7 Flash 4.0 4.5 55.6% 2 27.09s 1,085 5,597
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-9B 10.0 10.0 0.0% 0 0ms 0 0
GLM 4.7 Flash 10.0 2.1 33.3% 1 65.57s 2,585 20,648
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-9B 5.0 5.6 33.3% 1 87.31s 1,383 32,113
GLM 4.7 Flash 5.0 10.0 50.0% 0 1.51s 584 2,755
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-9B 10.0 7.2 22.2% 1 137.75s 11,549 48,475
GLM 4.7 Flash 10.0 4.4 33.3% 2 174.55s 33,000 25,394
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-9B 10.0 1.6 33.3% 1 226.38s 0 30,695
GLM 4.7 Flash 10.0 9.7 0.0% 0 18.14s 18 2,138
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-9B 5.5 5.8 66.7% 1 17.15s 599 4,517
GLM 4.7 Flash 5.0 5.8 66.7% 1 2.97s 388 2,181
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-9B 10.0 10.0 0.0% 0 33.38s 1,545 11,844
GLM 4.7 Flash 10.0 7.2 11.1% 1 12.90s 798 5,225
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-9B 10.0 10.0 100.0% 0 4.31s 444 1,149
GLM 4.7 Flash 10.0 10.0 100.0% 0 15.95s 224 1,014

Quick Compare

Switch Comparison Pair