Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

IBM: Granite 4.1 8B vs Z.ai: GLM 4.7 Flash

Summary

Granite 4.1 8B vs GLM 4.7 Flash benchmark comparison: GLM 4.7 Flash leads on average score with 4.3 vs 4.0. Granite 4.1 8B has the lower benchmark cost at $0.003 vs $0.054. Granite 4.1 8B is faster at 728ms vs 35.10s, with pass rates of 9.5% vs 33.3%.

Recommended model: Granite 4.1 8B - Its score stays close to the best score here (4.0 vs 4.3), while costing about 20.5x less than GLM 4.7 Flash.

Last updated at: 2026-06-12

Metric Granite 4.1 8B Granite 4.1 8B none Release: 2026-05-01 GLM 4.7 Flash GLM 4.7 Flash medium Release: 2026-01-19
Score 4.0 4.3
Rank #163 #159
Reliability 10.0 6.7
Consistency 10.0 6.8
Tests Correct
Attempt pass rate 9.5% 33.3%
Flaky tests 0 8
Total Runs 63 63
Cost per result 0.131 1.337
Total Cost $0.003 $0.054
Input Price $0.050 / 1M $0.060 / 1M
Output Price $0.100 / 1M $0.400 / 1M
Total Input Tokens 46,285 37,206
Output Tokens 2,911 43,754
Reasoning Tokens 0 89,079
Response Time (avg) 728ms 35.10s
Response Time (max) 2.17s 174.55s
Response Time (total) 15.29s 456.24s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#163 IBM: Granite 4.1 8B

none
Cost
$0.001
Time
3.2s
Tokens
491 tok

#159 GLM 4.7 Flash

medium
Invalid SVG
Cost
$0.000
Time
186.2s
Tokens
12,112 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Granite 4.1 8B 4.9 10.0 25.0% 0 844ms 645 903 0
GLM 4.7 Flash 4.7 5.9 41.7% 2 14.95s 555 1,122 6,110
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Granite 4.1 8B 4.5 10.0 0.0% 0 775ms 8,344 525 0
GLM 4.7 Flash 3.2 7.4 11.1% 1 55.33s 3,106 4,981 22,387
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Granite 4.1 8B 3.0 10.0 0.0% 0 1.88s 19,089 396 0
GLM 4.7 Flash 2.8 2.1 33.3% 1 65.57s 17,185 2,585 20,648
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Granite 4.1 8B 3.0 10.0 0.0% 0 575ms 7,617 195 0
GLM 4.7 Flash 6.3 10.0 50.0% 0 1.51s 7,107 584 2,755
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Granite 4.1 8B 3.0 10.0 0.0% 0 357ms 768 24 0
GLM 4.7 Flash 3.5 4.4 33.3% 2 174.55s 643 33,000 25,394
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Granite 4.1 8B 4.0 10.0 0.0% 0 499ms 528 115 0
GLM 4.7 Flash 3.6 9.7 0.0% 0 18.14s 318 18 2,138
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Granite 4.1 8B 3.6 9.9 0.0% 0 344ms 687 66 0
GLM 4.7 Flash 6.2 5.8 66.7% 1 2.97s 636 388 2,181
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Granite 4.1 8B 3.2 10.0 0.0% 0 608ms 672 432 0
GLM 4.7 Flash 2.9 7.2 11.1% 1 12.93s 521 781 5,255
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Granite 4.1 8B 10.0 10.0 100.0% 0 2.17s 7,719 243 0
GLM 4.7 Flash 10.0 10.0 100.0% 0 15.95s 6,949 224 1,014
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Granite 4.1 8B 3.0 10.0 0.0% 0 306ms 216 12 0
GLM 4.7 Flash 3.0 10.0 0.0% 0 11.13s 186 71 1,197

Quick Compare

Switch Comparison Pair