Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Anthropic: Claude Sonnet 4.6 vs Qwen: Qwen3.5-35B-A3B

Last updated at: 2026-05-29

Metric Claude Sonnet 4.6 Claude Sonnet 4.6 none Release: 2026-02-17 Qwen3.5-35B-A3B Qwen3.5-35B-A3B medium Release: 2026-02-24
Score 7.0 7.3
Rank #78 #68
Reliability 10.0 10.0
Consistency 9.7 7.5
Tests Correct
Attempt pass rate 58.3% 73.3%
Flaky tests 1 6
Total Runs 60 60
Cost per result 2.782 4.865
Total Cost $0.306 $0.536
Input Price $3.000 / 1M $0.139 / 1M
Output Price $15.000 / 1M $1.000 / 1M
Output Tokens 9,450 31,242
Reasoning Tokens 0 330,546
Response Time (avg) 5.27s 69.66s
Response Time (max) 23.84s 409.98s
Response Time (total) 68.50s 1393.17s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 4.8 10.0 25.0% 0 2.94s 1,214 0
Qwen3.5-35B-A3B 10.0 10.0 100.0% 0 21.13s 798 42,652
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 6.8 10.0 50.0% 0 6.73s 2,112 0
Qwen3.5-35B-A3B 6.5 10.0 50.0% 0 244.54s 14,456 88,431
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 9.5 10.0 100.0% 0 23.84s 3,766 0
Qwen3.5-35B-A3B 4.7 1.6 66.7% 1 75.34s 775 12,485
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 10.0 10.0 100.0% 0 3.43s 252 0
Qwen3.5-35B-A3B 7.3 5.9 83.3% 1 59.33s 235 19,493
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 7.7 10.0 66.7% 0 3.54s 413 0
Qwen3.5-35B-A3B 4.1 4.4 44.5% 2 88.34s 41 46,368
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 6.1 3.1 66.7% 1 2.56s 192 0
Qwen3.5-35B-A3B 2.8 1.6 33.3% 1 30.30s 20 3,753
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 6.5 10.0 50.0% 0 1.96s 90 0
Qwen3.5-35B-A3B 10.0 10.0 100.0% 0 24.45s 97 17,361
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 7.7 10.0 66.7% 0 2.53s 533 0
Qwen3.5-35B-A3B 8.2 7.2 88.9% 1 33.13s 3,592 26,585
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 10.0 10.0 100.0% 0 4.11s 447 0
Qwen3.5-35B-A3B 10.0 10.0 100.0% 0 4.65s 309 1,365
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 3.0 10.0 0.0% 0 4.67s 431 0
Qwen3.5-35B-A3B 3.0 10.0 0.0% 0 177.35s 10,919 72,053

Quick Compare

Switch Comparison Pair