Foutenranglijst voor Verkeerd antwoord

Zie welke AI-modellen het vaakst tegen Verkeerd antwoord aanlopen, zodat je betrouwbaarheidsrisico's ziet voordat je kiest. Sorteren op: Aantal fouten ↑.

Getoonde modellen

Totaal fouten

1642

Meest getroffen model

Gemini 3.6 Flash 1

Categorieën

In categorie Domeinspecifiek433 In categorie Anti-AI-trucs306 In categorie Programmeren266 In categorie Puzzeloplossing214 In categorie Algemene kennis176 In categorie Gecombineerd71 In categorie Algemene intelligentie66 In categorie Instructies opvolgen65 In categorie Gegevensparsering en extractie41 In categorie Toolaanroepen4

219/219

Rang	Model	Bedrijf	Verkeerd antwoord-aantal	Score	Totale kosten	Correcte tests	Responstijd (gem.)
#44	Claude Sonnet 4.6 medium	Anthropic	4	7.8	$2.057	14/22	25.9s
Totaal tests 22 Foute tests 8 Totale kosten $2.057 Responstijd (gem.) 25.9s
#45	Claude Opus 4.8 low	Anthropic	4	7.8	$2.077	16/22	12.7s
Totaal tests 22 Foute tests 6 Totale kosten $2.077 Responstijd (gem.) 12.7s
#53	GLM 5 Turbo medium	Z.ai	4	7.6	$0.323	14/21	23.0s
Totaal tests 21 Foute tests 7 Totale kosten $0.323 Responstijd (gem.) 23.0s
#61	Qwen3.5 Plus 2026-02-15 medium	Qwen	4	7.5	$0.437	14/22	89.2s
Totaal tests 22 Foute tests 8 Totale kosten $0.437 Responstijd (gem.) 89.2s
#62	Qwen3.5-27B medium	Qwen	4	7.4	$1.627	13/22	111.9s
Totaal tests 22 Foute tests 9 Totale kosten $1.627 Responstijd (gem.) 111.9s
#70	Claude Opus 4.8 none	Anthropic	4	7.3	$1.166	13/22	4.91s
Totaal tests 22 Foute tests 9 Totale kosten $1.166 Responstijd (gem.) 4.91s
#78	GLM 5.1 medium	Z.ai	4	7.1	$0.535	13/22	46.8s
Totaal tests 22 Foute tests 9 Totale kosten $0.535 Responstijd (gem.) 46.8s
#84	Seed-2.0-Mini medium	Bytedance Seed	4	7.0	$0.101	11/22	92.5s
Totaal tests 22 Foute tests 11 Totale kosten $0.101 Responstijd (gem.) 92.5s
#94	Qwen3.6 35B A3B medium	Qwen	4	6.7	$0.746	13/22	58.1s
Totaal tests 22 Foute tests 9 Totale kosten $0.746 Responstijd (gem.) 58.1s
#120	Qwen3.5-Flash medium	Qwen	4	6.2	$0.139	12/22	84.8s
Totaal tests 22 Foute tests 10 Totale kosten $0.139 Responstijd (gem.) 84.8s
#136	Step 3.5 Flash medium	Stepfun	4	6.0	$0.108	11/21	174.2s
Totaal tests 21 Foute tests 10 Totale kosten $0.108 Responstijd (gem.) 174.2s
#149	Gemini 3.1 Flash Lite high	Google	4	5.6	$2.044	10/18	62.0s
Totaal tests 18 Foute tests 8 Totale kosten $2.044 Responstijd (gem.) 62.0s
#159	Hy3 preview low	Tencent	4	5.5	$0.015	10/21	24.6s
Totaal tests 21 Foute tests 11 Totale kosten $0.015 Responstijd (gem.) 24.6s
#190	Grok 4.20 Multi Agent Beta medium	X AI	4	4.8	$5.599	8/18	9.69s
Totaal tests 18 Foute tests 10 Totale kosten $5.599 Responstijd (gem.) 9.69s
#193	Hunter Alpha medium	OpenRouter	4	4.7	$0.000	8/18	10.3s
Totaal tests 18 Foute tests 10 Totale kosten $0.000 Responstijd (gem.) 10.3s

Verkeerd antwoord-fouten

Modellen filteren

Topmodellen op Verkeerd antwoord-aantal

Verkeerd antwoord-aantal vs Score

Topmodellen op Responstijd (gem.)