Foutenranglijst voor Verkeerd antwoord

Zie welke AI-modellen het vaakst tegen Verkeerd antwoord aanlopen, zodat je betrouwbaarheidsrisico's ziet voordat je kiest. Sorteren op: Responstijd (gem.) ↑.

Getoonde modellen

Totaal fouten

1558

Meest getroffen model

Nemotron 3 Nano Omni 30b A3b Reasoning 9

Categorieën

In categorie Domeinspecifiek412 In categorie Anti-AI-trucs293 In categorie Programmeren252 In categorie Puzzeloplossing201 In categorie Algemene kennis168 In categorie Gecombineerd68 In categorie Instructies opvolgen61 In categorie Algemene intelligentie59 In categorie Gegevensparsering en extractie41 In categorie Toolaanroepen3

209/209

Rang	Model	Bedrijf	Verkeerd antwoord-aantal	Score	Totale kosten	Correcte tests	Responstijd (gem.)
#208	Nemotron 3 Nano Omni 30b A3b Reasoning none	NVIDIA	9	3.2	$0.000	2/19	728ms
Totaal tests 19 Foute tests 17 Totale kosten $0.000 Responstijd (gem.) 728ms
#210	LFM2-24B-A2B none	Liquid	9	2.2	$0.001	2/16	782ms
Totaal tests 16 Foute tests 14 Totale kosten $0.001 Responstijd (gem.) 782ms
#205	Laguna Xs.2 none	Poolside	8	3.8	$0.004	5/19	806ms
Totaal tests 19 Foute tests 14 Totale kosten $0.004 Responstijd (gem.) 806ms
#189	Mercury 2 none	Inception	17	4.6	$0.030	4/22	829ms
Totaal tests 22 Foute tests 18 Totale kosten $0.030 Responstijd (gem.) 829ms
#197	Grok 4.20 none	X AI	10	4.1	$0.057	6/18	1.11s
Totaal tests 18 Foute tests 12 Totale kosten $0.057 Responstijd (gem.) 1.11s
#191	Grok 4.20 Beta none	X AI	10	4.4	$0.087	6/18	1.19s
Totaal tests 18 Foute tests 12 Totale kosten $0.087 Responstijd (gem.) 1.19s
#165	Mistral Small 4 none	Mistral	16	5.1	$0.022	5/22	1.20s
Totaal tests 22 Foute tests 17 Totale kosten $0.022 Responstijd (gem.) 1.20s
#193	Elephant Alpha none	Openrouter	9	4.3	$0.000	5/21	1.22s
Totaal tests 21 Foute tests 16 Totale kosten $0.000 Responstijd (gem.) 1.22s
#195	Elephant Alpha medium	Openrouter	9	4.3	$0.000	6/21	1.27s
Totaal tests 21 Foute tests 15 Totale kosten $0.000 Responstijd (gem.) 1.27s
#201	Granite 4.1 8B none	IBM Granite	13	4.0	$0.007	2/22	1.45s
Totaal tests 22 Foute tests 20 Totale kosten $0.007 Responstijd (gem.) 1.45s
#159	GPT-5.6 Luna none	OpenAI	14	5.4	$0.142	6/22	1.50s
Totaal tests 22 Foute tests 16 Totale kosten $0.142 Responstijd (gem.) 1.50s
#136	GPT-5.4 Mini none	OpenAI	13	5.9	$0.095	6/22	1.53s
Totaal tests 22 Foute tests 16 Totale kosten $0.095 Responstijd (gem.) 1.53s
#160	Laguna XS 2.1 none	Poolside	14	5.3	$0.008	5/22	1.55s
Totaal tests 22 Foute tests 17 Totale kosten $0.008 Responstijd (gem.) 1.55s
#106	Gemini 3.1 Flash Lite Preview none	Google	7	6.4	$0.052	12/22	1.58s
Totaal tests 22 Foute tests 10 Totale kosten $0.052 Responstijd (gem.) 1.58s
#203	Grok 4.1 Fast none	X AI	13	3.8	$0.008	3/19	1.62s
Totaal tests 19 Foute tests 16 Totale kosten $0.008 Responstijd (gem.) 1.62s

Verkeerd antwoord-fouten

Modellen filteren

Topmodellen op Verkeerd antwoord-aantal

Verkeerd antwoord-aantal vs Score

Topmodellen op Responstijd (gem.)