Clasament al eșecurilor pentru Răspuns greșit

Vezi ce modele AI se lovesc cel mai des de Răspuns greșit, ca să identifici riscurile de fiabilitate înainte să alegi.

Modele afișate

Eșecuri totale

1558

Modelul cel mai afectat

Categorii

În categoria Specific domeniului412 În categoria Trucuri anti-AI293 În categoria Programare252 În categoria Rezolvare de puzzle-uri201 În categoria Cultură generală168 În categoria Combinat68 În categoria Respectarea instrucțiunilor61 În categoria Inteligență generală59 În categoria Parsare și extragere de date41 În categoria Apelare instrumente3

209/209

Rang	Model	Companie	Număr de Răspuns greșit	Scor	Cost total	Teste corecte	Timp de răspuns (mediu)
#145	GLM 5V Turbo none	Z.ai	11	5.6	$0.052	8/21	2.99s
Total teste 21 Teste greșite 13 Cost total $0.052 Timp de răspuns (mediu) 2.99s
#147	Mimo V2 PRO none	Xiaomi	11	5.6	$0.045	7/21	2.27s
Total teste 21 Teste greșite 14 Cost total $0.045 Timp de răspuns (mediu) 2.27s
#149	KAT-Coder-Air V2.5 medium	Kwaipilot	11	5.6	$0.048	8/22	8.42s
Total teste 22 Teste greșite 14 Cost total $0.048 Timp de răspuns (mediu) 8.42s
#152	Qwen3.6 27B none	Qwen	11	5.5	$0.087	7/22	10.7s
Total teste 22 Teste greșite 15 Cost total $0.087 Timp de răspuns (mediu) 10.7s
#154	MiMo-V2.5-Pro none	Xiaomi	11	5.5	$0.068	6/22	4.12s
Total teste 22 Teste greșite 16 Cost total $0.068 Timp de răspuns (mediu) 4.12s
#62	KAT-Coder-Pro V2.5 low	Kwaipilot	10	7.4	$0.387	11/22	19.5s
Total teste 22 Teste greșite 11 Cost total $0.387 Timp de răspuns (mediu) 19.5s
#69	KAT-Coder-Pro V2.5 high	Kwaipilot	10	7.2	$0.482	11/22	20.8s
Total teste 22 Teste greșite 11 Cost total $0.482 Timp de răspuns (mediu) 20.8s
#71	Qwen3.7 Plus none	Qwen	10	7.2	$0.106	11/22	12.1s
Total teste 22 Teste greșite 11 Cost total $0.106 Timp de răspuns (mediu) 12.1s
#83	GPT-5.6 Sol none	OpenAI	10	6.9	$0.524	11/22	2.16s
Total teste 22 Teste greșite 11 Cost total $0.524 Timp de răspuns (mediu) 2.16s
#92	KAT-Coder-Pro V2.5 none	Kwaipilot	10	6.7	$0.476	11/22	25.6s
Total teste 22 Teste greșite 11 Cost total $0.476 Timp de răspuns (mediu) 25.6s
#98	Qwen3.6 Max Preview none	Qwen	10	6.6	$0.231	12/22	7.82s
Total teste 22 Teste greșite 10 Cost total $0.231 Timp de răspuns (mediu) 7.82s
#117	GPT-5.6 Luna low	OpenAI	10	6.2	$0.249	10/22	5.04s
Total teste 22 Teste greșite 12 Cost total $0.249 Timp de răspuns (mediu) 5.04s
#146	Owl Alpha medium	Openrouter	10	5.6	$0.000	8/21	11.9s
Total teste 21 Teste greșite 13 Cost total $0.000 Timp de răspuns (mediu) 11.9s
#148	Owl Alpha none	Openrouter	10	5.6	$0.000	7/21	9.88s
Total teste 21 Teste greșite 14 Cost total $0.000 Timp de răspuns (mediu) 9.88s
#156	Gemma 4 26B A4B none	Google	10	5.5	$0.015	8/22	7.64s
Total teste 22 Teste greșite 14 Cost total $0.015 Timp de răspuns (mediu) 7.64s

Eșecuri Răspuns greșit

Filtrează modelele

Top modele după Număr de Răspuns greșit

Număr de Răspuns greșit vs Scor

Top modele după Timp de răspuns (mediu)