AI BENCHY
Advertise here

AI BENCHY Failures

API error Failures

See which AI models run into API error most often, so you can spot reliability risks before choosing one. Sort by: Tests Correct ↑.

Models Shown

9

Total Failures

144

Rank Model Company API error Count Score Tests Correct Response Time (avg)
#64 MiMo-V2-Flash medium Xiaomi 1 7.2 12/21 20.1s
#41 Nemotron 3 Ultra 550b A55b medium NVIDIA 1 7.5 13/21 15.1s
#46 Qwen3.6 35B A3B medium Qwen 2 7.4 13/21 18.1s
#25 Qwen3.5 Plus 2026-02-15 medium Qwen 1 7.9 14/21 73.8s
#26 Qwen3.6 Plus medium Qwen 1 7.9 14/21 30.7s
#27 Gemma 4 31B medium Google 2 7.8 14/21 56.5s
#33 Hy3 preview medium Tencent 3 7.7 14/21 16.3s
#35 Gemini 3 PRO Preview medium Google 4 7.6 14/21 9.05s
#20 Gemini 3.5 Flash none Google 3 8.1 15/21 9.93s

Top Models by API error Count

API error Count vs Score

Top Models by Response Time (avg)