Anti-AI Tricks: Wrong answer
Anti-AI Tricks
Wrong answer
See which AI models are most likely to hit Wrong answer on Anti-AI Tricks, so you can spot weak points faster. Sort by: Total Cost ↑.
Failure Reasons
140/140
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #18 | GPT-5.4 medium | OpenAI | 1 | 8.3 | $1.533 | 3/4 | 4.11s |
| #27 | Muse Spark 1.1 high | Meta | 1 | 7.5 | $1.694 | 2/4 | 8.60s |
| #143 | Gemini 3.1 Flash Lite high | 1 | 8.7 | $2.044 | 3/4 | 37.2s | |
| #40 | Claude Sonnet 4.6 medium | Anthropic | 1 | 6.5 | $2.057 | 2/4 | 2.98s |
| #181 | Grok 4.20 Multi Agent Beta medium | X AI | 1 | 6.9 | $5.599 | 2/4 | 3.46s |