Domain specific: Wrong answer
Domain specific
Wrong answer
See which AI models are most likely to hit Wrong answer on Domain specific, so you can spot weak points faster. Sort by: Total Cost ↑.
Failure Reasons
198/198
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #17 | Claude Fable 5 medium | Anthropic | 2 | 5.3 | $3.478 | 1/3 | 53.4s |
| #10 | GPT-5.5 medium | OpenAI | 2 | 5.3 | $4.137 | 1/3 | 164.1s |
| #181 | Grok 4.20 Multi Agent Beta medium | X AI | 2 | 2.9 | $5.599 | 0/3 | 24.7s |