| # | Model | Lab | Overall | Tier | 95% CI | Output price $/1M | Instructions | Truthfulness | Reasoning | Character | Expression | Spanish | Consistency |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 claude-fable-5-1 |
Anthropic | 97.3% | S | 95.7–98.8 | 50 | 98.3% | 100% | 91.7% | 100% | 98.4% | 95.3% | ±2.1 |
| 1 | Claude Opus 5 claude-opus-5 |
Anthropic | 97.1% | S | 95.4–98.7 | 25 | 94.4% | 99.2% | 100% | 99.4% | 98.1% | 91.7% | ±3.3 |
| 3 | Claude Fable 5 claude-fable-5 |
Anthropic | 94.3% | S | 92.5–95.9 | 50 | 93.1% | 95% | 95.8% | 100% | 82.9% | 99% | ±3.4 |
| 4 | Kimi K3 accounts/fireworks/models/kimi-k3 |
Moonshot | 89.2% | A | 86.5–91.7 | 15 | 84.9% | 97.5% | 77.1% | 96.4% | 94.8% | 84.4% | ±6.2 |
| 4 | GPT-5.6 Sol gpt-5.6-sol |
OpenAI | 87.2% | A | 85.2–89.2 | 20 | 100% | 93.3% | 83.3% | 75.6% | 86.6% | 84.4% | ±6 |
| 6 | Grok 4.6 grok-4.6 |
xAI | 72.1% | B | 70.2–73.9 | 6 | 65.3% | 85.8% | 83.3% | 60.7% | 64.4% | 72.9% | ±7 |
| 6 | DeepSeek V4 Flash deepseek-v4-flash |
DeepSeek | 72.1% | B | 69.1–75.4 | 1.3 | 71.4% | 80% | 71.9% | 30.4% | 96.2% | 82.8% | ±13 |
| 6 | Gemini 3.1 Pro gemini-3.1-pro-preview |
71.2% | B | 68.4–73.8 | 12 | 66% | 66.7% | 79.2% | 42.9% | 86.1% | 86.5% | ±7.5 | |
| 9 | DeepSeek V4 Pro deepseek-v4-pro |
DeepSeek | 67.3% | C | 63.2–71.3 | 4 | 82.3% | 61.3% | 72.9% | 32.6% | 81.3% | 73.4% | ±13.5 |