Independent benchmark · Spanish

The ranking of AI models in Spanish

9 frontier models, 25 real-world tests, 6 families. Pre-registered methodology, automated judges with recusal by lab, and every raw response published. Every cell in the table opens and can be read.
Models
9
Tests
25
Runs
4
Responses
900
1.0

Tiers

Sorted by overall score
ranking.sozpic.com
S
≥ 90
A
Claude Fable 5.1
Anthropic
97.3%
A
Claude Opus 5
Anthropic
97.1%
A
Claude Fable 5
Anthropic
94.3%
A
80–89
M
Kimi K3
Moonshot
89.2%
O
GPT-5.6 Sol
OpenAI
87.2%
B
70–79
x
Grok 4.6
xAI
72.1%
D
DeepSeek V4 Flash
DeepSeek
72.1%
G
Gemini 3.1 Pro
Google
71.2%
C
60–69
D
DeepSeek V4 Pro
DeepSeek
67.3%
D
< 60
S ≥ 90 · A 80–89 · B 70–79 · C 60–69 · D < 60
2.0

Table

No empty cells. Every one of them opens.
Overall score: the unweighted mean of the 6 families.
Ranking IA — overall and per-family scores of 9 frontier models in Spanish. Run 2026-09-01_7e3c679, data as of September 1, 2026. Scores as a percentage of the rubric maximum.
#ModelLabOverallTier95% CIOutput price $/1MInstructionsTruthfulnessReasoningCharacterExpressionSpanishConsistency
1 Claude Fable 5.1
claude-fable-5-1
Anthropic 97.3% S 95.7–98.8 50 98.3%100%91.7%100%98.4%95.3% ±2.1
1 Claude Opus 5
claude-opus-5
Anthropic 97.1% S 95.4–98.7 25 94.4%99.2%100%99.4%98.1%91.7% ±3.3
3 Claude Fable 5
claude-fable-5
Anthropic 94.3% S 92.5–95.9 50 93.1%95%95.8%100%82.9%99% ±3.4
4 Kimi K3
accounts/fireworks/models/kimi-k3
Moonshot 89.2% A 86.5–91.7 15 84.9%97.5%77.1%96.4%94.8%84.4% ±6.2
4 GPT-5.6 Sol
gpt-5.6-sol
OpenAI 87.2% A 85.2–89.2 20 100%93.3%83.3%75.6%86.6%84.4% ±6
6 Grok 4.6
grok-4.6
xAI 72.1% B 70.2–73.9 6 65.3%85.8%83.3%60.7%64.4%72.9% ±7
6 DeepSeek V4 Flash
deepseek-v4-flash
DeepSeek 72.1% B 69.1–75.4 1.3 71.4%80%71.9%30.4%96.2%82.8% ±13
6 Gemini 3.1 Pro
gemini-3.1-pro-preview
Google 71.2% B 68.4–73.8 12 66%66.7%79.2%42.9%86.1%86.5% ±7.5
9 DeepSeek V4 Pro
deepseek-v4-pro
DeepSeek 67.3% C 63.2–71.3 4 82.3%61.3%72.9%32.6%81.3%73.4% ±13.5
≥ 75 40 – 75 < 40Consistency: deviation across the runs of each test, in percentage points.
3.0

Head to head

Tests won by family, and the same test side by side.
VS
4.0

Score versus cost

Each dot is a model: top left means a higher score for less money.
5.0

Model pages

Each model, test by test, on its own page.