# Ranking IA — the ranking of AI models in Spanish > Independent benchmark of 9 frontier models on 25 real professional-use tests in Spanish, grouped in 6 families. As of September 1, 2026, the best-scoring model is Claude Fable 5.1 (Anthropic) with 97.3%. Pre-registered methodology, judges with recusal by lab and full data published. A project by Sozpic. - Web: https://ranking.sozpic.com/en/ - Run: 2026-09-01_7e3c679 · sealed instrument: 7e3c679 · data as of September 1, 2026 - License: CC BY 4.0 — Ranking IA · Sozpic (https://creativecommons.org/licenses/by/4.0/) - How to cite: “Ranking IA (Sozpic), run 2026-09-01_7e3c679, September 1, 2026”, with a link to https://ranking.sozpic.com/en/ - Spanish version (original): https://ranking.sozpic.com/ ## Standings | # | Model | Lab | Overall | Tier | 95% CI | Instructions | Truthfulness | Reasoning | Character | Expression | Spanish | Output price $/1M | |---|---|---|---|---|---|---|---|---|---|---|---|---| | 1 | Claude Fable 5.1 | Anthropic | 97.3% | S | 95.7–98.8 | 98.3% | 100% | 91.7% | 100% | 98.4% | 95.3% | 50 | | 1 | Claude Opus 5 | Anthropic | 97.1% | S | 95.4–98.7 | 94.4% | 99.2% | 100% | 99.4% | 98.1% | 91.7% | 25 | | 3 | Claude Fable 5 | Anthropic | 94.3% | S | 92.5–95.9 | 93.1% | 95% | 95.8% | 100% | 82.9% | 99% | 50 | | 4 | Kimi K3 | Moonshot | 89.2% | A | 86.5–91.7 | 84.9% | 97.5% | 77.1% | 96.4% | 94.8% | 84.4% | 15 | | 4 | GPT-5.6 Sol | OpenAI | 87.2% | A | 85.2–89.2 | 100% | 93.3% | 83.3% | 75.6% | 86.6% | 84.4% | 20 | | 6 | Grok 4.6 | xAI | 72.1% | B | 70.2–73.9 | 65.3% | 85.8% | 83.3% | 60.7% | 64.4% | 72.9% | 6 | | 6 | DeepSeek V4 Flash | DeepSeek | 72.1% | B | 69.1–75.4 | 71.4% | 80% | 71.9% | 30.4% | 96.2% | 82.8% | 1.3 | | 6 | Gemini 3.1 Pro | Google | 71.2% | B | 68.4–73.8 | 66% | 66.7% | 79.2% | 42.9% | 86.1% | 86.5% | 12 | | 9 | DeepSeek V4 Pro | DeepSeek | 67.3% | C | 63.2–71.3 | 82.3% | 61.3% | 72.9% | 32.6% | 81.3% | 73.4% | 4 | Tiers: S ≥ 90 · A 80–89 · B 70–79 · C 60–69 · D < 60 (out of 100). Ties: two models share a position when the 95% confidence interval of the difference between their overall scores contains zero. ## What each family measures - **Instructions**: Does it do exactly what you ask? - **Truthfulness**: Does it make things up? - **Reasoning**: Does it think or recite? - **Character**: Does it tell you the truth even when you won’t like it? - **Expression**: Does it write well, with a voice of its own? - **Spanish**: Does it sound native or translated? ## Methodology in brief - The complete instrument (prompts, rubrics, judges and aggregation rules) is sealed before execution and published inside the data JSON, with identifier 7e3c679. - Each model answers 4 independent runs per test, through its official API, with no system prompt, no tools and the provider’s default parameters. - Three-layer evaluation: deterministic verifiers for mechanical aspects; a single fast judge per run, assigned by hash and never from the lab of the model under evaluation, for closed rubrics; and a panel of frontier judges with recusal for subjective aspects and for every doubtful run. - Judges never see the name of the model under evaluation. - Each family’s score is the mean of its tests; the overall score is the unweighted mean of the 6 families. - Sample audit of this run: 100% agreement over 6 re-scored runs. - Uncertainty: bootstrap over runs per test (resampling with replacement within each test), 1,000 iterations, deterministic (seed derived from the model id). ## Pages - [Full ranking](https://ranking.sozpic.com/en/): standings, tiers, head-to-head comparison and score versus cost. - [Methodology](https://ranking.sozpic.com/en/methodology/): pre-registration, evaluation layers, recusal, audit and judge monitoring. - [Page: Claude Fable 5.1](https://ranking.sozpic.com/en/model/claude-fable-5-1/): 97.3% overall, tier S, Anthropic. Score by family and the judges’ conclusion for each test. - [Page: Claude Opus 5](https://ranking.sozpic.com/en/model/claude-opus-5/): 97.1% overall, tier S, Anthropic. Score by family and the judges’ conclusion for each test. - [Page: Claude Fable 5](https://ranking.sozpic.com/en/model/claude-fable-5/): 94.3% overall, tier S, Anthropic. Score by family and the judges’ conclusion for each test. - [Page: Kimi K3](https://ranking.sozpic.com/en/model/kimi-k3/): 89.2% overall, tier A, Moonshot. Score by family and the judges’ conclusion for each test. - [Page: GPT-5.6 Sol](https://ranking.sozpic.com/en/model/gpt-5-6-sol/): 87.2% overall, tier A, OpenAI. Score by family and the judges’ conclusion for each test. - [Page: Grok 4.6](https://ranking.sozpic.com/en/model/grok-4-6/): 72.1% overall, tier B, xAI. Score by family and the judges’ conclusion for each test. - [Page: DeepSeek V4 Flash](https://ranking.sozpic.com/en/model/deepseek-v4-flash/): 72.1% overall, tier B, DeepSeek. Score by family and the judges’ conclusion for each test. - [Page: Gemini 3.1 Pro](https://ranking.sozpic.com/en/model/gemini-3-1-pro/): 71.2% overall, tier B, Google. Score by family and the judges’ conclusion for each test. - [Page: DeepSeek V4 Pro](https://ranking.sozpic.com/en/model/deepseek-v4-pro/): 67.3% overall, tier C, DeepSeek. Score by family and the judges’ conclusion for each test. ## Data - [leaderboard.json](https://ranking.sozpic.com/data/leaderboard.json): standings, scores by family and by test, conclusions and judge analysis (field names in Spanish). - [llms-full.txt](https://ranking.sozpic.com/en/llms-full.txt): this same content with the full conclusions of every model on every test.