Issue archive · updated 2026-06-10

LLM Leaderboard Archive — 2026-06

Archived 2026-06 LLM leaderboard: category leaders and full rankings for that issue.

01

Text Generation & Reasoning

Text Generation & Reasoning

Current leader Claude Fable 5 Anthropic previously: Claude Opus 4.8

June 9 — Anthropic ships Claude Fable 5, the first publicly available Mythos-class model, and it debuts at #1 on the Artificial Analysis Intelligence Index v4.0 with a score of 65, clearing Opus 4.8 by 4 points. Anthropic now holds the entire reasoning frontier.

#ModelCompanyScoreStrengths
1 Claude Fable 5Released June 9. First public Mythos-class model — debuts #1 on the AA Intelligence Index v4.0 at 65 and scores 53% on Humanity's Last Exam (7+ points clear). Anthropic 65 Intelligence Index v4.0: 65 (#1) · Humanity's Last Exam: 53% · Highest AA-Omniscience score to date · 1M+ context · Adaptive Reasoning · Released 2026-06-09
2 Claude Opus 4.8Released May 28. Adaptive reasoning + max-effort mode; now #2 on the AA Intelligence Index v4.0 behind Fable 5. Anthropic 61.4 Intelligence Index v4.0: 61.4 · Adaptive reasoning mode · Coding + agentic upgrade over 4.7 · Long-running work consistency · Released 2026-05-28
3 GPT-5.5Apr 24 release. 1M-token context, native MCP + Skills + computer use + hosted shell. OpenAI 60.2 Intelligence Index v4.0 (xhigh): 60.2 · 1M token context · Native MCP + Skills · Computer use built-in · Tool search + web search · Released 2026-04-24
4 Gemini 3.1 ProStrong agentic frontier with Antigravity 2.0 platform integration. Google 57 Intelligence Index v4.0: 57 · Agentic platform (Antigravity 2.0) integration · Tied with GPT-5.5 medium and Qwen3.7 Max
5 Grok 4.3Released Apr 30. AA Intelligence Index 53 with always-on reasoning, 1M context, and aggressive pricing — the cheapest of the frontier four. xAI 53 Intelligence Index v4.0: 53 · Always-on reasoning · 1M context · ~40% / 60% cheaper input/output vs Grok 4.20 · Strong agentic + tool use

Released June 9, 2026. AA Intelligence Index v4.0: 65 (Adaptive Reasoning, max effort) — 4 points ahead of Opus 4.8 (61.4) and GPT-5.5 xhigh (60.2); Gemini 3.1 Pro (57) and Grok 4.3 (53) follow.

Fable 5 is the Mythos-class tier above Opus, gated with extra safeguards; Opus 4.8 stays the default workhorse for cost-sensitive high-volume work.

As of 2026-06-10, Claude Fable 5 (Anthropic, released 2026-06-09) leads the Artificial Analysis Intelligence Index v4.0 with a score of 65 in Adaptive-Reasoning max-effort mode, ahead of Claude Opus 4.8 (61.4), GPT-5.5 xhigh (OpenAI, 60.2), Gemini 3.1 Pro (Google, 57), and Grok 4.3 (xAI, 53). Fable 5 is the first publicly released Mythos-class model and posts 53% on Humanity's Last Exam — more than 7 points ahead of the next model.
02

Image Generation

Image Generation

Current leader GPT Image 2 OpenAI previously: GPT Image-2

OpenAI ships GPT Image 2 with token-based pricing and 50% Batch API discount. Recraft V4.1 holds the Artificial Analysis quality leaderboard, while Adobe Firefly enterprise mode remains the rights-cleared default.

#ModelCompanyScoreStrengths
1 GPT Image 2Released April 21. State-of-the-art quality with token pricing and Batch API support. OpenAI None Token-based pricing · Batch API at 50% discount · Flexible image sizes · High-fidelity inputs · Released 2026-04-21
2 Recraft V4.1Leads Artificial Analysis text-to-image arena on raw output quality. Recraft None Top of AA text-to-image quality · Strong control / style transfer · Designer-grade output
3 Adobe Firefly Image 4IP-cleared training data; the enterprise-safe choice for commercial use. Adobe None Trained on licensed assets · Indemnification for enterprise · Native Adobe Creative Cloud integration

Released April 21, 2026. Token-based pricing, Batch API at 50% off.

Recraft V4.1 leads AA's text-to-image arena on raw quality; GPT Image 2 wins on ecosystem and pricing transparency.

As of 2026-06-10, GPT Image 2 (OpenAI, released 2026-04-21) is the most-shipped foundation image model, offering token-based pricing and 50% Batch API discount. Recraft V4.1 leads Artificial Analysis text-to-image quality rankings, and Adobe Firefly remains the IP-cleared default for enterprise.
03

Video Generation

Video Generation

Current leader Seedance 2.0 ByteDance previously: Veo 3.5

Seedance 2.0 (ByteDance) tops the Artificial Analysis text-to-video Arena with audio at ELO 1215 — the first model to make synced audio-visual generation state of the art. Google's Veo 3.5 leads the silent-cinematic tier; Kling 4 holds the value end. OpenAI's Sora exited the field after its app was discontinued on April 26, 2026.

#ModelCompanyScoreStrengths
1 Seedance 2.0Tops the Artificial Analysis text-to-video Arena (with audio) at ELO 1215 — ahead of Veo. Native audio-visual generation: 15-second multi-shot clips with synced sound from text, image, audio and video inputs. ByteDance 1215 AA T2V Arena (w/ audio) #1 · ELO 1215 · 15s multi-shot · synced audio · Multimodal input (text/image/audio/video) · Dual-Branch Diffusion Transformer
2 Veo 3.5Production-grade film output with strongest temporal coherence. Google None 1080p output · Strong physics simulation · Long-shot temporal coherence · Native Gemini API integration
3 Kling 4Dominant in APAC short-form ad creative; fastest iteration cycle in the space. Kuaishou 快手 None 9:16 vertical native · Fastest editorial iteration · TikTok / Douyin native style · Low-latency generation

Seedance 2.0 holds #1 on the with-audio text-to-video Arena; Sora 2 drops out after OpenAI discontinued the Sora app (Apr 26, 2026).

Seedance 2.0 for synced audio-visual and multi-shot narrative; Veo 3.5 for cinematic fidelity; Kling 4 for cost.

As of 2026-06-10, Seedance 2.0 (ByteDance) leads text-to-video generation, ranking #1 on the Artificial Analysis text-to-video Arena (with audio) at ELO 1215 — ahead of Google Veo. It generates 15-second multi-shot clips with natively synced audio from text, image, audio and video inputs. Veo 3.5 (Google) leads silent cinematic quality and Kling 4 (Kuaishou) anchors the value tier; OpenAI's Sora left the ranking after the Sora app was discontinued on April 26, 2026.
04

Code Generation & Agentic Coding

Code Generation & Agentic Coding

Current leader Claude Fable 5 Anthropic previously: Claude Opus 4.8

Claude Fable 5 launches June 9 and immediately tops SWE-bench Pro at 80.3% — about 11 points clear of Opus 4.8 (69.2%) on the same uncontaminated multi-language benchmark. Pro scores run far below the contaminated Verified set, so treat 80% as the honest new frontier.

#ModelCompanyScoreStrengths
1 Claude Fable 5#1 on SWE-bench Pro at 80.3% (June 9 launch) — 11 points ahead of the field. Mythos-class agentic coding for the hardest multi-file, long-horizon work. Anthropic 80.3 SWE-bench Pro: 80.3% (#1) · 11 points clear of Opus 4.8 · Long-horizon agentic edits · 1M+ context · Mythos-class · Released 2026-06-09
2 Claude Opus 4.8#2 on SWE-bench Pro at 69.2% (updated June 8), behind Fable 5. The cost-sensitive default for multi-file refactoring, code review, and long-horizon agentic edits. Anthropic 69.2 SWE-bench Pro: 69.2% (#2) · Multi-file refactor SOTA · Code review top · Vision-aware coding
3 GPT-5.3 Codex (xhigh)#3 on SWE-bench Pro (June 8). Specialized terminal-code agent; AA Intelligence Index 54. OpenAI 54 AA Intelligence Index: 54 · Sandboxed shell built-in · Strong agentic loops · Terminal-Bench best
4 Cursor Composer 2.5Ranks #3 on AA Coding Agent Index. IDE-native pair programming with multi-file context. Cursor None AA Coding Agent Index: #3 · IDE-native context · Multi-file edits · Inline diff workflow

Fable 5 debuts June 9 at #1 on SWE-bench Pro (80.3%); Opus 4.8 (69.2%) #2, GPT-5.3 Codex #3.

Fable 5 for the hardest multi-file refactors and long-horizon agents; Opus 4.8 as the cost-sensitive default; GPT-5.3 Codex for sandboxed terminal agents; Cursor Composer 2.5 for IDE-native pair programming.

As of 2026-06-10, Claude Fable 5 (Anthropic, released 2026-06-09) leads agentic coding, ranking #1 on the SWE-bench Pro leaderboard at 80.3% — about 11 points ahead of Claude Opus 4.8 (69.2%) and GPT-5.3 Codex (OpenAI). SWE-bench Pro uses 1,865 uncontaminated multi-language tasks, so scores sit well below the Python-only Verified set. Cursor Composer 2.5 ranks among the top IDE-native coding agents on AA's Coding Agent Index.
05

Voice / Speech

Voice / Speech

Current leader Realtime 2 OpenAI previously: ElevenLabs v3

OpenAI's Realtime 2 (May 7) brings configurable-reasoning speech-to-speech to general availability; AA's text-to-speech crown goes to Fun-Realtime-TTS, and MAI-Transcribe-1.5 wins STT on accuracy-speed.

#ModelCompanyScoreStrengths
1 Realtime 2GA on May 7. Configurable-reasoning speech-to-speech with realtime translate + Whisper variants. OpenAI None Configurable reasoning · Speech-to-speech agents · Streaming translate variant · Streaming STT variant · Released 2026-05-07
2 ElevenLabs v3Industry default for character voice cloning and audiobook production. ElevenLabs None Character voice cloning SOTA · 100+ languages · Long-form audiobook quality · Emotion control
3 Fun-Realtime-TTSTops AA text-to-speech leaderboard on quality metrics. Fun (Alibaba DAMO) None AA TTS leaderboard #1 · Sub-200ms latency · Multi-speaker streaming · Strong CJK

Realtime 2 family shipped May 7 (gpt-realtime-2 / -translate / -whisper).

Realtime 2 for agentic voice; ElevenLabs v3 for character voice cloning; Fun-Realtime-TTS for raw TTS quality; MAI-Transcribe-1.5 for transcription.

As of 2026-06-10, OpenAI Realtime 2 (released 2026-05-07) is the leading speech-to-speech model for agentic voice with configurable reasoning. Fun-Realtime-TTS tops the Artificial Analysis text-to-speech leaderboard; MAI-Transcribe-1.5 leads accuracy-speed on speech-to-text. ElevenLabs v3 remains dominant for character voice cloning.
06

Music Generation

Music Generation

Current leader Suno v6 Suno previously: Suno v5.5

Suno v6 widens the gap on full-song coherence and lyric prosody; Udio v3 keeps pushing studio-grade mixing; Lyria (Google) integrates into Gemini Omni for any-to-music workflows.

#ModelCompanyScoreStrengths
1 Suno v6Rolling release expected late June. Best full-song coherence and lyric prosody. Suno None Full-song coherence SOTA · Lyric prosody best · Multilingual vocal · Style transfer
2 Udio v3Studio-grade mixing with stem-level output for producers. Udio None Stem-level export · Studio-grade mixing · Strong electronic genres · DAW-friendly workflow
3 Lyria (via Gemini Omni)Folded into Gemini Omni for any-to-music + cross-modal generation. Google None Gemini Omni native · Cross-modal generation · Image / video → music workflows

Suno v6 expected late June; v5.5 remains the deployed default.

Suno v6 for full-song generation; Udio v3 for mixed studio-grade stems; Lyria via Gemini Omni for cross-modal generation.

As of 2026-06-10, Suno v6 (rolling release expected late June) leads full-song generation with strong lyric prosody; v5.5 remains the deployed default. Udio v3 leads on studio-grade mixing and stem-level output. Lyria from Google is integrated into Gemini Omni for cross-modal (image / video / music) workflows.
07

Vision / Multimodal Understanding

Vision / Multimodal Understanding

Current leader Claude Opus 4.7-thinking Anthropic previously: GPT-4o Vision

Anthropic sweeps LMArena Vision top-3 with Opus 4.7-thinking, 4.6-thinking, and 4.7. Opus 4.8 is too freshly released to appear on Arena ELO but is expected to consolidate the lead by end of June.

#ModelCompanyScoreStrengths
1 Claude Opus 4.7-thinkingTops LMArena Vision. Best OCR + chart + document understanding. Anthropic 1309 LMArena Vision ELO: 1309 · OCR SOTA · Chart understanding · Document Q&A
2 GPT-5.5Strongest image-grounded reasoning chains; native computer-use vision pipeline. OpenAI None Image-grounded reasoning best · Computer use vision · 1M token multimodal · Released 2026-04-24
3 Gemini 3.1 ProBest video understanding and long-form temporal reasoning. Google None Video understanding SOTA · Long-form temporal · Robotics-ER 1.6 integration · Multimodal context 2M+

LMArena Vision top-3 all Anthropic (1309 / 1303 / 1298 ELO).

Anthropic for OCR + document Q&A + chart understanding; GPT-5.5 for image-grounded reasoning; Gemini 3.1 Pro for video understanding.

As of 2026-06-10, Claude Opus 4.7-thinking (Anthropic, LMArena Vision ELO 1309) leads vision and multimodal understanding, followed by Opus 4.6-thinking (1303) and Opus 4.7 (1298) — Anthropic sweeps the Vision Arena top-3. Opus 4.8 is too newly released to appear on Arena ELO. GPT-5.5 leads image-grounded reasoning chains; Gemini 3.1 Pro leads video understanding.
08

Open-Source / Open-Weights

Open-Source / Open-Weights

Current leader Kimi K2.6 Moonshot AI previously: Llama 4

Kimi K2.6 (Moonshot) leads open weights at AA Intelligence Index 54 — within 7 points of frontier closed models. DeepSeek V4 Pro (MIT, 52) is the #2 open reasoning model, and Google's new Gemma 4 12B (Apache 2.0, 2026-06-03) packs native multimodal into a 16GB-laptop footprint. The closed-vs-open gap is the narrowest it has ever been.

#ModelCompanyScoreStrengths
1 Kimi K2.6Open-weights leader on AA Intelligence Index. Closes the closed-source gap to 7 points. Moonshot AI 54 AA Intelligence Index: 54 · Open weights · Strong Chinese + English · Long-context retention
2 DeepSeek V4 ProMIT-licensed open weights at AA Intelligence Index 52 — #3 of 89 overall and the #2 open reasoning model behind only Kimi K2.6. DeepSeek 52 AA Intelligence Index: 52 (#3/89) · MIT license · open weights · MoE 1.6T total / 49B active · 1M-token context
3 Gemma 4 12BReleased 2026-06-03 under Apache 2.0. Encoder-free native multimodal (text/image/audio/video), 256K context, runs on a 16GB laptop — performance nearing last-gen 27B. Google Apache 2.0 · open weights · 256K context · native multimodal · Runs on 16GB VRAM · MMLU-Pro 77.2 · GPQA-Diamond 78.8
4 Qwen3.7 PlusBest open-source for Chinese-language self-host deployment. Alibaba 53 AA Intelligence Index: 53 · Best Chinese open-source · Strong tool use · Open weights

Kimi K2.6 (54) tops open weights; DeepSeek V4 Pro (52) #2; Gemma 4 12B brings 16GB-laptop multimodal.

Kimi K2.6 for general-purpose open deployment; DeepSeek V4 Pro for cheap frontier-grade reasoning; Gemma 4 12B for on-device multimodal; Qwen3.7 Plus for Chinese self-host.

As of 2026-06-10, Kimi K2.6 (Moonshot AI) leads open-weights models with an Artificial Analysis Intelligence Index of 54, 7 points behind the closed-source frontier (Claude Opus 4.8, 61). DeepSeek V4 Pro (DeepSeek, MIT license) ranks second among open models at Intelligence Index 52 (#3 of 89 overall). Google released Gemma 4 12B on 2026-06-03 under Apache 2.0 — an encoder-free multimodal model with 256K context that runs on 16GB of memory. Qwen3.7 Plus (Alibaba, 53) rounds out the Chinese open-source top tier.
09

Cost-Effectiveness / Value

Intelligence per Dollar

Current leader DeepSeek V4 Flash DeepSeek

The 2026 value war is led by China's open-source camp. DeepSeek V4 Flash delivers near-flagship intelligence (Index 47) at roughly a tenth of the price — a blended cost near $0.06 per 1M tokens. The leaderboard makes one warning explicit: 'Flash' and 'mini' branding does not mean cheap. Gemini 3.5 Flash scores a strong 55 but costs over 20× more per full Intelligence-Index run.

#ModelCompanyScoreStrengths
1 DeepSeek V4 FlashThe intelligence-per-dollar king. AA Intelligence Index 47 at $0.14/$0.28 per 1M tokens — about a tenth of comparable Flash flagships, with the lowest cache-hit price of any 2026 frontier model. DeepSeek 47 AA Intelligence Index: 47 · $0.14 in / $0.28 out per 1M · Blended ≈ $0.06 / 1M · MIT open weights · 1M context
2 DeepSeek V4 ProBest balance in the high-intelligence + low-price quadrant. Intelligence Index 52 (top-3 overall) at $0.435/$0.87 — a fraction of same-tier flagships like GPT-5.5 and Claude Opus. DeepSeek 52 AA Intelligence Index: 52 (#3/89) · $0.435 in / $0.87 out per 1M · Top-3 intelligence overall · MIT open weights
3 Qwen3.7 PlusKeeps the value race from being a single-vendor story. Intelligence Index 53 at $0.40 input — a higher score than V4 Pro; the trade-off is pricier output and slower generation. Alibaba 53 AA Intelligence Index: 53 · $0.40 in / $1.16 out per 1M · Highest score in the value tier · Strong Chinese + tool use

New category. DeepSeek V4 Flash leads intelligence-per-dollar; open-source models sweep the value tier.

DeepSeek V4 Flash for highest-volume low-cost workloads; V4 Pro when you need stronger reasoning but still want to save; Qwen3.7 Plus to diversify away from a single vendor.

As of 2026-06-10, DeepSeek V4 Flash (DeepSeek) leads cost-effectiveness with an Artificial Analysis Intelligence Index of 47 at $0.14 input / $0.28 output per 1M tokens — a blended cost near $0.06 per 1M, roughly one-tenth of comparable Flash-tier flagships. DeepSeek V4 Pro (Index 52, $0.435/$0.87) and Qwen3.7 Plus (Alibaba, Index 53, $0.40/$1.16) complete the value top tier. Note that 'Flash'/'mini' branding does not imply low cost: Gemini 3.5 Flash scores 55 but costs over 20× more per full Intelligence-Index run.

This month's trends

Sources & methodology

Data as of 2026-06-10

Monthly archive