Issue archive · updated 2026-08-09

LLM Leaderboard Archive — 2026-08

Archived 2026-08 LLM leaderboard: category leaders and full rankings for that issue.

01

Text Generation & Reasoning

Text Generation & Reasoning

Current leader Claude Opus 5 Anthropic

August 9 scrape of Artificial Analysis shows Claude Opus 5 still #1, but the whole frontier band rebased upward: Opus 5 max 63 (July issue 61), Fable 5 with fallback 62, GPT-5.6 Sol max 61, open-weights Kimi K3 max 60. Muse Spark 1.2 (xhigh) enters at 57; Grok 4.5 high sits at 56.

#ModelCompanyScoreStrengths
1 Claude Opus 5Still #1 on AA Intelligence Index — now 63 max, up from 61 in the July issue. Anthropic 63 Intelligence Index: 63 (max) · SWE-bench Verified ~96% · $5 / $25 per 1M tokens · 1M context · Released 2026-07-24
2 Claude Fable 5AA lists Fable 5 with fallback at 62 — #2 overall after the August rebase. Anthropic 62 Intelligence Index: 62 (with fallback) · Mythos-class safeguards · 1M+ context · Adaptive Reasoning · Released 2026-06-09
3 GPT-5.6 SolSol max climbs to 61 on AA; still the OpenAI flagship tier. OpenAI 61 Intelligence Index: 61 (max) · Terminal-Bench 2.1: 88.8% · $5 / $30 per 1M tokens · 1M context · GA 2026-07-09
4 Kimi K3Open-weights K3 max hits 60 — global #4, only 3 points off Opus 5. Moonshot AI 60 Intelligence Index: 60 (open #1) · 2.8T MoE · 1M context · LMArena WebDev #1 (July) · Open weights · $3/$15 per 1M · Released 2026-07-17
5 Muse Spark 1.2New mid-frontier entrant at AA Index 57 (xhigh) — displaces Grok from the July top-5 shape. Muse 57 Intelligence Index: 57 (xhigh) · New vs July top tier · Closed-weights contender
6 Grok 4.5AA high mode at 56; still the cheap always-on reasoning pick at $2/$6. xAI 56 Intelligence Index: 56 (high) · Aggressive pricing: $2 / $6 per 1M · Always-on reasoning · 1M context · Released 2026-07-08

AA Intelligence Index top-4 all +2–3 vs July issue. Muse Spark 1.2 is new in the mid-frontier; DeepSeek V4 Flash 0731 lands at 52.

Opus 5 remains the default frontier workhorse at $5/$25; Fable 5 is the Mythos-class ceiling; Sol is OpenAI’s tiered answer; Kimi K3 is still the open model inside the frontier band.

As of 2026-08-09, Claude Opus 5 (Anthropic) leads the Artificial Analysis Intelligence Index at 63 in max-effort mode, ahead of Claude Fable 5 (62), GPT-5.6 Sol (OpenAI, 61), and open-weights Kimi K3 (Moonshot AI, 60). Opus 5 is priced at $5 input / $25 output per 1M tokens.
02

Image Generation

Image Generation

Current leader GPT Image 2 OpenAI

GPT Image 2 (high) extends its Artificial Analysis Image Arena lead to ELO 1357 (July issue 1338). Reve 2.1 holds #2 at 1314. Google’s Nano Banana 2 (Gemini 3.1 Flash Image Preview) debuts in the top 3 at 1307. Seedream 5.0 Pro sits at 1266.

#ModelCompanyScoreStrengths
1 GPT Image 2AA Image Arena ELO 1357 (high) — quality and ecosystem crown extended. OpenAI 1357 AA T2I Arena ELO: 1357 (#1) · Token-based pricing · Batch API at 50% discount · High-fidelity inputs
2 Reve 2.1Holds #2 at ELO 1314 with layered generation. Reve 1314 AA T2I Arena ELO: 1314 (#2) · Layered generation · $24 per 1k images · Day-one top-tier quality
3 Nano Banana 2Gemini 3.1 Flash Image Preview branding — new AA top-3 entrant at 1307. Google 1307 AA T2I Arena ELO: 1307 (#3) · Gemini 3.1 Flash Image Preview · Google stack integration
4 MAI-Image-2.5Enterprise image default at ELO 1298. Microsoft 1298 AA T2I Arena ELO: 1298 · Enterprise integration · Microsoft stack
5 Seedream 5.0 ProTop Chinese production pick at ELO 1266 with strong text rendering. ByteDance 1266 AA T2I Arena ELO: 1266 · Layer separation · Strong text rendering

GPT Image 2 high 1338→1357; Nano Banana 2 enters top 3; Seedream 5.0 Pro still the leading Chinese production option in the top tier.

Quality + ecosystem: GPT Image 2; layered generation value: Reve 2.1; Google preview stack: Nano Banana 2; Chinese text/layout: Seedream 5.0 Pro.

As of 2026-08-09, GPT Image 2 (OpenAI, high) leads the Artificial Analysis text-to-image Arena at ELO 1357, ahead of Reve 2.1 (1314) and Nano Banana 2 / Gemini 3.1 Flash Image Preview (Google, 1307). Seedream 5.0 Pro (ByteDance) scores 1266.
03

Video Generation

Video Generation

Current leader Gemini Omni Flash Google

Gemini Omni Flash keeps the AA text-to-video (with audio) crown at ELO 1244 (July 1240). MiniMax H3 surges to #2 at 1240. Dreamina Seedance 2.0 720p is #3 at 1224; Wan2.7-260612 and HappyHorse-1.1 round out the top 5. Sora API sunset remains 2026-09-24.

#ModelCompanyScoreStrengths
1 Gemini Omni FlashAA T2V with audio #1 at ELO 1244 — still the synced A/V crown. Google 1244 AA T2V Arena (w/ audio) #1 · ELO 1244 · Without-audio arena also #1 · ELO 1324 · Native audio-visual sync · Gemini Omni stack
2 MiniMax H3New #2 on AA T2V with audio at ELO 1240 — the biggest August video mover. MiniMax 1240 AA T2V Arena (w/ audio) #2 · ELO 1240 · Without-audio #2 · ELO 1306 · Fast climb vs July field
3 Seedance 2.0Dreamina Seedance 2.0 720p at ELO 1224 — still the ByteDance production workhorse. ByteDance 1224 AA T2V Arena (w/ audio) #3 · ELO 1224 · Production pipeline maturity · Multimodal references
4 Wan2.7-260612Alibaba Wan2.7 build 260612 at ELO 1161 on the with-audio arena. Alibaba 1161 AA T2V Arena (w/ audio) ELO: 1161 · Alibaba Cloud native
5 Kling 3.0 ProStill the APAC short-form value line (July ELO band); not displaced by H3/Seedance fight at the very top. Kuaishou 快手 1110 Native 4K/60fps line · Turbo iteration speed · TikTok / Douyin native style

With-audio arena: Gemini Omni Flash 1240→1244; MiniMax H3 takes #2 from Seedance; Seedance 2.0 still #3.

Synced A/V arena: Gemini Omni Flash; aggressive #2 challenger: MiniMax H3; ByteDance production: Seedance 2.0; APAC short-form still Kling-family adjacent.

As of 2026-08-09, Gemini Omni Flash (Google) leads text-to-video with audio on Artificial Analysis at ELO 1244, ahead of MiniMax H3 (1240) and Dreamina Seedance 2.0 720p (ByteDance, 1224). OpenAI’s Sora API is still scheduled to shut down on 2026-09-24.
04

Code Generation & Agentic Coding

Code Generation & Agentic Coding

Current leader Claude Opus 5 Anthropic

Claude Opus 5 remains the agentic-coding leader on the July SWE-bench Verified high-water mark (~96%). Kimi K3 keeps the open-weights / LMArena WebDev crowd-vote crown (1679). GPT-5.6 Sol still leads Terminal-Bench 2.1 at 88.8%. No cleaner public number beating Opus 5 was confirmed on the August 9 pass.

#ModelCompanyScoreStrengths
1 Claude Opus 5SWE-bench Verified ~96% high-water mark still stands; AA Index now 63. Anthropic 96 SWE-bench Verified: ~96% (#1) · AA Intelligence Index: 63 · $5 / $25 per 1M tokens · Long-horizon agentic edits
2 Kimi K3Open-weights WebDev arena king at 1679; AA Index 60. Moonshot AI 1679 LMArena WebDev: 1679 (#1) · AA Intelligence Index: 60 · Open weights · 2.8T MoE · $3 / $15 per 1M tokens
3 Claude Fable 5Mythos-class coding ceiling; AA Index 62 overall. Anthropic 95 SWE-bench Verified: 95% · Mythos-class agentic coding · 1M+ context
4 GPT-5.6 SolBest sandboxed terminal agent — Terminal-Bench 2.1 88.8%. OpenAI 88.8 Terminal-Bench 2.1: 88.8% (#1) · SWE-bench Pro: 64.6% · Sandboxed shell built-in

No crown flip: Opus 5 still SWE-bench leader; Kimi K3 still WebDev #1; AA overall Index now 63/60 for Opus/Kimi.

Hardest refactors: Opus 5; open frontend + deploy: Kimi K3; Mythos ceiling: Fable 5; sandboxed terminal agents: GPT-5.6 Sol.

As of 2026-08-09, Claude Opus 5 (Anthropic) remains the agentic-coding leader with ~96% on SWE-bench Verified from its July 24 launch measurements. Kimi K3 (Moonshot AI) holds the LMArena WebDev crown at ELO 1679, and GPT-5.6 Sol leads Terminal-Bench 2.1 at 88.8%.
05

Voice / Speech

Voice / Speech

Current leader Realtime 2 OpenAI

No new speech-to-speech king — OpenAI Realtime 2 still leads agentic voice. On AA TTS, Qwen-Audio-3.0-TTS-Plus remains #1 at ELO 1229 (July 1234; same crown, slight Elo drift). Speechify Simba 3.2 is #2 at 1227; Gemini 3.1 Flash TTS #3 at 1210.

#ModelCompanyScoreStrengths
1 Realtime 2Still the agentic voice default — configurable-reasoning speech-to-speech. OpenAI Configurable reasoning · Speech-to-speech agents · Streaming translate + STT variants · Released 2026-05-07
2 Qwen-Audio-3.0-TTS-PlusAA TTS #1 at ELO 1229 — China still holds the component TTS crown. Alibaba 1229 AA TTS ELO: 1229 (#1) · Strong multilingual + CJK · Alibaba Cloud native
3 Simba 3.2AA TTS #2 at ELO 1227 — inches behind Qwen. Speechify 1227 AA TTS ELO: 1227 (#2) · Consumer reading voice quality
4 ElevenLabs v3Still the character voice cloning default; Scribe v2 realtime STT remains in stack. ElevenLabs Character voice cloning SOTA · Scribe v2 realtime STT · Music v2 API · 100+ languages

TTS crown unchanged (Qwen-Audio-3.0-TTS-Plus); Elo 1234→1229. Agentic S2S still Realtime 2.

Agentic voice: Realtime 2; raw TTS quality: Qwen-Audio-3.0-TTS-Plus; character cloning: ElevenLabs v3.

As of 2026-08-09, OpenAI Realtime 2 remains the leading speech-to-speech model for agentic voice. Alibaba’s Qwen-Audio-3.0-TTS-Plus tops the Artificial Analysis text-to-speech leaderboard at ELO 1229, ahead of Speechify Simba 3.2 (1227) and Gemini 3.1 Flash TTS (1210).
06

Music Generation

Music Generation

Current leader Suno v5.5 Suno

Suno v6 still has not shipped as of the August 9 check. Suno v5.5 remains the full-song default. ElevenLabs Music v2 (API since June 15) and Udio v3 stems stay the secondary studio picks; Lyria remains on the Gemini Omni cross-modal path.

#ModelCompanyScoreStrengths
1 Suno v5.5Still the latest shipped Suno — best full-song coherence and lyric prosody. Suno Full-song coherence SOTA · Lyric prosody best · Multilingual vocal · Style transfer
2 Udio v3Studio-grade mixing with stem-level export. Udio Stem-level export · Studio-grade mixing · Strong electronic genres · DAW-friendly workflow
3 ElevenLabs Music v2API-first licensed-catalog music generation since June 15. ElevenLabs API-first music generation · Licensed catalog · Music Finetunes API · Voice stack integration

No crown change. Suno v6 still unannounced/unshipped.

Full songs: Suno v5.5; stems/studio: Udio v3; licensed API music: ElevenLabs Music v2; cross-modal: Lyria via Gemini Omni.

As of 2026-08-09, Suno v5.5 remains the leading music generation model — Suno v6 has still not shipped. ElevenLabs Music v2 continues on API (since 2026-06-15), and Udio v3 leads stem-level studio workflows.
07

Vision / Multimodal Understanding

Vision / Multimodal Understanding

Current leader Claude Fable 5 Anthropic

Claude Fable 5 remains the LMArena Vision leader at ELO 1327 from the July recalculation — no higher public Vision number was confirmed on August 9. Overall AA Intelligence still has Opus 5 at 63 and Fable 5 at 62, so multimodal understanding stays Anthropic-heavy. Gemini 3.6 Flash (Index 52) is the fast Google multimodal pick.

#ModelCompanyScoreStrengths
1 Claude Fable 5LMArena Vision #1 at 1327; AA overall Index 62. Anthropic 1327 LMArena Vision ELO: 1327 (#1) · AA Intelligence Index: 62 · OCR + document SOTA · 1M+ context
2 Claude Opus 5Multimodal flagship with AA overall #1 at 63. Anthropic 63 AA Intelligence Index: 63 (#1 overall) · Document Q&A + charts · $5 / $25 per 1M tokens
3 GPT-5.6 SolImage-grounded reasoning at AA Index 61. OpenAI 61 AA Intelligence Index: 61 · Image-grounded reasoning · 1M context
4 Gemini 3.6 FlashNew on AA overall at 52 — fast Google multimodal stack pick. Google 52 AA Intelligence Index: 52 · High output speed on AA · Gemini multimodal stack

Vision crown unchanged (Fable 5 @ 1327). Gemini 3.6 Flash appears on AA overall at 52.

OCR/docs/charts: Anthropic; image-grounded reasoning: GPT-5.6 Sol; video understanding: Gemini line.

As of 2026-08-09, Claude Fable 5 (Anthropic) remains the LMArena Vision leader at ELO 1327, with Claude Opus 5 (AA Index 63) and GPT-5.6 Sol (AA Index 61) close on overall intelligence. Gemini 3.6 Flash scores 52 on the Artificial Analysis Intelligence Index.
08

Open-Source / Open-Weights

Open-Source / Open-Weights

Current leader Kimi K3 Moonshot AI

Kimi K3 max reaches AA Intelligence Index 60 — still open #1 and global #4, only 3 points behind Opus 5 (63). DeepSeek V4 Flash 0731 jumps to 52 without a price change ($0.14/$0.28). GLM-5.2 max is ~53; MiniMax-M3 is 45.

#ModelCompanyScoreStrengths
1 Kimi K3Open #1 at AA Index 60 — global #4, 3 points off Opus 5. Moonshot AI 60 AA Intelligence Index: 60 (open #1) · 2.8T MoE · 1M context · LMArena WebDev #1 · $3 / $15 per 1M tokens
2 DeepSeek V4 Flash 07310731 refresh to Index 52 at unchanged $0.14/$0.28 — open value king. DeepSeek 52 AA Intelligence Index: 52 · $0.14 in / $0.28 out per 1M · MIT open weights · 1M context
3 GLM-5.2AA max around 53 — strong open Chinese alternative. Zhipu AI 53 AA Intelligence Index: ~53 (max) · Pay-as-you-go API available
4 Gemma 4 12BApache 2.0 on-device multimodal default. Google Apache 2.0 · open weights · 256K context · native multimodal · Runs on 16GB VRAM

Kimi K3 57→60; DeepSeek V4 Flash 0731 47→52; open crown unchanged.

Frontier open deploy: Kimi K3; pure $/IQ: DeepSeek V4 Flash 0731; on-device multimodal: Gemma 4 12B.

As of 2026-08-09, Kimi K3 (Moonshot AI) leads open-weights models with an Artificial Analysis Intelligence Index of 60 — global #4 and 3 points behind closed frontier Claude Opus 5 (63). DeepSeek V4 Flash 0731 scores 52 at $0.14/$0.28 per 1M tokens.
09

Cost-Effectiveness / Value

Intelligence per Dollar

Current leader DeepSeek V4 Flash 0731 DeepSeek previously: DeepSeek V4 Flash

DeepSeek V4 Flash 0731 is the new value headline: Intelligence Index 52 at the same $0.14/$0.28 pricing that made Flash the July king at 47. MiniMax-M3 remains the cheap high-coding tier at Index 45. Kimi K3 offers Index 60 but at $3/$15 — quality, not pure $/IQ.

#ModelCompanyScoreStrengths
1 DeepSeek V4 Flash 0731Index 52 at $0.14/$0.28 — still the intelligence-per-dollar king after the 0731 refresh. DeepSeek 52 AA Intelligence Index: 52 · $0.14 in / $0.28 out per 1M · Blended cost still ~1/10 of flagships · MIT open weights · 1M context
2 MiniMax-M3AA Index 45 — cheap high-coding value tier. MiniMax 45 AA Intelligence Index: 45 · Strong coding-per-dollar · Standard API pricing tier
3 GLM-5.2Higher absolute price than Flash, stronger index (~53 max). Zhipu AI 53 AA Intelligence Index: ~53 (max) · Pay-as-you-go API
4 Kimi K3Not the cheapest — but best open quality at $3/$15 with Index 60. Moonshot AI 60 AA Intelligence Index: 60 · $3 / $15 per 1M tokens · Open weights frontier-adjacent

Leader name updates to V4 Flash 0731; Index 47→52 with unchanged price. MiniMax-M3 still the SWE-bench value pick from July narrative.

Highest volume low-cost: V4 Flash 0731; stronger open reasoning on a budget: Kimi K2.6-class / GLM; quality-per-dollar frontier-adjacent: Kimi K3 if you can pay $15 out.

As of 2026-08-09, DeepSeek V4 Flash 0731 leads cost-effectiveness with an Artificial Analysis Intelligence Index of 52 at $0.14 input / $0.28 output per 1M tokens — prices unchanged while the index rose from the July issue’s 47.

This month's trends

Sources & methodology

Data as of 2026-08-09

Monthly archive