Issue archive · updated 2026-07-27

LLM Leaderboard Archive — 2026-07

Archived 2026-07 LLM leaderboard: category leaders and full rankings for that issue.

01

Text Generation & Reasoning

توليد النصوص والاستدلال

Current leader Claude Opus 5 Anthropic previously: Claude Fable 5

July 24 — Anthropic ships Claude Opus 5 and retakes the Artificial Analysis Intelligence Index crown at 61 (max effort), a point ahead of its own Mythos-class Fable 5 (60). OpenAI's GPT-5.6 Sol lands at 59, and Moonshot's open-weights Kimi K3 breaks into the frontier band at 57.

#ModelCompanyScoreStrengths
1 Claude Opus 5Released July 24. Retakes #1 on the AA Intelligence Index at 61 — at half the Mythos-class price. Anthropic 61 Intelligence Index: 61 (max) / 60 (xhigh) · SWE-bench Verified ~96% · $5 / $25 per 1M tokens · 1M context · Released 2026-07-24
2 Claude Fable 5June 9 Mythos-class debut; now #2 overall at 60 after Opus 5. Briefly export-paused in June, restored June 23. Anthropic 60 Intelligence Index: 60 · Humanity's Last Exam: 53% · Mythos-class safeguards · 1M+ context · Adaptive Reasoning · Released 2026-06-09
3 GPT-5.6 SolJune 26 preview, GA July 9. Sol/Terra/Luna tiered family; Sol max scores 59 with Terminal-Bench 2.1 at 88.8%. OpenAI 59 Intelligence Index: 59 (max) · Terminal-Bench 2.1: 88.8% · $5 / $30 per 1M tokens · 1M context · GA 2026-07-09
4 Kimi K3July 17. 2.8T-parameter open-weights MoE — the first open model inside the frontier band, and #1 on LMArena WebDev. Moonshot AI 57 Intelligence Index: 57 (open #1) · 2.8T MoE · 1M context · LMArena WebDev #1 · Open weights · Released 2026-07-17
5 Grok 4.5July 8 release at $2/$6 per 1M — the cheapest frontier-tier reasoning model. xAI 53 Aggressive pricing: $2 / $6 per 1M · Always-on reasoning · 1M context · Strong agentic + tool use · Released 2026-07-08

Released July 24, 2026. AA Intelligence Index: 61 (max) / 60 (xhigh) at $5/$25 per 1M tokens — half the Mythos-class price.

Opus 5 becomes the default frontier workhorse; Fable 5 stays the Mythos-class ceiling; GPT-5.6 Sol is OpenAI's tiered answer; Kimi K3 is the first open model inside the frontier band.

As of 2026-07-27, Claude Opus 5 (Anthropic, released 2026-07-24) leads the Artificial Analysis Intelligence Index at 61 in max-effort mode, ahead of Claude Fable 5 (60), GPT-5.6 Sol (OpenAI, 59), and open-weights Kimi K3 (Moonshot AI, 57). Opus 5 is priced at $5 input / $25 output per 1M tokens — roughly half the Mythos-class tier.
02

Image Generation

توليد الصور

Current leader GPT Image 2 OpenAI previously: GPT Image-2

GPT Image 2 now tops the Artificial Analysis text-to-image Arena on raw quality too (ELO 1338) — quality and ecosystem crown in one. Reve 2.1 debuts at #2 (ELO 1301) on launch day, ByteDance's Seedream 5.0 Pro enters the top 10, and Black Forest's FLUX 3 lands in limited release as the first omni image-video-audio model.

#ModelCompanyScoreStrengths
1 GPT Image 2Now #1 on the AA text-to-image Arena (ELO 1338) on top of its ecosystem and pricing lead. OpenAI 1338 AA T2I Arena ELO: 1338 (#1) · Token-based pricing · Batch API at 50% discount · High-fidelity inputs · Released 2026-04-21
2 Reve 2.1July 9 launch. Debuts #2 on the AA arena with layered generation at a disruptive $24 per 1k images. Reve 1301 AA T2I Arena ELO: 1301 (#2) · Layered generation · $24 per 1k images · Day-one top-tier quality
3 Seedream 5.0 ProJuly 8. Layer separation and best-in-class text rendering push it into the AA top 10. ByteDance 1239 AA T2I Arena ELO: 1239 (top 10) · Layer separation · Strong text rendering · $90 per 1k images

Recraft V4.1 falls out of the top tier; Reve 2.1 (Jul 9) and Seedream 5.0 Pro (Jul 8) enter; FLUX 3 (Jul 23) limited release.

GPT Image 2 for quality + ecosystem; Reve 2.1 for layered generation at $24/1k images; Seedream 5.0 Pro for text rendering; Firefly remains the IP-cleared enterprise default.

As of 2026-07-27, GPT Image 2 (OpenAI) leads the Artificial Analysis text-to-image Arena at ELO 1338, ahead of Reve 2.1 (1301, launched 2026-07-09) and MAI-Image-2.5 (Microsoft, 1269), with Seedream 5.0 Pro (ByteDance, 1239) in the top 10. Black Forest Labs released FLUX 3 on 2026-07-23 in limited release — an omni model spanning image, 20-second video, and audio.
03

Video Generation

توليد الفيديو

Current leader Gemini Omni Flash Google previously: Seedance 2.0

Google's Gemini Omni Flash takes the Artificial Analysis text-to-video Arena (with audio) at ELO 1240, ending Seedance 2.0's reign (1225). ByteDance's answer is already rolling out: Seedance 2.5 (July 16) brings native 30-second 4K/10-bit generation, though it has no Arena score yet. OpenAI exits the field — the Sora API shuts down September 24.

#ModelCompanyScoreStrengths
1 Gemini Omni FlashNew #1 on the AA text-to-video Arena (with audio) at ELO 1240 — the first Google model to hold the synced audio-visual crown. Google 1240 AA T2V Arena (w/ audio) #1 · ELO 1240 · Native audio-visual sync · Gemini Omni stack integration · API-available today
2 Seedance 2.0 / 2.52.0 now #2 at ELO 1225; 2.5 (native 30s, 4K/10-bit, 50 reference inputs) rolling out since July 16, not yet arena-listed. ByteDance 1225 AA T2V Arena ELO: 1225 (#2) · 2.5: native 30s 4K/10-bit · Multimodal input (text/image/audio/video) · Dual-Branch Diffusion Transformer
3 Kling 3.0 ProJune 17 Turbo update keeps it the APAC short-form value pick; native 4K/60fps from the 3.0 line. Kuaishou 快手 1110 AA T2V Arena ELO: 1110 · 3.0 Turbo: fastest iteration · Native 4K/60fps (3.0 line) · TikTok / Douyin native style

Arena crown passes Seedance 2.0 → Gemini Omni Flash (1240 vs 1225); Seedance 2.5 rolling out since July 16; Sora API sunset set for Sep 24.

Gemini Omni Flash for arena-topping synced audio-visual; Seedance 2.5 for 30s 4K production; Kling 3.0 for APAC short-form value.

As of 2026-07-27, Gemini Omni Flash (Google) leads text-to-video generation, ranking #1 on the Artificial Analysis text-to-video Arena (with audio) at ELO 1240, ahead of Seedance 2.0 (ByteDance, 1225) and Wan 2.7 (Alibaba, 1160). ByteDance began rolling out Seedance 2.5 on 2026-07-16 with native 30-second 4K/10-bit output. OpenAI's Sora API is scheduled to shut down on 2026-09-24.
04

Code Generation & Agentic Coding

توليد الكود والبرمجة الوكيلة

Current leader Claude Opus 5 Anthropic previously: Claude Fable 5

Claude Opus 5 (July 24) pushes SWE-bench Verified to ~96% — a new high-water mark — while Moonshot's open-weights Kimi K3 storms LMArena WebDev at 1679, the first Chinese model to top the frontend code arena. The coding frontier splits in two: Anthropic owns the benchmark, Kimi owns the crowd vote.

#ModelCompanyScoreStrengths
1 Claude Opus 5SWE-bench Verified ~96% on launch day (independent runs measure up to 97.0%) — the new coding frontier. Anthropic 96 SWE-bench Verified: ~96% (#1) · $5 / $25 per 1M tokens · Long-horizon agentic edits · 1M context · Released 2026-07-24
2 Kimi K3Tops LMArena WebDev at 1679 — the first Chinese model to rank #1 on the frontend arena. Open weights, SWE-bench Verified 93.4% (independent). Moonshot AI 1679 LMArena WebDev: 1679 (#1) · Open weights · 2.8T MoE · SWE-bench Verified: 93.4% · $3 / $15 per 1M tokens · Released 2026-07-17
3 Claude Fable 5SWE-bench Verified 95%; the Mythos-class ceiling for the hardest multi-file, long-horizon work. Anthropic 95 SWE-bench Verified: 95% · Mythos-class agentic coding · 1M+ context · Restored after June export pause
4 GPT-5.6 SolTerminal-Bench 2.1 at 88.8% — best sandboxed terminal agent; SWE-bench Pro 64.6%. OpenAI 88.8 Terminal-Bench 2.1: 88.8% (#1) · SWE-bench Pro: 64.6% · Sandboxed shell built-in · GA 2026-07-09

Opus 5 #1 on SWE-bench Verified (~96%, Jul 24); Kimi K3 #1 on LMArena WebDev (1679, Jul 16) ahead of Fable 5 (1631) and GPT-5.6 Sol (1618).

Opus 5 for the hardest refactors and long-horizon agents; Fable 5 as the Mythos-class ceiling; Kimi K3 for frontend work and open deployment; GPT-5.6 Sol for sandboxed terminal agents.

As of 2026-07-27, Claude Opus 5 (Anthropic, released 2026-07-24) leads agentic coding with ~96% on SWE-bench Verified, the highest score on record. Kimi K3 (Moonshot AI, open weights) tops the LMArena WebDev arena at ELO 1679 — the first Chinese model to rank #1 there — ahead of Claude Fable 5 (1631) and GPT-5.6 Sol (OpenAI, 1618). GPT-5.6 Sol also leads Terminal-Bench 2.1 at 88.8% for sandboxed terminal agents.
05

Voice / Speech

الصوت / الكلام

Current leader Realtime 2 OpenAI previously: ElevenLabs v3

No new speech-to-speech king — OpenAI Realtime 2 still leads agentic voice. But the component crowns moved to China: Alibaba's Qwen-Audio-3.0-TTS-Plus tops the AA text-to-speech arena (ELO 1234), and Fun-Realtime-ASR-preview leads speech-to-text at 1.7% WER. ElevenLabs answers with Music v2 and Scribe v2 realtime.

#ModelCompanyScoreStrengths
1 Realtime 2Still the agentic voice default. Configurable-reasoning speech-to-speech with realtime translate + Whisper variants. OpenAI Configurable reasoning · Speech-to-speech agents · Streaming translate variant · Streaming STT variant · Released 2026-05-07
2 Qwen-Audio-3.0-TTS-PlusNew #1 on the AA text-to-speech arena at ELO 1234 — China's first TTS crown. Alibaba 1234 AA TTS ELO: 1234 (#1) · July 2026 release · Strong multilingual + CJK · Alibaba Cloud native
3 ElevenLabs v3Slipped to #11 on AA TTS but remains the character voice cloning default; new Scribe v2 realtime + Music v2 APIs. ElevenLabs Character voice cloning SOTA · Scribe v2 realtime STT (2.2% WER) · Music v2 API (June 15) · 100+ languages

TTS crown: Fun-Realtime-TTS → Qwen-Audio-3.0-TTS-Plus (1234); STT crown: MAI-Transcribe-1.5 → Fun-Realtime-ASR-preview (1.7% WER).

Realtime 2 for agentic voice; Qwen-Audio-3.0-TTS-Plus for raw TTS quality; Fun-Realtime-ASR for transcription; ElevenLabs v3 for character voice cloning.

As of 2026-07-27, OpenAI Realtime 2 remains the leading speech-to-speech model for agentic voice. Alibaba's Qwen-Audio-3.0-TTS-Plus tops the Artificial Analysis text-to-speech leaderboard at ELO 1234, ahead of Speechify Simba 3.2 (1230) and Gemini 3.1 Flash TTS (1215). Fun-Realtime-ASR-preview leads speech-to-text at 1.7% WER, ahead of ElevenLabs Scribe v2 (2.2%).
06

Music Generation

توليد الموسيقى

Current leader Suno v5.5 Suno

Correction to last month: Suno v6 never shipped — v5.5 remains the deployed default and category leader. The real action moved elsewhere: ElevenLabs Music v2 hit the API on June 15, and Google's Lyria keeps riding Gemini Omni for cross-modal any-to-music workflows.

#ModelCompanyScoreStrengths
1 Suno v5.5Still the latest Suno release — v6 has not shipped. Best full-song coherence and lyric prosody. Suno Full-song coherence SOTA · Lyric prosody best · Multilingual vocal · Style transfer
2 Udio v3Studio-grade mixing with stem-level output for producers. Udio Stem-level export · Studio-grade mixing · Strong electronic genres · DAW-friendly workflow
3 ElevenLabs Music v2June 15 API release. Licensed-catalog music generation with fine-tuning support. ElevenLabs API-first music generation · Licensed catalog · Music Finetunes API · ElevenLabs voice stack integration

Suno v6 missed its rumored June window and remains unannounced; v5.5 stays #1. ElevenLabs Music v2 enters via API.

Suno v5.5 for full-song generation; Udio v3 for studio-grade stems; ElevenLabs Music v2 for licensed catalog workflows; Lyria via Gemini Omni for cross-modal.

As of 2026-07-27, Suno v5.5 remains the leading music generation model — the rumored Suno v6 has not shipped. ElevenLabs released Music v2 to its API on 2026-06-15, and Udio v3 continues to lead on studio-grade mixing and stem-level output. Google's Lyria is integrated into Gemini Omni for cross-modal (image / video / music) workflows.
07

Vision / Multimodal Understanding

الرؤية / الفهم متعدد الوسائط

Current leader Claude Fable 5 Anthropic previously: Claude Opus 4.7-thinking

Fable 5 takes the LMArena Vision crown at 1327 ELO, ending the Opus 4.7-generation sweep. Opus 5 and GPT-5.6 Sol cluster within ~15 points behind — the top of the vision arena is now a genuine three-lab race.

#ModelCompanyScoreStrengths
1 Claude Fable 5Tops LMArena Vision at 1327 (recalculated July 12 on fresh votes) — leads Text arena at the same time. Anthropic 1327 LMArena Vision ELO: 1327 (#1) · Also #1 on Text arena · OCR + document SOTA · 1M+ context
2 Claude Opus 5New multimodal flagship, clustering within ~15 ELO of Fable 5 on the Vision arena. Anthropic Frontier multimodal understanding · Document Q&A + charts · $5 / $25 per 1M tokens · Released 2026-07-24
3 Gemini 3.1 ProBest video understanding and long-form temporal reasoning. Google Video understanding SOTA · Long-form temporal · Robotics-ER 1.6 integration · Multimodal context 2M+

Vision #1 passes Opus 4.7-thinking (1309) → Fable 5 (1327, recalculated July 12 on fresh votes).

Anthropic for OCR + document Q&A + chart understanding; GPT-5.6 Sol for image-grounded reasoning; Gemini 3.1 Pro for video understanding.

As of 2026-07-27, Claude Fable 5 (Anthropic) leads the LMArena Vision arena at ELO 1327, with Claude Opus 5 and GPT-5.6 Sol (OpenAI) clustered within roughly 15 points behind. Fable 5 simultaneously holds the top spot on the Text arena. Gemini 3.1 Pro (Google) remains the leader for video understanding.
08

Open-Source / Open-Weights

مفتوح المصدر / أوزان مفتوحة

Current leader Kimi K3 Moonshot AI previously: Kimi K2.6

Moonshot's Kimi K3 (July 17) is the new open-weights king: a 2.8T-parameter MoE at AA Intelligence Index 57 — global #3 and the first open model inside the frontier band, just 4 points behind Opus 5. It also tops LMArena WebDev outright. The catch: output pricing jumped to $15/1M, nearly 4× K2.6.

#ModelCompanyScoreStrengths
1 Kimi K3Open-weights leader at AA Intelligence Index 57 — global #3, 4 points behind the closed frontier. Also #1 on LMArena WebDev. Moonshot AI 57 AA Intelligence Index: 57 (open #1) · 2.8T MoE · 1M context · LMArena WebDev #1 · $3 / $15 per 1M tokens · Released 2026-07-17
2 Kimi K2.6Former open king, now the value pick of the K-line: Index 54 at a quarter of K3's output price. Moonshot AI 54 AA Intelligence Index: 54 · Open weights · ~$4 output per 1M (vs K3's $15) · Strong Chinese + English
3 DeepSeek V4 ProMIT-licensed open weights at AA Intelligence Index 52 — the cheap frontier-grade reasoning default. Full V4 still rumored. DeepSeek 52 AA Intelligence Index: 52 · MIT license · open weights · $0.435 / $0.87 per 1M · 1M-token context
4 Gemma 4 12BApache 2.0 open weights. Encoder-free native multimodal, 256K context, runs on a 16GB laptop. Google Apache 2.0 · open weights · 256K context · native multimodal · Runs on 16GB VRAM · MMLU-Pro 77.2 · GPQA-Diamond 78.8

Kimi K3 (57) takes the open crown from K2.6 (54); K2.7 Code (1T/32B MoE) arrived mid-June for the coding line; DeepSeek's full V4 remains rumor only.

Kimi K3 for frontier-grade open deployment; K2.6 for the same ecosystem at a quarter of the output price; DeepSeek V4 Pro for cheap reasoning; Gemma 4 12B for on-device multimodal.

As of 2026-07-27, Kimi K3 (Moonshot AI, released 2026-07-17) leads open-weights models with an Artificial Analysis Intelligence Index of 57 — global #3 and 4 points behind the closed frontier (Claude Opus 5, 61). The 2.8T-parameter MoE offers 1M-token context at $3/$15 per 1M tokens and ranks #1 on LMArena WebDev. Kimi K2.6 (54) and DeepSeek V4 Pro (MIT, 52) complete the open top tier; Google's Gemma 4 12B (Apache 2.0) remains the on-device multimodal pick.
09

Cost-Effectiveness / Value

الذكاء مقابل الدولار

Current leader DeepSeek V4 Flash DeepSeek

The value war froze in place: DeepSeek hasn't touched prices since April, and V4 Flash still delivers Intelligence Index 47 at a blended ~$0.06 per 1M tokens — a tenth of comparable flagships. Watch two things: the old deepseek-chat/reasoner aliases retired July 24 (migrate to deepseek-v4-flash/pro), and MiniMax M3's launch promo ended, back to $0.60/$2.40 standard.

#ModelCompanyScoreStrengths
1 DeepSeek V4 FlashThe intelligence-per-dollar king, unmoved since April. Index 47 at $0.14/$0.28 per 1M — about a tenth of comparable Flash flagships. DeepSeek 47 AA Intelligence Index: 47 · $0.14 in / $0.28 out per 1M · Blended ≈ $0.06 / 1M · MIT open weights · 1M context
2 DeepSeek V4 ProBest balance in the high-intelligence + low-price quadrant. Index 52 at $0.435/$0.87 — a fraction of same-tier flagships. DeepSeek 52 AA Intelligence Index: 52 · $0.435 in / $0.87 out per 1M · Frontier-tier reasoning · MIT open weights
3 MiniMax M3The cheapest SWE-bench-Verified-80%+ model. Launch promo ended — now $0.60/$2.40 standard. MiniMax SWE-bench Verified: 80%+ · $0.60 in / $2.40 out per 1M · Cheapest agentic-capable tier · Open weights

No price moves from the leaders; DeepSeek aliases retired July 24; MiniMax M3 promo over; GLM-5.2 opened pay-as-you-go API.

V4 Flash for highest-volume low-cost work; V4 Pro for stronger reasoning on a budget; MiniMax M3 as the cheapest SWE-bench-80%+ option; Qwen3.7 Plus to diversify vendors.

As of 2026-07-27, DeepSeek V4 Flash leads cost-effectiveness with an Artificial Analysis Intelligence Index of 47 at $0.14 input / $0.28 output per 1M tokens — a blended cost near $0.06 per 1M, roughly one-tenth of comparable flagships. Prices have held since April; the legacy deepseek-chat and deepseek-reasoner aliases were retired on 2026-07-24. MiniMax M3 ($0.60/$2.40) is the cheapest model scoring above 80% on SWE-bench Verified; GLM-5.2 opened pay-as-you-go API at $1.40/$4.40.

This month's trends

Sources & methodology

Data as of 2026-07-27

Monthly archive