VOL. 2026ISSUE 07Actualizado el 2026-07-27

Leaderboard Mensual de LLMs

julio de 2026

Ocho categorías. Veinticuatro modelos líderes. Actualizado mensualmente. Con citas amigables para IA.

9
categories
31
models
10
sources
Compartir esta ediciónXLinkedIn
Explorar archivo mensual
01
Text Generation & Reasoning

Text Generation & Reasoning

July 24 — Anthropic ships Claude Opus 5 and retakes the Artificial Analysis Intelligence Index crown at 61 (max effort), a point ahead of its own Mythos-class Fable 5 (60). OpenAI's GPT-5.6 Sol lands at 59, and Moonshot's open-weights Kimi K3 breaks into the frontier band at 57.

Previously: Claude Fable 5

Líder actual
Claude Opus 5
Anthropic

Released July 24. Retakes #1 on the AA Intelligence Index at 61 — at half the Mythos-class price.

Puntuación
61
  • 01Intelligence Index: 61 (max) / 60 (xhigh)
  • 02SWE-bench Verified ~96%
  • 03$5 / $25 per 1M tokens
  • 041M context
  • 05Released 2026-07-24
Runners-up
2

Claude Fable 5

Anthropic

June 9 Mythos-class debut; now #2 overall at 60 after Opus 5. Briefly export-paused in June, restored June 23.

  • Intelligence Index: 60
  • Humanity's Last Exam: 53%
  • Mythos-class safeguards
  • 1M+ context · Adaptive Reasoning
  • Released 2026-06-09
60
3

GPT-5.6 Sol

OpenAI

June 26 preview, GA July 9. Sol/Terra/Luna tiered family; Sol max scores 59 with Terminal-Bench 2.1 at 88.8%.

  • Intelligence Index: 59 (max)
  • Terminal-Bench 2.1: 88.8%
  • $5 / $30 per 1M tokens
  • 1M context
  • GA 2026-07-09
59
4

Kimi K3

Moonshot AI

July 17. 2.8T-parameter open-weights MoE — the first open model inside the frontier band, and #1 on LMArena WebDev.

  • Intelligence Index: 57 (open #1)
  • 2.8T MoE · 1M context
  • LMArena WebDev #1
  • Open weights
  • Released 2026-07-17
57
5

Grok 4.5

xAI

July 8 release at $2/$6 per 1M — the cheapest frontier-tier reasoning model.

  • Aggressive pricing: $2 / $6 per 1M
  • Always-on reasoning · 1M context
  • Strong agentic + tool use
  • Released 2026-07-08
53
Change

Released July 24, 2026. AA Intelligence Index: 61 (max) / 60 (xhigh) at $5/$25 per 1M tokens — half the Mythos-class price.

Market

Opus 5 becomes the default frontier workhorse; Fable 5 stays the Mythos-class ceiling; GPT-5.6 Sol is OpenAI's tiered answer; Kimi K3 is the first open model inside the frontier band.

02
Image Generation

Image Generation

GPT Image 2 now tops the Artificial Analysis text-to-image Arena on raw quality too (ELO 1338) — quality and ecosystem crown in one. Reve 2.1 debuts at #2 (ELO 1301) on launch day, ByteDance's Seedream 5.0 Pro enters the top 10, and Black Forest's FLUX 3 lands in limited release as the first omni image-video-audio model.

Previously: GPT Image-2

Líder actual
GPT Image 2
OpenAI

Now #1 on the AA text-to-image Arena (ELO 1338) on top of its ecosystem and pricing lead.

Puntuación
1338
  • 01AA T2I Arena ELO: 1338 (#1)
  • 02Token-based pricing
  • 03Batch API at 50% discount
  • 04High-fidelity inputs
  • 05Released 2026-04-21
Runners-up
2

Reve 2.1

Reve

July 9 launch. Debuts #2 on the AA arena with layered generation at a disruptive $24 per 1k images.

  • AA T2I Arena ELO: 1301 (#2)
  • Layered generation
  • $24 per 1k images
  • Day-one top-tier quality
1301
3

Seedream 5.0 Pro

ByteDance

July 8. Layer separation and best-in-class text rendering push it into the AA top 10.

  • AA T2I Arena ELO: 1239 (top 10)
  • Layer separation
  • Strong text rendering
  • $90 per 1k images
1239
Change

Recraft V4.1 falls out of the top tier; Reve 2.1 (Jul 9) and Seedream 5.0 Pro (Jul 8) enter; FLUX 3 (Jul 23) limited release.

Market

GPT Image 2 for quality + ecosystem; Reve 2.1 for layered generation at $24/1k images; Seedream 5.0 Pro for text rendering; Firefly remains the IP-cleared enterprise default.

03
Video Generation

Video Generation

Google's Gemini Omni Flash takes the Artificial Analysis text-to-video Arena (with audio) at ELO 1240, ending Seedance 2.0's reign (1225). ByteDance's answer is already rolling out: Seedance 2.5 (July 16) brings native 30-second 4K/10-bit generation, though it has no Arena score yet. OpenAI exits the field — the Sora API shuts down September 24.

Previously: Seedance 2.0

Líder actual
Gemini Omni Flash
Google

New #1 on the AA text-to-video Arena (with audio) at ELO 1240 — the first Google model to hold the synced audio-visual crown.

Puntuación
1240
  • 01AA T2V Arena (w/ audio) #1 · ELO 1240
  • 02Native audio-visual sync
  • 03Gemini Omni stack integration
  • 04API-available today
Runners-up
2

Seedance 2.0 / 2.5

ByteDance

2.0 now #2 at ELO 1225; 2.5 (native 30s, 4K/10-bit, 50 reference inputs) rolling out since July 16, not yet arena-listed.

  • AA T2V Arena ELO: 1225 (#2)
  • 2.5: native 30s 4K/10-bit
  • Multimodal input (text/image/audio/video)
  • Dual-Branch Diffusion Transformer
1225
3

Kling 3.0 Pro

Kuaishou 快手

June 17 Turbo update keeps it the APAC short-form value pick; native 4K/60fps from the 3.0 line.

  • AA T2V Arena ELO: 1110
  • 3.0 Turbo: fastest iteration
  • Native 4K/60fps (3.0 line)
  • TikTok / Douyin native style
1110
Change

Arena crown passes Seedance 2.0 → Gemini Omni Flash (1240 vs 1225); Seedance 2.5 rolling out since July 16; Sora API sunset set for Sep 24.

Market

Gemini Omni Flash for arena-topping synced audio-visual; Seedance 2.5 for 30s 4K production; Kling 3.0 for APAC short-form value.

04
Code Generation & Agentic Coding

Code Generation & Agentic Coding

Claude Opus 5 (July 24) pushes SWE-bench Verified to ~96% — a new high-water mark — while Moonshot's open-weights Kimi K3 storms LMArena WebDev at 1679, the first Chinese model to top the frontend code arena. The coding frontier splits in two: Anthropic owns the benchmark, Kimi owns the crowd vote.

Previously: Claude Fable 5

Líder actual
Claude Opus 5
Anthropic

SWE-bench Verified ~96% on launch day (independent runs measure up to 97.0%) — the new coding frontier.

Puntuación
96
  • 01SWE-bench Verified: ~96% (#1)
  • 02$5 / $25 per 1M tokens
  • 03Long-horizon agentic edits
  • 041M context
  • 05Released 2026-07-24
Runners-up
2

Kimi K3

Moonshot AI

Tops LMArena WebDev at 1679 — the first Chinese model to rank #1 on the frontend arena. Open weights, SWE-bench Verified 93.4% (independent).

  • LMArena WebDev: 1679 (#1)
  • Open weights · 2.8T MoE
  • SWE-bench Verified: 93.4%
  • $3 / $15 per 1M tokens
  • Released 2026-07-17
1679
3

Claude Fable 5

Anthropic

SWE-bench Verified 95%; the Mythos-class ceiling for the hardest multi-file, long-horizon work.

  • SWE-bench Verified: 95%
  • Mythos-class agentic coding
  • 1M+ context
  • Restored after June export pause
95
4

GPT-5.6 Sol

OpenAI

Terminal-Bench 2.1 at 88.8% — best sandboxed terminal agent; SWE-bench Pro 64.6%.

  • Terminal-Bench 2.1: 88.8% (#1)
  • SWE-bench Pro: 64.6%
  • Sandboxed shell built-in
  • GA 2026-07-09
88.8
Change

Opus 5 #1 on SWE-bench Verified (~96%, Jul 24); Kimi K3 #1 on LMArena WebDev (1679, Jul 16) ahead of Fable 5 (1631) and GPT-5.6 Sol (1618).

Market

Opus 5 for the hardest refactors and long-horizon agents; Fable 5 as the Mythos-class ceiling; Kimi K3 for frontend work and open deployment; GPT-5.6 Sol for sandboxed terminal agents.

05
Voice / Speech

Voice / Speech

No new speech-to-speech king — OpenAI Realtime 2 still leads agentic voice. But the component crowns moved to China: Alibaba's Qwen-Audio-3.0-TTS-Plus tops the AA text-to-speech arena (ELO 1234), and Fun-Realtime-ASR-preview leads speech-to-text at 1.7% WER. ElevenLabs answers with Music v2 and Scribe v2 realtime.

Previously: ElevenLabs v3

Líder actual
Realtime 2
OpenAI

Still the agentic voice default. Configurable-reasoning speech-to-speech with realtime translate + Whisper variants.

  • 01Configurable reasoning
  • 02Speech-to-speech agents
  • 03Streaming translate variant
  • 04Streaming STT variant
  • 05Released 2026-05-07
Runners-up
2

Qwen-Audio-3.0-TTS-Plus

Alibaba

New #1 on the AA text-to-speech arena at ELO 1234 — China's first TTS crown.

  • AA TTS ELO: 1234 (#1)
  • July 2026 release
  • Strong multilingual + CJK
  • Alibaba Cloud native
1234
3

ElevenLabs v3

ElevenLabs

Slipped to #11 on AA TTS but remains the character voice cloning default; new Scribe v2 realtime + Music v2 APIs.

  • Character voice cloning SOTA
  • Scribe v2 realtime STT (2.2% WER)
  • Music v2 API (June 15)
  • 100+ languages
Change

TTS crown: Fun-Realtime-TTS → Qwen-Audio-3.0-TTS-Plus (1234); STT crown: MAI-Transcribe-1.5 → Fun-Realtime-ASR-preview (1.7% WER).

Market

Realtime 2 for agentic voice; Qwen-Audio-3.0-TTS-Plus for raw TTS quality; Fun-Realtime-ASR for transcription; ElevenLabs v3 for character voice cloning.

06
Music Generation

Music Generation

Correction to last month: Suno v6 never shipped — v5.5 remains the deployed default and category leader. The real action moved elsewhere: ElevenLabs Music v2 hit the API on June 15, and Google's Lyria keeps riding Gemini Omni for cross-modal any-to-music workflows.

Líder actual
Suno v5.5
Suno

Still the latest Suno release — v6 has not shipped. Best full-song coherence and lyric prosody.

  • 01Full-song coherence SOTA
  • 02Lyric prosody best
  • 03Multilingual vocal
  • 04Style transfer
Runners-up
2

Udio v3

Udio

Studio-grade mixing with stem-level output for producers.

  • Stem-level export
  • Studio-grade mixing
  • Strong electronic genres
  • DAW-friendly workflow
3

ElevenLabs Music v2

ElevenLabs

June 15 API release. Licensed-catalog music generation with fine-tuning support.

  • API-first music generation
  • Licensed catalog
  • Music Finetunes API
  • ElevenLabs voice stack integration
Change

Suno v6 missed its rumored June window and remains unannounced; v5.5 stays #1. ElevenLabs Music v2 enters via API.

Market

Suno v5.5 for full-song generation; Udio v3 for studio-grade stems; ElevenLabs Music v2 for licensed catalog workflows; Lyria via Gemini Omni for cross-modal.

07
Vision / Multimodal Understanding

Vision / Multimodal Understanding

Fable 5 takes the LMArena Vision crown at 1327 ELO, ending the Opus 4.7-generation sweep. Opus 5 and GPT-5.6 Sol cluster within ~15 points behind — the top of the vision arena is now a genuine three-lab race.

Previously: Claude Opus 4.7-thinking

Líder actual
Claude Fable 5
Anthropic

Tops LMArena Vision at 1327 (recalculated July 12 on fresh votes) — leads Text arena at the same time.

Puntuación
1327
  • 01LMArena Vision ELO: 1327 (#1)
  • 02Also #1 on Text arena
  • 03OCR + document SOTA
  • 041M+ context
Runners-up
2

Claude Opus 5

Anthropic

New multimodal flagship, clustering within ~15 ELO of Fable 5 on the Vision arena.

  • Frontier multimodal understanding
  • Document Q&A + charts
  • $5 / $25 per 1M tokens
  • Released 2026-07-24
3

Gemini 3.1 Pro

Google

Best video understanding and long-form temporal reasoning.

  • Video understanding SOTA
  • Long-form temporal
  • Robotics-ER 1.6 integration
  • Multimodal context 2M+
Change

Vision #1 passes Opus 4.7-thinking (1309) → Fable 5 (1327, recalculated July 12 on fresh votes).

Market

Anthropic for OCR + document Q&A + chart understanding; GPT-5.6 Sol for image-grounded reasoning; Gemini 3.1 Pro for video understanding.

08
Open-Source / Open-Weights

Open-Source / Open-Weights

Moonshot's Kimi K3 (July 17) is the new open-weights king: a 2.8T-parameter MoE at AA Intelligence Index 57 — global #3 and the first open model inside the frontier band, just 4 points behind Opus 5. It also tops LMArena WebDev outright. The catch: output pricing jumped to $15/1M, nearly 4× K2.6.

Previously: Kimi K2.6

Líder actual
Kimi K3
Moonshot AI

Open-weights leader at AA Intelligence Index 57 — global #3, 4 points behind the closed frontier. Also #1 on LMArena WebDev.

Puntuación
57
  • 01AA Intelligence Index: 57 (open #1)
  • 022.8T MoE · 1M context
  • 03LMArena WebDev #1
  • 04$3 / $15 per 1M tokens
  • 05Released 2026-07-17
Runners-up
2

Kimi K2.6

Moonshot AI

Former open king, now the value pick of the K-line: Index 54 at a quarter of K3's output price.

  • AA Intelligence Index: 54
  • Open weights
  • ~$4 output per 1M (vs K3's $15)
  • Strong Chinese + English
54
3

DeepSeek V4 Pro

DeepSeek

MIT-licensed open weights at AA Intelligence Index 52 — the cheap frontier-grade reasoning default. Full V4 still rumored.

  • AA Intelligence Index: 52
  • MIT license · open weights
  • $0.435 / $0.87 per 1M
  • 1M-token context
52
4

Gemma 4 12B

Google

Apache 2.0 open weights. Encoder-free native multimodal, 256K context, runs on a 16GB laptop.

  • Apache 2.0 · open weights
  • 256K context · native multimodal
  • Runs on 16GB VRAM
  • MMLU-Pro 77.2 · GPQA-Diamond 78.8
Change

Kimi K3 (57) takes the open crown from K2.6 (54); K2.7 Code (1T/32B MoE) arrived mid-June for the coding line; DeepSeek's full V4 remains rumor only.

Market

Kimi K3 for frontier-grade open deployment; K2.6 for the same ecosystem at a quarter of the output price; DeepSeek V4 Pro for cheap reasoning; Gemma 4 12B for on-device multimodal.

09
Intelligence per Dollar

Cost-Effectiveness / Value

The value war froze in place: DeepSeek hasn't touched prices since April, and V4 Flash still delivers Intelligence Index 47 at a blended ~$0.06 per 1M tokens — a tenth of comparable flagships. Watch two things: the old deepseek-chat/reasoner aliases retired July 24 (migrate to deepseek-v4-flash/pro), and MiniMax M3's launch promo ended, back to $0.60/$2.40 standard.

Líder actual
DeepSeek V4 Flash
DeepSeek

The intelligence-per-dollar king, unmoved since April. Index 47 at $0.14/$0.28 per 1M — about a tenth of comparable Flash flagships.

Puntuación
47
  • 01AA Intelligence Index: 47
  • 02$0.14 in / $0.28 out per 1M
  • 03Blended ≈ $0.06 / 1M
  • 04MIT open weights · 1M context
Runners-up
2

DeepSeek V4 Pro

DeepSeek

Best balance in the high-intelligence + low-price quadrant. Index 52 at $0.435/$0.87 — a fraction of same-tier flagships.

  • AA Intelligence Index: 52
  • $0.435 in / $0.87 out per 1M
  • Frontier-tier reasoning
  • MIT open weights
52
3

MiniMax M3

MiniMax

The cheapest SWE-bench-Verified-80%+ model. Launch promo ended — now $0.60/$2.40 standard.

  • SWE-bench Verified: 80%+
  • $0.60 in / $2.40 out per 1M
  • Cheapest agentic-capable tier
  • Open weights
Change

No price moves from the leaders; DeepSeek aliases retired July 24; MiniMax M3 promo over; GLM-5.2 opened pay-as-you-go API.

Market

V4 Flash for highest-volume low-cost work; V4 Pro for stronger reasoning on a budget; MiniMax M3 as the cheapest SWE-bench-80%+ option; Qwen3.7 Plus to diversify vendors.

Editorial · 07 observations

Qué cambió este mes

What changed across the AI model landscape this month — distilled from the data above.

01

Opus 5 Retakes the Crown at Half the Price

Six weeks after the Mythos-class Fable 5, Anthropic ships Opus 5 (July 24): Intelligence Index 61 at $5/$25 — near-Mythos intelligence at Opus-tier pricing. The frontier cadence is now measured in weeks, and Anthropic holds the top 3 overall (Opus 5, Opus 5 xhigh, Fable 5).

02

Kimi K3: Open Source Enters the Frontier

Moonshot's 2.8T-param Kimi K3 (July 17) hits Intelligence Index 57 — global #3, 4 points behind Opus 5 — and tops LMArena WebDev, the first Chinese model to lead that arena. Open weights are no longer a tier below; they're inside the band.

03

GPT-5.6: OpenAI's Tiered Answer

The Sol/Terra/Luna family (preview June 26, GA July 9) splits the frontier into price-performance tiers. Sol max scores 59 on the Intelligence Index and leads Terminal-Bench 2.1 (88.8%) — competitive, but for the first time OpenAI is chasing, not setting, the pace.

04

Google Takes the Video Crown; ByteDance Counters

Gemini Omni Flash ends Seedance 2.0's reign on the AA video arena (1240 vs 1225). ByteDance's Seedance 2.5 (30s native 4K) is already in staged rollout, and OpenAI is leaving the field entirely — Sora's API sunsets September 24.

05

The Music Model That Wasn't

Suno v6, widely expected in June, never shipped — v5.5 is still the leader. Meanwhile ElevenLabs Music v2 quietly entered via API with licensed-catalog positioning. Music is the one category where the frontier stalled this month.

06

Chinese Labs Sweep the Voice Component Crowns

Alibaba's Qwen-Audio-3.0-TTS-Plus takes the AA TTS crown (ELO 1234) and Fun-Realtime-ASR-preview leads STT (1.7% WER). OpenAI still owns the speech-to-speech agent layer, but the picks-and-shovels of voice now come from China.

07

The Value War Hits a Stalemate

DeepSeek prices haven't moved since April — V4 Flash at ~$0.06 blended still leads intelligence-per-dollar. The news is structural instead: legacy aliases retired July 24, MiniMax M3's promo ended, and GLM-5.2 opened pay-as-you-go. OpenAI is reportedly weighing major price cuts; none have landed.

Fuentes
  1. [01]
  2. [02]
    LMArena Leaderboardcommunity leaderboard
  3. [03]
  4. [04]
    OpenAI Changelogofficial changelog
  5. [05]
    Anthropic Newsofficial changelog
  6. [06]
    Google DeepMind Blogofficial changelog
  7. [07]
    DeepSeek API Pricingofficial changelog
  8. [08]
    Moonshot AI / Kimiofficial changelog
  9. [09]
  10. [10]