Issue archive · updated 2026-04-29

LLM Leaderboard Archive — 2026-04

Archived 2026-04 LLM leaderboard: category leaders and full rankings for that issue.

01

Text Generation & Reasoning

Text Generation & Reasoning

Current leader GPT-5.5 OpenAI previously: GPT-5.4

2026 enters a tri-titan era — no single dominant model, the best choice depends on the task at hand.

#ModelCompanyScoreStrengths
1 GPT-5.5Released April 23, the first fully retrained foundation model since GPT-5. OpenAI 89 Terminal-Bench 2.0: 82.7% · OSWorld-Verified: 78.7% · GDPval: 84.9% · ARC-AGI-2: 85.0% · 1M-token context
2 Claude Opus 4.7Released April 16, strongest at long-context and code review. Anthropic 86 SWE-Bench Pro: 64.3% · MCP-Atlas: 79.1% · Most reliable multi-step reasoning · Most thorough code-logic review · 1M-token context
3 Gemini 3.1 ProIn preview, strongest at math and algorithmic competition. Google ~85 LiveCodeBench Elo: 2887 · 1M-token context · Lowest API price ($2/$12) · Leading video understanding · Best price-to-performance

Released April 23, 2026. SOTA on 14 benchmarks, composite score 89.

GPT-5.5 leads in agentic and terminal coding; Claude Opus 4.7 leads in multi-file refactoring and code review; Gemini 3.1 Pro leads in algorithmic competition.

As of 2026-04-29, GPT-5.5 (OpenAI) leads text generation and reasoning with composite score 89, followed by Claude Opus 4.7 (Anthropic, 86) and Gemini 3.1 Pro (Google, ~85). GPT-5.5 wins on Terminal-Bench 2.0 (82.7%) and OSWorld-Verified (78.7%); Claude Opus 4.7 wins on SWE-Bench Pro (64.3%) and MCP-Atlas (79.1%); Gemini 3.1 Pro wins on LiveCodeBench Elo (2887). All three support 1M-token context.
02

Text-to-Image

Image Generation

Current leader GPT Image-2 OpenAI previously: Nano Banana 2

GPT Image-2 takes the throne with 99.2% text-rendering accuracy, while Nano Banana 2 keeps an edge in real-time generation.

#ModelCompanyScoreStrengths
1 GPT Image-2Highest text-rendering accuracy. OpenAI 99.2% Text-rendering accuracy 99.2% · Chinese / Arabic support · Spatial logic & anatomical correctness · Character consistency · Thinking-mode reasoning engine
2 Nano Banana 2Ultra-fast 4K generation with live web search. Google 4-15s Flash architecture, ultra-fast generation · 4K image in 4-15s · Live web-search integration · Fastest on the market · Deep Gemini-ecosystem integration
3 Flux ProStrongest open-source ecosystem. Black Forest Labs Open-source, commercial use · Rich community ecosystem · Style diversity · Local deployment

Released April 2026, with major leads in text rendering and spatial reasoning.

GPT Image-2 wins on typography and physical correctness; Nano Banana 2 wins on speed and live web grounding — they complement rather than replace each other.

As of 2026-04-29, GPT Image-2 (OpenAI) leads text-to-image generation with 99.2% text-rendering accuracy, followed by Nano Banana 2 (Google, sub-15s 4K output) and Flux Pro (Black Forest Labs, strongest open-source ecosystem). GPT Image-2 wins on typography precision, character consistency, and spatial logic; Nano Banana 2 wins on Flash-architecture speed and real-time search integration.
03

Text-to-Video

Video Generation

Current leader Veo 3.1 Google previously: Sora 2

Sora 2 has exited; Google Veo 3.1 now leads in overall capability, while Seedance 2.0 and Kling 3.0 lead in specific niches.

#ModelCompanyScoreStrengths
1 Veo 3.1Native audio + multi-shot, strongest overall. Google Native audio generation · Multi-shot narrative · Excellent physics simulation · YouTube-ecosystem integration
2 Seedance 2.0Strongest multi-shot storyboarding. ByteDance Multi-shot storyboarding · Professional cinematic language · Leading domestic Chinese model · Douyin/TikTok ecosystem integration
3 Kling 3.0 OmniCinematic-grade visuals + most accurate lip-sync. Kuaishou Cinematic-grade visuals · Most accurate lip-sync · Kuaishou ecosystem integration · Optimized for Chinese scenarios

Sora 2 deprecated. Veo 3.1 takes the overall lead.

Veo 3.1 best overall; Seedance for multi-shot storyboarding; Kling for cinematic visuals and lip-sync; Pika for social creators.

As of 2026-04-29, Google Veo 3.1 leads text-to-video generation with native audio and multi-shot narrative, followed by Seedance 2.0 (ByteDance, strongest multi-shot storyboarding) and Kling 3.0 Omni (Kuaishou, cinematic-grade visuals and most accurate lip-sync). Sora 2 has been deprecated.
04

Code Generation

Code Generation & Agentic Coding

Current leader GPT-5.5 (Agentic) OpenAI previously: Claude Opus 4.6

GPT-5.5 retakes the lead in terminal-agent coding; Claude Opus 4.7 still owns multi-file refactoring and tool orchestration.

#ModelCompanyScoreStrengths
1 GPT-5.5Terminal-Bench 2.0 #1, strongest agentic coding. OpenAI 82.7% Terminal-Bench 2.0: 82.7% · Expert-SWE: 73.1% · Autonomous coding judgment · Fewer tokens for the same task
2 Claude Opus 4.7SWE-Bench Pro #1, strongest multi-file refactoring. Anthropic 64.3% SWE-Bench Pro: 64.3% · MCP-Atlas: 79.1% · Multi-file logic review · Code-vulnerability detection
3 Gemini 3.1 ProLiveCodeBench #1, strongest in algorithmic competition. Google 2887 Elo LiveCodeBench Elo: 2887 · 1M-context whole-repo analysis · Lowest price · Best for algorithmic competition

GPT-5.5 released April 23, leading Terminal-Bench 2.0 by 13 percentage points.

Use GPT-5.5 for terminal agentic coding, Claude Opus 4.7 for multi-file refactoring and review, Gemini for whole-repo analysis.

As of 2026-04-29, GPT-5.5 (OpenAI) leads code generation with 82.7% on Terminal-Bench 2.0, followed by Claude Opus 4.7 (Anthropic, SWE-Bench Pro 64.3%) and Gemini 3.1 Pro (Google, LiveCodeBench Elo 2887). Choose GPT-5.5 for autonomous agentic coding, Claude Opus 4.7 for multi-file refactoring, Gemini for whole-repo analysis with 1M-token context.
05

Text-to-Speech

Voice / Speech

Current leader ElevenLabs v3 ElevenLabs previously: ElevenLabs v2

ElevenLabs remains the industry benchmark for voice realism and cloning; Hume AI leads in emotional voice.

#ModelCompanyScoreStrengths
1 ElevenLabs v3Industry-benchmark voice realism. ElevenLabs 9.2/10 Realism score 9.2/10 · 75ms ultra-low latency · 29+ languages · Professional Clone quality · Enterprise-grade API
2 Hume AI OctaveTop of the emotional-voice leaderboard. Hume AI 9.3/10 Emotion recognition 9.3/10 · Emotional response capability · Empathetic interaction · Precise affect awareness
3 GPT-4o VoiceBest real-time conversational experience. OpenAI Low-latency real-time conversation · Natural voice output · Multilingual real-time translation · Deep ChatGPT integration

Continues to lead. v3 ships at 75ms ultra-low latency.

ElevenLabs v3 for professional voiceover and cloning; Hume Octave for emotional interaction; GPT-4o Voice for real-time conversation.

As of 2026-04-29, ElevenLabs v3 leads text-to-speech with a 9.2/10 realism score and 75ms latency across 29+ languages, followed by Hume AI Octave (highest emotional-voice rating 9.3/10) and GPT-4o Voice (best real-time conversational experience).
06

AI Music Generation

Music Generation

Current leader Suno v5.5 Suno previously: Suno v5

Suno v5.5 remains the most-used platform; tools differentiate on speed, post-production, and enterprise deployment.

#ModelCompanyScoreStrengths
1 Suno v5.5Most widely used AI music platform. Suno Largest user base · Studio multi-track editing · MIDI export · Fastest to a finished song
2 Udio v1.5Strongest post-production and stem control. Udio Stem download · Mix control · Key adjustment · Professional post-production
3 Lyria 3 ProBest for enterprise / API deployment. Google DeepMind Vertex AI delivery · Structured generation · Clear copyright posture · Enterprise-grade deployment

Continuous iteration. v5.5 Studio adds multi-track editing and MIDI export.

Suno is fastest to a finished song, Udio is strongest in editing, Lyria is safest for enterprise deployment, ElevenMusic / StableAudio are clearest on commercial rights.

As of 2026-04-29, Suno v5.5 leads AI music generation by user adoption with multi-track Studio editing and MIDI export, followed by Udio v1.5 (strongest stem-level editing and post-production) and Lyria 3 Pro (Google DeepMind, best for enterprise / API deployment via Vertex AI).
07

Vision Understanding

Vision / Multimodal Understanding

Current leader GPT-4o Vision OpenAI previously: GPT-4o Vision

GPT-4o Vision keeps the strongest general-purpose lead; Gemini Vision leads on video understanding and long-document parsing.

#ModelCompanyScoreStrengths
1 GPT-4o VisionStrongest general-purpose vision understanding. OpenAI UI parsing · Chart understanding · Live visual conversation · Multimodal fusion
2 Gemini VisionLeader for video and long-document understanding. Google 1M-token long documents · Leading video understanding · Multi-frame analysis · Search integration
3 Qwen-VLTop open-source Chinese-scenario vision model. Alibaba Optimized for Chinese scenarios · Open-source, commercial use · Multimodal reasoning · Local deployment

Continues to lead — strongest on UI parsing and live visual conversation.

GPT-4o Vision for general-purpose; Gemini Vision for long video / documents; Qwen-VL for the strongest open-source Chinese option.

As of 2026-04-29, GPT-4o Vision (OpenAI) leads general-purpose vision understanding with strongest UI parsing and live visual conversation, followed by Gemini Vision (Google, leading 1M-token long-document and video understanding) and Qwen-VL (Alibaba, strongest open-source Chinese-scenario model).
08

Open Source

Open-Source / Open-Weights

Current leader Llama 4 Meta previously: Llama 3

Open-source models are closing the gap with closed-source on several benchmarks. Llama 4, DeepSeek V4, and Qwen3 form the leading tier.

#ModelCompanyScoreStrengths
1 Llama 4Most complete open-source ecosystem. Meta Multimodal support · Largest community ecosystem · Commercial-use license · Multiple sizes
2 DeepSeek V4Strongest open-source reasoning, upgraded architecture. DeepSeek Superior math and reasoning · Best-in-class coding ability · Efficient MoE architecture · Extremely low API price
3 Qwen3Top open-source Chinese model. Alibaba Strongest Chinese understanding · Multimodal support · Agent capability · Full size coverage

Released in 2026, with major multimodal improvements.

Llama 4 has the largest ecosystem; DeepSeek is strongest at reasoning and the cheapest; Qwen3 leads in Chinese and agent capability.

As of 2026-04-29, Llama 4 (Meta) leads open-source models by ecosystem with multimodal support and commercial-friendly licensing, followed by DeepSeek V4 (DeepSeek, strongest open-source reasoning with MoE architecture and lowest API price) and Qwen3 (Alibaba, leading open-source Chinese-language and agent-capability model).

This month's trends

Sources & methodology

Data as of 2026-04-29

Monthly archive