Issue archive · updated 2026-04-29

LLM Leaderboard Archive — 2026-04

Archived 2026-04 LLM leaderboard: category leaders and full rankings for that issue.

01

Pembuatan Teks & Penalaran

Generasi teks & penalaran

Current leader GPT-5.5 OpenAI previously: GPT-5.4

2026 memasuki era tiga raksasa — tidak ada model dominan tunggal, pilihan terbaik bergantung pada tugas yang dihadapi.

#ModelCompanyScoreStrengths
1 GPT-5.5Released April 23, the first fully retrained foundation model since GPT-5. OpenAI 89 Terminal-Bench 2.0: 82.7% · OSWorld-Verified: 78.7% · GDPval: 84.9% · ARC-AGI-2: 85.0% · 1M-token context
2 Claude Opus 4.7Released April 16, strongest at long-context and code review. Anthropic 86 SWE-Bench Pro: 64.3% · MCP-Atlas: 79.1% · Most reliable multi-step reasoning · Most thorough code-logic review · 1M-token context
3 Gemini 3.1 ProIn preview, strongest at math and algorithmic competition. Google ~85 LiveCodeBench Elo: 2887 · 1M-token context · Lowest API price ($2/$12) · Leading video understanding · Best price-to-performance
Per 2026-04-29, GPT-5.5 (OpenAI) memimpin pembuatan teks dan penalaran dengan skor komposit 89, diikuti Claude Opus 4.7 (Anthropic, 86) dan Gemini 3.1 Pro (Google, ~85). GPT-5.5 menang di Terminal-Bench 2.0 (82.7%) dan OSWorld-Verified (78.7%); Claude Opus 4.7 menang di SWE-Bench Pro (64.3%) dan MCP-Atlas (79.1%); Gemini 3.1 Pro menang di LiveCodeBench Elo (2887). Ketiganya mendukung konteks 1M-token.
02

Teks ke Gambar

Generasi gambar

Current leader GPT Image-2 OpenAI previously: Nano Banana 2

GPT Image-2 mengambil takhta dengan akurasi rendering teks 99,2%, sementara Nano Banana 2 mempertahankan keunggulan dalam pembuatan real-time.

#ModelCompanyScoreStrengths
1 GPT Image-2Highest text-rendering accuracy. OpenAI 99.2% Text-rendering accuracy 99.2% · Chinese / Arabic support · Spatial logic & anatomical correctness · Character consistency · Thinking-mode reasoning engine
2 Nano Banana 2Ultra-fast 4K generation with live web search. Google 4-15s Flash architecture, ultra-fast generation · 4K image in 4-15s · Live web-search integration · Fastest on the market · Deep Gemini-ecosystem integration
3 Flux ProStrongest open-source ecosystem. Black Forest Labs Open-source, commercial use · Rich community ecosystem · Style diversity · Local deployment
Per 2026-04-29, GPT Image-2 (OpenAI) memimpin pembuatan teks-ke-gambar dengan akurasi rendering teks 99,2%, diikuti Nano Banana 2 (Google, output 4K kurang dari 15 detik) dan Flux Pro (Black Forest Labs, ekosistem open-source terkuat). GPT Image-2 menang pada tipografi, konsistensi karakter, dan logika spasial; Nano Banana 2 menang pada kecepatan arsitektur Flash dan integrasi pencarian real-time.
03

Teks ke Video

Generasi video

Current leader Veo 3.1 Google previously: Sora 2

Sora 2 telah keluar; Google Veo 3.1 kini memimpin kemampuan keseluruhan, sementara Seedance 2.0 dan Kling 3.0 memimpin di niche tertentu.

#ModelCompanyScoreStrengths
1 Veo 3.1Native audio + multi-shot, strongest overall. Google Native audio generation · Multi-shot narrative · Excellent physics simulation · YouTube-ecosystem integration
2 Seedance 2.0Strongest multi-shot storyboarding. ByteDance Multi-shot storyboarding · Professional cinematic language · Leading domestic Chinese model · Douyin/TikTok ecosystem integration
3 Kling 3.0 OmniCinematic-grade visuals + most accurate lip-sync. Kuaishou Cinematic-grade visuals · Most accurate lip-sync · Kuaishou ecosystem integration · Optimized for Chinese scenarios
Per 2026-04-29, Google Veo 3.1 memimpin pembuatan teks-ke-video dengan audio native dan narasi multi-shot, diikuti Seedance 2.0 (ByteDance, storyboarding multi-shot terkuat) dan Kling 3.0 Omni (Kuaishou, visual kualitas sinematik dan lip-sync paling akurat). Sora 2 telah dihentikan.
04

Pembuatan Kode

Generasi kode & pengodean agen

Current leader GPT-5.5 (Agentic) OpenAI previously: Claude Opus 4.6

GPT-5.5 merebut kembali kepemimpinan dalam coding agen-terminal; Claude Opus 4.7 masih menguasai refactoring multi-file dan orkestrasi tool.

#ModelCompanyScoreStrengths
1 GPT-5.5Terminal-Bench 2.0 #1, strongest agentic coding. OpenAI 82.7% Terminal-Bench 2.0: 82.7% · Expert-SWE: 73.1% · Autonomous coding judgment · Fewer tokens for the same task
2 Claude Opus 4.7SWE-Bench Pro #1, strongest multi-file refactoring. Anthropic 64.3% SWE-Bench Pro: 64.3% · MCP-Atlas: 79.1% · Multi-file logic review · Code-vulnerability detection
3 Gemini 3.1 ProLiveCodeBench #1, strongest in algorithmic competition. Google 2887 Elo LiveCodeBench Elo: 2887 · 1M-context whole-repo analysis · Lowest price · Best for algorithmic competition
Per 2026-04-29, GPT-5.5 (OpenAI) memimpin pembuatan kode dengan 82,7% pada Terminal-Bench 2.0, diikuti Claude Opus 4.7 (Anthropic, SWE-Bench Pro 64,3%) dan Gemini 3.1 Pro (Google, LiveCodeBench Elo 2887). Pilih GPT-5.5 untuk coding agentik otonom, Claude Opus 4.7 untuk refactoring multi-file, Gemini untuk analisis repo penuh dengan konteks 1M-token.
05

Teks ke Suara

Suara / Bicara

Current leader ElevenLabs v3 ElevenLabs previously: ElevenLabs v2

ElevenLabs tetap menjadi tolok ukur industri untuk realisme suara dan kloning; Hume AI memimpin dalam suara emosional.

#ModelCompanyScoreStrengths
1 ElevenLabs v3Industry-benchmark voice realism. ElevenLabs 9.2/10 Realism score 9.2/10 · 75ms ultra-low latency · 29+ languages · Professional Clone quality · Enterprise-grade API
2 Hume AI OctaveTop of the emotional-voice leaderboard. Hume AI 9.3/10 Emotion recognition 9.3/10 · Emotional response capability · Empathetic interaction · Precise affect awareness
3 GPT-4o VoiceBest real-time conversational experience. OpenAI Low-latency real-time conversation · Natural voice output · Multilingual real-time translation · Deep ChatGPT integration
Per 2026-04-29, ElevenLabs v3 memimpin teks-ke-suara dengan skor realisme 9,2/10 dan latensi 75ms di 29+ bahasa, diikuti Hume AI Octave (peringkat suara emosional tertinggi 9,3/10) dan GPT-4o Voice (pengalaman percakapan real-time terbaik).
06

Pembuatan Musik AI

Generasi musik

Current leader Suno v5.5 Suno previously: Suno v5

Suno v5.5 tetap menjadi platform yang paling banyak digunakan; tool-tool berbeda dalam kecepatan, pasca-produksi, dan deployment enterprise.

#ModelCompanyScoreStrengths
1 Suno v5.5Most widely used AI music platform. Suno Largest user base · Studio multi-track editing · MIDI export · Fastest to a finished song
2 Udio v1.5Strongest post-production and stem control. Udio Stem download · Mix control · Key adjustment · Professional post-production
3 Lyria 3 ProBest for enterprise / API deployment. Google DeepMind Vertex AI delivery · Structured generation · Clear copyright posture · Enterprise-grade deployment
Per 2026-04-29, Suno v5.5 memimpin pembuatan musik AI berdasarkan adopsi pengguna dengan editing multi-track Studio dan ekspor MIDI, diikuti Udio v1.5 (editing tingkat stem dan pasca-produksi terkuat) dan Lyria 3 Pro (Google DeepMind, terbaik untuk deployment enterprise / API via Vertex AI).
07

Pemahaman Visual

Visi / Pemahaman multimodal

Current leader GPT-4o Vision OpenAI previously: GPT-4o Vision

GPT-4o Vision mempertahankan kepemimpinan tujuan umum; Gemini Vision memimpin dalam pemahaman video dan parsing dokumen panjang.

#ModelCompanyScoreStrengths
1 GPT-4o VisionStrongest general-purpose vision understanding. OpenAI UI parsing · Chart understanding · Live visual conversation · Multimodal fusion
2 Gemini VisionLeader for video and long-document understanding. Google 1M-token long documents · Leading video understanding · Multi-frame analysis · Search integration
3 Qwen-VLTop open-source Chinese-scenario vision model. Alibaba Optimized for Chinese scenarios · Open-source, commercial use · Multimodal reasoning · Local deployment
Per 2026-04-29, GPT-4o Vision (OpenAI) memimpin pemahaman visual tujuan umum dengan parsing UI dan percakapan visual real-time terkuat, diikuti Gemini Vision (Google, pemimpin pemahaman dokumen panjang dan video dengan 1M-token) dan Qwen-VL (Alibaba, model open-source terkuat untuk skenario bahasa Mandarin).
08

Sumber Terbuka

Sumber terbuka / Bobot terbuka

Current leader Llama 4 Meta previously: Llama 3

Model open-source mengejar cepat closed-source di beberapa benchmark. Llama 4, DeepSeek V4, dan Qwen3 membentuk tier pertama.

#ModelCompanyScoreStrengths
1 Llama 4Most complete open-source ecosystem. Meta Multimodal support · Largest community ecosystem · Commercial-use license · Multiple sizes
2 DeepSeek V4Strongest open-source reasoning, upgraded architecture. DeepSeek Superior math and reasoning · Best-in-class coding ability · Efficient MoE architecture · Extremely low API price
3 Qwen3Top open-source Chinese model. Alibaba Strongest Chinese understanding · Multimodal support · Agent capability · Full size coverage
Per 2026-04-29, Llama 4 (Meta) memimpin model open-source berdasarkan ekosistem dengan dukungan multimodal dan lisensi commercial-friendly, diikuti DeepSeek V4 (DeepSeek, penalaran open-source terkuat dengan arsitektur MoE dan harga API terendah) dan Qwen3 (Alibaba, model open-source pemimpin untuk bahasa Mandarin dan kemampuan agen).

This month's trends

Sources & methodology

Data as of 2026-04-29

Monthly archive