Issue archive · updated 2026-09-24

LLM Leaderboard Archive — 2026-09

Archived 2026-09 LLM leaderboard: category leaders and full rankings for that issue.

RANKING MENSUAL DE LLMS · 2026-09

Llévate este número

Guarda una imagen limpia o escanea la página en otro dispositivo.

01

Reasoning

Generación de texto y razonamiento

Current leader Claude Opus 5.5 Anthropic previously: Claude Opus 5

The latest leaders in reasoning capabilities, showcasing the most advanced models in logical processing and decision-making.

#ModelCompanyScoreStrengths
1 Claude Opus 5.5Leading in reasoning with a robust score. Anthropic 58 Logical processing · Decision-making
2 Claude Opus 5.5 (xhigh with fallback)Strong performance with fallback capabilities. Anthropic 56 Fallback logic · Robust reasoning
3 Claude Opus 5.5 (high with fallback)High reasoning capacity with fallback. Anthropic 54 High capacity · Fallback support
4 Claude Fable 5.1 (max with fallback)Maximized reasoning with fallback. Anthropic 53 Maximized logic · Fallback
5 Claude Fable 5.1 (xhigh with fallback)Xhigh reasoning with fallback. Anthropic 53 Xhigh logic · Fallback
6 GPT-6 Astra (max)Maximized reasoning capabilities. OpenAI 53 Maximized reasoning · Advanced logic

Claude Opus 5.5 has taken the lead from Claude Opus 5, with an Intelligence Index score of 58.

Anthropic continues to dominate the reasoning category with its Claude Opus series.

As of 2026-09-24, Claude Opus 5.5 leads with a score of 58 among 169 ranked models.
02

Text to Image

Generación de imágenes

Current leader GPT Image 2.5 Sunburst (max) OpenAI previously: GPT Image 2 / Nano Banana 2

GPT Image 2.5 Sunburst (max) leads the current Text to Image Arena at Elo 1197, followed by GPT Image 2.5 Flare (max) at 1191 and GPT Image 2 (high) at 1171. Sunburst is listed with 13,433 samples.

#ModelCompanyScoreStrengths
1 GPT Image 2.5 Sunburst (max)Current Text to Image Arena leader at Elo 1197. OpenAI 1197 13,433 samples · API pricing: $210.7 per 1,000 images
2 GPT Image 2.5 Flare (max)Second in the current Text to Image Arena at Elo 1191. OpenAI 1191 Top-two placement
3 GPT Image 2 (high)Third in the current Text to Image Arena at Elo 1171. OpenAI 1171 Top-three placement · API pricing: $211.0 per 1,000 images
4 Grok Imagine Image 2.0Fourth in the current Text to Image Arena at Elo 1150. xAI 1150 API pricing: $60.0 per 1,000 images
5 MAI-Image-2.6Fifth in the current Text to Image Arena at Elo 1147. Microsoft 1147 API pricing: $38.9 per 1,000 images
6 Nano Banana 2 (Gemini 3.1 Flash Image)Sixth in the current Text to Image Arena at Elo 1122. Google 1122 API pricing: $67.0 per 1,000 images

The leader changed from the August issue's GPT Image 2 / Nano Banana 2 pairing.

Among open-weight image models, Ideogram 4.0 (Quality) leads at 1011, followed by Ideogram 4.0 at 1003 and FLUX.2 [dev] at 1000.

As of 2026-09-24, GPT Image 2.5 Sunburst (max) ranks #1 at Elo 1197 in the Text to Image Arena, with 13,433 samples.
03

Video Generation

Generación de video

Current leader Gemini Omni Flash Google

Gemini Omni Flash retains first place in the audio-enabled Text to Video arena with 1233 Elo, narrowly ahead of Wan 3.0 at 1229. The ordering differs without audio, where Wan 3.0 leads with 1336.

#ModelCompanyScoreStrengths
1 Gemini Omni FlashRetains the audio-video lead, four Elo points ahead of Wan 3.0. Google 1233 Elo No. 1 with audio at 1233 Elo · No. 2 without audio at 1330 · Listed price of $6.00 per minute
2 Wan 3.0Places second with audio and leads the separate no-audio ranking. Alibaba 1229 Elo No. 2 with audio at 1229 Elo · No. 1 without audio at 1336 · Listed price of $12.00 per minute
3 Minimax H3 Max post-trained by falRanks only two Elo points behind Wan 3.0 while posting the lowest listed price among the top five. fal 1227 Elo No. 3 with audio at 1227 Elo · Two Elo points behind second place · Listed price of $2.40 per minute
4 MiniMax H3 Open WeightsThe highest-ranked open-weights video model with audio. MiniMax 1220 Elo Open-weights leader with audio at 1220 Elo · No. 4 in the overall audio ranking · Listed price of $7.80 per minute
5 Dreamina Seedance 2.0 720pCompletes a tightly grouped top five, 23 Elo points behind the leader. ByteDance 1210 Elo No. 5 with audio at 1210 Elo · Within 23 Elo points of first place · Listed price of $9.07 per minute

The leader is unchanged from August, but Wan 3.0 has moved into second place while two MiniMax H3 variants occupy ranks three and four.

Competition is tight at the top: only 23 Elo points separate first-ranked Gemini Omni Flash from fifth-ranked Dreamina Seedance 2.0 720p. Pricing is more dispersed, ranging from $2.40 to $12.00 per minute among these five models.

As of 2026-09-24, Wan 3.0 ranks No. 2 with audio at 1229 Elo and No. 1 without audio at 1336.
04

Code and Terminal Tasks

Generación de código y programación agéntica

Current leader Claude Fable 5.1 Adaptive Reasoning Max Effort with Default Fallback Anthropic previously: Claude Opus 5

Terminal-Bench 2.1 evaluates software engineering, system administration, data processing, model training, and security. This independently run 89-task refresh includes environment and instruction fixes, with pass@1 averaged over three repeats.

#ModelCompanyScoreStrengths
1 Claude Fable 5.1 Adaptive Reasoning Max Effort with Default FallbackThe highest verified Terminal-Bench 2.1 result. Anthropic 91.4% Ranks first with 91.4% pass@1 · Uses the Max Effort adaptive-reasoning configuration · Evaluated across five operational and technical task areas
2 Claude Fable 5.1 Adaptive Reasoning Xhigh with Default FallbackFinishes only 0.4 percentage points behind the leading configuration. Anthropic 91.0% Ranks second with 91.0% pass@1 · Uses the Xhigh adaptive-reasoning configuration · Maintains a result above 90% on the 89-task refresh
3 Claude Fable 5.1 Adaptive Reasoning High with Default FallbackCompletes an all-Claude Fable 5.1 top three. Anthropic 89.9% Ranks third with 89.9% pass@1 · Uses the High adaptive-reasoning configuration · Trails the leading Max Effort variant by 1.5 percentage points

Claude Fable 5.1 Adaptive Reasoning Max Effort with Default Fallback replaces the August issue’s Claude Opus 5 as the category leader. Because Terminal-Bench 2.1 is a refreshed evaluation, the September percentage should not be treated as a direct comparison with the prior issue’s score.

The leaderboard’s top three results are effort variants of Claude Fable 5.1, spanning 89.9% to 91.4%. For enterprise buyers, this highlights the growing importance of selecting a reasoning-effort tier rather than evaluating only the base model name.

As of 2026-09-24, Claude Fable 5.1 Adaptive Reasoning Max Effort with Default Fallback ranks No. 1 on Terminal-Bench 2.1 with a 91.4% pass@1 result.
05

Voice and Text-to-Speech

Voz / Habla

Current leader Cartesia Sonic 3.6 Cartesia previously: Realtime 2

Released in August 2026, Cartesia Sonic 3.6 leads the Provider Voice Arena at 1278 Elo. Inworld Realtime TTS-2 follows at 1243, ahead of SpeechifyAI Simba 3.2 at 1238 and Alibaba Qwen-Audio-3.0-TTS-Plus at 1235.

#ModelCompanyScoreStrengths
1 Cartesia Sonic 3.6The August 2026 release is the new Provider Voice Arena leader. Cartesia 1278 Elo Ranks first with 1278 Elo · API price listed at $49.0 per 1M characters
2 Inworld Realtime TTS-2Second overall in the current provider voice ranking. Inworld 1243 Elo Scores 1243 Elo · API price listed at $20.8 per 1M characters
3 SpeechifyAI Simba 3.2Places third, five Elo points behind Realtime TTS-2. SpeechifyAI 1238 Elo Scores 1238 Elo · API price listed at $6.6 per 1M characters
4 Alibaba Qwen-Audio-3.0-TTS-PlusAlibaba's entry places fourth in a tightly grouped top tier. Alibaba 1235 Elo Scores 1235 Elo · API price listed at $27.6 per 1M characters
5 VUI Labs Luna TTSLuna TTS rounds out the arena's top five. VUI Labs 1230 Elo Scores 1230 Elo · API price listed at $80.0 per 1M characters
6 Inworld Realtime TTS-2 FlashInworld holds two of the top six positions with its standard and Flash variants. Inworld 1214 Elo Scores 1214 Elo · API price listed at $10.4 per 1M characters

Cartesia Sonic 3.6 replaces August leader Realtime 2, while Inworld Realtime TTS-2 enters the current table in second place.

API pricing varies substantially across the top six: SpeechifyAI Simba 3.2 is listed at $6.6 per 1M characters, compared with $49.0 for Sonic 3.6 and $80.0 for VUI Labs Luna TTS. Breeze TTS 2 is the highest-ranked open-weights TTS model at 1203.

As of 2026-09-24, Cartesia Sonic 3.6 ranks No. 1 in the Provider Voice Arena with an Elo score of 1278.
06

Instrumental Music

Generación de música

Current leader Mureka V9 Mureka previously: Suno v5.5

Mureka V9 leads the Instrumental Music Arena with 1177 Elo, followed by Suno V5.5 at 1171 and Mureka V8 at 1143. The leaderboard is based on blind preference votes and evaluates instrumental and vocals modes separately.

#ModelCompanyScoreStrengths
1 Mureka V9Instrumental Music Arena leader, supported by 2,237 samples. Mureka 1177 Elo Highest blind-preference Elo among the ranked instrumental models · Ranks ahead of Suno V5.5
2 Suno V5.5Runner-up with 1171 Elo across 2,222 samples. Suno 1171 Elo Second-highest blind-preference score · One of two Suno models in the top four
3 Mureka V8Third place with 1143 Elo and 3,130 samples. Mureka 1143 Elo Second Mureka model in the top three · Largest sample count among the six listed leaders
4 Suno V5Fourth place with 1137 Elo across 2,713 samples. Suno 1137 Elo Ranks within the arena's top four · Second Suno model among the four leaders
5 StepAudio 3 MusicFifth place with 1113 Elo and 2,212 samples. StepAudio 1113 Elo Top-five position in blind instrumental preference voting · Ranks ahead of Google Lyria 3 Pro
6 Google Lyria 3 ProSixth place with 1102 Elo across 2,959 samples. Google 1102 Elo Included among the six leading instrumental models · Evaluation supported by 2,959 samples

Mureka V9 replaces August leader Suno v5.5, while Suno V5.5 now ranks second.

Mureka and Suno each hold two of the top four positions. Sample counts range from 2,212 for StepAudio 3 Music to 3,130 for Mureka V8; the source does not provide music-model pricing.

As of 2026-09-24, Mureka V9 ranks No. 1 in the Instrumental Music Arena with 1177 Elo.
07

Vision

Visión / Comprensión multimodal

Current leader Claude Fable 5 High Anthropic previously: Claude Fable 5

The Vision Arena table dated September 13, 2026 covers 152 models and 1,314,119 votes. Claude Fable 5 High leads at 1310±8, narrowly ahead of Qwen3.8 Max at 1302±8 and two Claude Opus 4.7 configurations.

#ModelCompanyScoreStrengths
1 Claude Fable 5 HighLeads the September Vision Arena. Anthropic 1310±8 Ranked No. 1 among 152 models · Backed by 1,314,119 arena votes across the full leaderboard
2 Qwen3.8 MaxFinishes eight Elo points behind the leader. Alibaba 1302±8 Ranked No. 2 overall · Listed at $2 per 1M input tokens and $6 per 1M output tokens
3 Claude Opus 4.7 HighPlaces third, one point behind Qwen3.8 Max. Anthropic 1301±7 Ranked No. 3 overall · Score interval of 1301±7
4 Claude Opus 4.7The standard configuration follows its High variant by one point. Anthropic 1300±7 Ranked No. 4 overall · Only 10 points behind the leader
5 Claude Opus 4.6 HighExtends Anthropic's presence to four of the top five positions. Anthropic 1299±7 Ranked No. 5 overall · Only one point behind Claude Opus 4.7

Claude Fable 5 High takes the top position, succeeding the August issue's Claude Fable 5 listing.

The leader is priced at $10 per 1M input tokens and $50 per 1M output tokens. Runner-up Qwen3.8 Max is listed at $2 and $6 respectively, creating a substantial price gap between the top two Vision Arena models.

As of 2026-09-24, Claude Fable 5 High ranks No. 1 at 1310±8 in a Vision Arena covering 152 models and 1,314,119 votes.
08

Open-Weights Models

Código abierto / Pesos abiertos

Current leader MiMo-V2.6-Pro Xiaomi previously: Kimi K3

MiMo-V2.6-Pro ranks first among open-weights models with an Artificial Analysis Intelligence Index of 46, ahead of GLM-5.3 max at 45 and Kimi K3 max at 44.

#ModelCompanyScoreStrengths
1 MiMo-V2.6-ProThe highest-ranked open-weights model on the current Intelligence Index. Xiaomi 46 Intelligence Index of 46 · Ranks first among open-weights models · Listed with a 1M context window
2 GLM-5.3 maxSecond among current open-weights models, one index point behind the leader. Zhipu AI 45 Intelligence Index of 45 · Ranks second among open-weights models · GLM-5.3 is listed with a 1M context window
3 Kimi K3 maxThe August leader now ranks third in the open-weights table. Moonshot AI 44 Intelligence Index of 44 · Ranks third among open-weights models · Listed with a 1M context window
4 GLM-5.3-FlashGLM-5.3-Flash holds fourth place in the current open-weights ranking. Zhipu AI 42 Intelligence Index of 42 · Ranks fourth among open-weights models
5 Qwen3.8 2.4T A95BQwen3.8 2.4T A95B records an Intelligence Index of 40. Alibaba 40 Intelligence Index of 40 · Listed among the six leading open-weights models
6 Qwen3.8-Flash-NextQwen3.8-Flash-Next matches Qwen3.8 2.4T A95B with an index score of 40. Alibaba 40 Intelligence Index of 40 · Listed among the six leading open-weights models

MiMo-V2.6-Pro replaces August leader Kimi K3 at the top of the open-weights ranking.

Open weights remain competitive but trail the absolute proprietary frontier: MiMo-V2.6-Pro scores 46, versus 58 for Claude Opus 5.5 max with fallback.

As of 2026-09-24, 68 of the 169 models ranked by Artificial Analysis are open weights, led by MiMo-V2.6-Pro with an Intelligence Index of 46.
09

Cost-effectiveness

Inteligencia por dólar

Current leader GPT-6 Luna (low) OpenAI previously: DeepSeek V4 Flash 0731

GPT-6 Luna (low) is the value leader at $0.0045 per task and an Intelligence Index of 21. This designation is an editorial inference from the Artificial Analysis comparison table: it combines the lowest stated positive task cost with higher intelligence than the table’s $0.00 utility models.

#ModelCompanyScoreStrengths
1 GPT-6 Luna (low)The lowest stated positive cost per task, paired with an Intelligence Index of 21. OpenAI $0.0045/task $0.0045 cost per task · Intelligence Index of 21 · Best overall value by editorial comparison of cost and intelligence
2 Granite 4.2 3BListed among the current Cost per Task leaders at $0.01 per task. IBM $0.01/task $0.01 cost per task · 3B model designation · Second model in the published cost-leader ordering
3 Ministral 3 3BMatches Granite 4.2 3B at a stated cost of $0.01 per task. Mistral AI $0.01/task $0.01 cost per task · 3B model designation · Included among the three current cost leaders

GPT-6 Luna (low) replaces August leader DeepSeek V4 Flash 0731.

Command A+ and North Mini Code are listed at $0.00 per task, but their Intelligence Index scores are only 13 and 10 respectively. They should not be treated as the best overall value without that performance qualification.

As of 2026-09-24, GPT-6 Luna (low) ranks No. 1 for cost-effectiveness at $0.0045 per task, with an Intelligence Index of 21.

This month's trends

Sources & methodology

Data as of 2026-09-24

Monthly archive