The latest leaders in reasoning capabilities, showcasing the most advanced models in logical processing and decision-making.
| # | Model | Company | Score | Strengths |
| 1 | Claude Opus 5.5Leading in reasoning with a robust score. | Anthropic | 58 | Logical processing · Decision-making |
| 2 | Claude Opus 5.5 (xhigh with fallback)Strong performance with fallback capabilities. | Anthropic | 56 | Fallback logic · Robust reasoning |
| 3 | Claude Opus 5.5 (high with fallback)High reasoning capacity with fallback. | Anthropic | 54 | High capacity · Fallback support |
| 4 | Claude Fable 5.1 (max with fallback)Maximized reasoning with fallback. | Anthropic | 53 | Maximized logic · Fallback |
| 5 | Claude Fable 5.1 (xhigh with fallback)Xhigh reasoning with fallback. | Anthropic | 53 | Xhigh logic · Fallback |
| 6 | GPT-6 Astra (max)Maximized reasoning capabilities. | OpenAI | 53 | Maximized reasoning · Advanced logic |
Claude Opus 5.5 has taken the lead from Claude Opus 5, with an Intelligence Index score of 58.
Anthropic continues to dominate the reasoning category with its Claude Opus series.
As of 2026-09-24, Claude Opus 5.5 leads with a score of 58 among 169 ranked models.
GPT Image 2.5 Sunburst (max) leads the current Text to Image Arena at Elo 1197, followed by GPT Image 2.5 Flare (max) at 1191 and GPT Image 2 (high) at 1171. Sunburst is listed with 13,433 samples.
| # | Model | Company | Score | Strengths |
| 1 | GPT Image 2.5 Sunburst (max)Current Text to Image Arena leader at Elo 1197. | OpenAI | 1197 | 13,433 samples · API pricing: $210.7 per 1,000 images |
| 2 | GPT Image 2.5 Flare (max)Second in the current Text to Image Arena at Elo 1191. | OpenAI | 1191 | Top-two placement |
| 3 | GPT Image 2 (high)Third in the current Text to Image Arena at Elo 1171. | OpenAI | 1171 | Top-three placement · API pricing: $211.0 per 1,000 images |
| 4 | Grok Imagine Image 2.0Fourth in the current Text to Image Arena at Elo 1150. | xAI | 1150 | API pricing: $60.0 per 1,000 images |
| 5 | MAI-Image-2.6Fifth in the current Text to Image Arena at Elo 1147. | Microsoft | 1147 | API pricing: $38.9 per 1,000 images |
| 6 | Nano Banana 2 (Gemini 3.1 Flash Image)Sixth in the current Text to Image Arena at Elo 1122. | Google | 1122 | API pricing: $67.0 per 1,000 images |
The leader changed from the August issue's GPT Image 2 / Nano Banana 2 pairing.
Among open-weight image models, Ideogram 4.0 (Quality) leads at 1011, followed by Ideogram 4.0 at 1003 and FLUX.2 [dev] at 1000.
As of 2026-09-24, GPT Image 2.5 Sunburst (max) ranks #1 at Elo 1197 in the Text to Image Arena, with 13,433 samples.
Gemini Omni Flash retains first place in the audio-enabled Text to Video arena with 1233 Elo, narrowly ahead of Wan 3.0 at 1229. The ordering differs without audio, where Wan 3.0 leads with 1336.
| # | Model | Company | Score | Strengths |
| 1 | Gemini Omni FlashRetains the audio-video lead, four Elo points ahead of Wan 3.0. | Google | 1233 Elo | No. 1 with audio at 1233 Elo · No. 2 without audio at 1330 · Listed price of $6.00 per minute |
| 2 | Wan 3.0Places second with audio and leads the separate no-audio ranking. | Alibaba | 1229 Elo | No. 2 with audio at 1229 Elo · No. 1 without audio at 1336 · Listed price of $12.00 per minute |
| 3 | Minimax H3 Max post-trained by falRanks only two Elo points behind Wan 3.0 while posting the lowest listed price among the top five. | fal | 1227 Elo | No. 3 with audio at 1227 Elo · Two Elo points behind second place · Listed price of $2.40 per minute |
| 4 | MiniMax H3 Open WeightsThe highest-ranked open-weights video model with audio. | MiniMax | 1220 Elo | Open-weights leader with audio at 1220 Elo · No. 4 in the overall audio ranking · Listed price of $7.80 per minute |
| 5 | Dreamina Seedance 2.0 720pCompletes a tightly grouped top five, 23 Elo points behind the leader. | ByteDance | 1210 Elo | No. 5 with audio at 1210 Elo · Within 23 Elo points of first place · Listed price of $9.07 per minute |
The leader is unchanged from August, but Wan 3.0 has moved into second place while two MiniMax H3 variants occupy ranks three and four.
Competition is tight at the top: only 23 Elo points separate first-ranked Gemini Omni Flash from fifth-ranked Dreamina Seedance 2.0 720p. Pricing is more dispersed, ranging from $2.40 to $12.00 per minute among these five models.
As of 2026-09-24, Wan 3.0 ranks No. 2 with audio at 1229 Elo and No. 1 without audio at 1336.
Terminal-Bench 2.1 evaluates software engineering, system administration, data processing, model training, and security. This independently run 89-task refresh includes environment and instruction fixes, with pass@1 averaged over three repeats.
| # | Model | Company | Score | Strengths |
| 1 | Claude Fable 5.1 Adaptive Reasoning Max Effort with Default FallbackThe highest verified Terminal-Bench 2.1 result. | Anthropic | 91.4% | Ranks first with 91.4% pass@1 · Uses the Max Effort adaptive-reasoning configuration · Evaluated across five operational and technical task areas |
| 2 | Claude Fable 5.1 Adaptive Reasoning Xhigh with Default FallbackFinishes only 0.4 percentage points behind the leading configuration. | Anthropic | 91.0% | Ranks second with 91.0% pass@1 · Uses the Xhigh adaptive-reasoning configuration · Maintains a result above 90% on the 89-task refresh |
| 3 | Claude Fable 5.1 Adaptive Reasoning High with Default FallbackCompletes an all-Claude Fable 5.1 top three. | Anthropic | 89.9% | Ranks third with 89.9% pass@1 · Uses the High adaptive-reasoning configuration · Trails the leading Max Effort variant by 1.5 percentage points |
Claude Fable 5.1 Adaptive Reasoning Max Effort with Default Fallback replaces the August issue’s Claude Opus 5 as the category leader. Because Terminal-Bench 2.1 is a refreshed evaluation, the September percentage should not be treated as a direct comparison with the prior issue’s score.
The leaderboard’s top three results are effort variants of Claude Fable 5.1, spanning 89.9% to 91.4%. For enterprise buyers, this highlights the growing importance of selecting a reasoning-effort tier rather than evaluating only the base model name.
As of 2026-09-24, Claude Fable 5.1 Adaptive Reasoning Max Effort with Default Fallback ranks No. 1 on Terminal-Bench 2.1 with a 91.4% pass@1 result.
Released in August 2026, Cartesia Sonic 3.6 leads the Provider Voice Arena at 1278 Elo. Inworld Realtime TTS-2 follows at 1243, ahead of SpeechifyAI Simba 3.2 at 1238 and Alibaba Qwen-Audio-3.0-TTS-Plus at 1235.
| # | Model | Company | Score | Strengths |
| 1 | Cartesia Sonic 3.6The August 2026 release is the new Provider Voice Arena leader. | Cartesia | 1278 Elo | Ranks first with 1278 Elo · API price listed at $49.0 per 1M characters |
| 2 | Inworld Realtime TTS-2Second overall in the current provider voice ranking. | Inworld | 1243 Elo | Scores 1243 Elo · API price listed at $20.8 per 1M characters |
| 3 | SpeechifyAI Simba 3.2Places third, five Elo points behind Realtime TTS-2. | SpeechifyAI | 1238 Elo | Scores 1238 Elo · API price listed at $6.6 per 1M characters |
| 4 | Alibaba Qwen-Audio-3.0-TTS-PlusAlibaba's entry places fourth in a tightly grouped top tier. | Alibaba | 1235 Elo | Scores 1235 Elo · API price listed at $27.6 per 1M characters |
| 5 | VUI Labs Luna TTSLuna TTS rounds out the arena's top five. | VUI Labs | 1230 Elo | Scores 1230 Elo · API price listed at $80.0 per 1M characters |
| 6 | Inworld Realtime TTS-2 FlashInworld holds two of the top six positions with its standard and Flash variants. | Inworld | 1214 Elo | Scores 1214 Elo · API price listed at $10.4 per 1M characters |
Cartesia Sonic 3.6 replaces August leader Realtime 2, while Inworld Realtime TTS-2 enters the current table in second place.
API pricing varies substantially across the top six: SpeechifyAI Simba 3.2 is listed at $6.6 per 1M characters, compared with $49.0 for Sonic 3.6 and $80.0 for VUI Labs Luna TTS. Breeze TTS 2 is the highest-ranked open-weights TTS model at 1203.
As of 2026-09-24, Cartesia Sonic 3.6 ranks No. 1 in the Provider Voice Arena with an Elo score of 1278.
Mureka V9 leads the Instrumental Music Arena with 1177 Elo, followed by Suno V5.5 at 1171 and Mureka V8 at 1143. The leaderboard is based on blind preference votes and evaluates instrumental and vocals modes separately.
| # | Model | Company | Score | Strengths |
| 1 | Mureka V9Instrumental Music Arena leader, supported by 2,237 samples. | Mureka | 1177 Elo | Highest blind-preference Elo among the ranked instrumental models · Ranks ahead of Suno V5.5 |
| 2 | Suno V5.5Runner-up with 1171 Elo across 2,222 samples. | Suno | 1171 Elo | Second-highest blind-preference score · One of two Suno models in the top four |
| 3 | Mureka V8Third place with 1143 Elo and 3,130 samples. | Mureka | 1143 Elo | Second Mureka model in the top three · Largest sample count among the six listed leaders |
| 4 | Suno V5Fourth place with 1137 Elo across 2,713 samples. | Suno | 1137 Elo | Ranks within the arena's top four · Second Suno model among the four leaders |
| 5 | StepAudio 3 MusicFifth place with 1113 Elo and 2,212 samples. | StepAudio | 1113 Elo | Top-five position in blind instrumental preference voting · Ranks ahead of Google Lyria 3 Pro |
| 6 | Google Lyria 3 ProSixth place with 1102 Elo across 2,959 samples. | Google | 1102 Elo | Included among the six leading instrumental models · Evaluation supported by 2,959 samples |
Mureka V9 replaces August leader Suno v5.5, while Suno V5.5 now ranks second.
Mureka and Suno each hold two of the top four positions. Sample counts range from 2,212 for StepAudio 3 Music to 3,130 for Mureka V8; the source does not provide music-model pricing.
As of 2026-09-24, Mureka V9 ranks No. 1 in the Instrumental Music Arena with 1177 Elo.
The Vision Arena table dated September 13, 2026 covers 152 models and 1,314,119 votes. Claude Fable 5 High leads at 1310±8, narrowly ahead of Qwen3.8 Max at 1302±8 and two Claude Opus 4.7 configurations.
| # | Model | Company | Score | Strengths |
| 1 | Claude Fable 5 HighLeads the September Vision Arena. | Anthropic | 1310±8 | Ranked No. 1 among 152 models · Backed by 1,314,119 arena votes across the full leaderboard |
| 2 | Qwen3.8 MaxFinishes eight Elo points behind the leader. | Alibaba | 1302±8 | Ranked No. 2 overall · Listed at $2 per 1M input tokens and $6 per 1M output tokens |
| 3 | Claude Opus 4.7 HighPlaces third, one point behind Qwen3.8 Max. | Anthropic | 1301±7 | Ranked No. 3 overall · Score interval of 1301±7 |
| 4 | Claude Opus 4.7The standard configuration follows its High variant by one point. | Anthropic | 1300±7 | Ranked No. 4 overall · Only 10 points behind the leader |
| 5 | Claude Opus 4.6 HighExtends Anthropic's presence to four of the top five positions. | Anthropic | 1299±7 | Ranked No. 5 overall · Only one point behind Claude Opus 4.7 |
Claude Fable 5 High takes the top position, succeeding the August issue's Claude Fable 5 listing.
The leader is priced at $10 per 1M input tokens and $50 per 1M output tokens. Runner-up Qwen3.8 Max is listed at $2 and $6 respectively, creating a substantial price gap between the top two Vision Arena models.
As of 2026-09-24, Claude Fable 5 High ranks No. 1 at 1310±8 in a Vision Arena covering 152 models and 1,314,119 votes.
MiMo-V2.6-Pro ranks first among open-weights models with an Artificial Analysis Intelligence Index of 46, ahead of GLM-5.3 max at 45 and Kimi K3 max at 44.
| # | Model | Company | Score | Strengths |
| 1 | MiMo-V2.6-ProThe highest-ranked open-weights model on the current Intelligence Index. | Xiaomi | 46 | Intelligence Index of 46 · Ranks first among open-weights models · Listed with a 1M context window |
| 2 | GLM-5.3 maxSecond among current open-weights models, one index point behind the leader. | Zhipu AI | 45 | Intelligence Index of 45 · Ranks second among open-weights models · GLM-5.3 is listed with a 1M context window |
| 3 | Kimi K3 maxThe August leader now ranks third in the open-weights table. | Moonshot AI | 44 | Intelligence Index of 44 · Ranks third among open-weights models · Listed with a 1M context window |
| 4 | GLM-5.3-FlashGLM-5.3-Flash holds fourth place in the current open-weights ranking. | Zhipu AI | 42 | Intelligence Index of 42 · Ranks fourth among open-weights models |
| 5 | Qwen3.8 2.4T A95BQwen3.8 2.4T A95B records an Intelligence Index of 40. | Alibaba | 40 | Intelligence Index of 40 · Listed among the six leading open-weights models |
| 6 | Qwen3.8-Flash-NextQwen3.8-Flash-Next matches Qwen3.8 2.4T A95B with an index score of 40. | Alibaba | 40 | Intelligence Index of 40 · Listed among the six leading open-weights models |
MiMo-V2.6-Pro replaces August leader Kimi K3 at the top of the open-weights ranking.
Open weights remain competitive but trail the absolute proprietary frontier: MiMo-V2.6-Pro scores 46, versus 58 for Claude Opus 5.5 max with fallback.
As of 2026-09-24, 68 of the 169 models ranked by Artificial Analysis are open weights, led by MiMo-V2.6-Pro with an Intelligence Index of 46.
GPT-6 Luna (low) is the value leader at $0.0045 per task and an Intelligence Index of 21. This designation is an editorial inference from the Artificial Analysis comparison table: it combines the lowest stated positive task cost with higher intelligence than the table’s $0.00 utility models.
| # | Model | Company | Score | Strengths |
| 1 | GPT-6 Luna (low)The lowest stated positive cost per task, paired with an Intelligence Index of 21. | OpenAI | $0.0045/task | $0.0045 cost per task · Intelligence Index of 21 · Best overall value by editorial comparison of cost and intelligence |
| 2 | Granite 4.2 3BListed among the current Cost per Task leaders at $0.01 per task. | IBM | $0.01/task | $0.01 cost per task · 3B model designation · Second model in the published cost-leader ordering |
| 3 | Ministral 3 3BMatches Granite 4.2 3B at a stated cost of $0.01 per task. | Mistral AI | $0.01/task | $0.01 cost per task · 3B model designation · Included among the three current cost leaders |
GPT-6 Luna (low) replaces August leader DeepSeek V4 Flash 0731.
Command A+ and North Mini Code are listed at $0.00 per task, but their Intelligence Index scores are only 13 and 10 respectively. They should not be treated as the best overall value without that performance qualification.
As of 2026-09-24, GPT-6 Luna (low) ranks No. 1 for cost-effectiveness at $0.0045 per task, with an Intelligence Index of 21.