VOL. 2026ISSUE 09Updated as of 2026-09-24

LLM Monthly Leaderboard

Eight categories. Twenty-four leading models. Updated monthly. AI-friendly citations included.

9
categories
46
models
7
sources
Share this issueXLinkedIn
01
Top LLMs in Reasoning

Reasoning

The latest leaders in reasoning capabilities, showcasing the most advanced models in logical processing and decision-making.

Previously: Claude Opus 5

Current leader
Claude Opus 5.5
Anthropic

Leading in reasoning with a robust score.

Score
58
  • 01Logical processing
  • 02Decision-making
Runners-up
2

Claude Opus 5.5 (xhigh with fallback)

Anthropic

Strong performance with fallback capabilities.

  • Fallback logic
  • Robust reasoning
56
3

Claude Opus 5.5 (high with fallback)

Anthropic

High reasoning capacity with fallback.

  • High capacity
  • Fallback support
54
4

Claude Fable 5.1 (max with fallback)

Anthropic

Maximized reasoning with fallback.

  • Maximized logic
  • Fallback
53
5

Claude Fable 5.1 (xhigh with fallback)

Anthropic

Xhigh reasoning with fallback.

  • Xhigh logic
  • Fallback
53
6

GPT-6 Astra (max)

OpenAI

Maximized reasoning capabilities.

  • Maximized reasoning
  • Advanced logic
53
Change

Claude Opus 5.5 has taken the lead from Claude Opus 5, with an Intelligence Index score of 58.

Market

Anthropic continues to dominate the reasoning category with its Claude Opus series.

02
GPT Image 2.5 Sunburst (max) takes the top spot at Elo 1197

Text to Image

GPT Image 2.5 Sunburst (max) leads the current Text to Image Arena at Elo 1197, followed by GPT Image 2.5 Flare (max) at 1191 and GPT Image 2 (high) at 1171. Sunburst is listed with 13,433 samples.

Previously: GPT Image 2 / Nano Banana 2

Current leader
GPT Image 2.5 Sunburst (max)
OpenAI

Current Text to Image Arena leader at Elo 1197.

Score
1197
  • 0113,433 samples
  • 02API pricing: $210.7 per 1,000 images
Runners-up
2

GPT Image 2.5 Flare (max)

OpenAI

Second in the current Text to Image Arena at Elo 1191.

  • Top-two placement
1191
3

GPT Image 2 (high)

OpenAI

Third in the current Text to Image Arena at Elo 1171.

  • Top-three placement
  • API pricing: $211.0 per 1,000 images
1171
4

Grok Imagine Image 2.0

xAI

Fourth in the current Text to Image Arena at Elo 1150.

  • API pricing: $60.0 per 1,000 images
1150
5

MAI-Image-2.6

Microsoft

Fifth in the current Text to Image Arena at Elo 1147.

  • API pricing: $38.9 per 1,000 images
1147
6

Nano Banana 2 (Gemini 3.1 Flash Image)

Google

Sixth in the current Text to Image Arena at Elo 1122.

  • API pricing: $67.0 per 1,000 images
1122
Change

The leader changed from the August issue's GPT Image 2 / Nano Banana 2 pairing.

Market

Among open-weight image models, Ideogram 4.0 (Quality) leads at 1011, followed by Ideogram 4.0 at 1003 and FLUX.2 [dev] at 1000.

03
Text-to-video models with audio

Video Generation

Gemini Omni Flash retains first place in the audio-enabled Text to Video arena with 1233 Elo, narrowly ahead of Wan 3.0 at 1229. The ordering differs without audio, where Wan 3.0 leads with 1336.

Current leader
Gemini Omni Flash
Google

Retains the audio-video lead, four Elo points ahead of Wan 3.0.

Score
1233 Elo
  • 01No. 1 with audio at 1233 Elo
  • 02No. 2 without audio at 1330
  • 03Listed price of $6.00 per minute
Runners-up
2

Wan 3.0

Alibaba

Places second with audio and leads the separate no-audio ranking.

  • No. 2 with audio at 1229 Elo
  • No. 1 without audio at 1336
  • Listed price of $12.00 per minute
1229 Elo
3

Minimax H3 Max post-trained by fal

fal

Ranks only two Elo points behind Wan 3.0 while posting the lowest listed price among the top five.

  • No. 3 with audio at 1227 Elo
  • Two Elo points behind second place
  • Listed price of $2.40 per minute
1227 Elo
4

MiniMax H3 Open Weights

MiniMax

The highest-ranked open-weights video model with audio.

  • Open-weights leader with audio at 1220 Elo
  • No. 4 in the overall audio ranking
  • Listed price of $7.80 per minute
1220 Elo
5

Dreamina Seedance 2.0 720p

ByteDance

Completes a tightly grouped top five, 23 Elo points behind the leader.

  • No. 5 with audio at 1210 Elo
  • Within 23 Elo points of first place
  • Listed price of $9.07 per minute
1210 Elo
Change

The leader is unchanged from August, but Wan 3.0 has moved into second place while two MiniMax H3 variants occupy ranks three and four.

Market

Competition is tight at the top: only 23 Elo points separate first-ranked Gemini Omni Flash from fifth-ranked Dreamina Seedance 2.0 720p. Pricing is more dispersed, ranging from $2.40 to $12.00 per minute among these five models.

04
Claude Fable 5.1 takes all three verified leading positions

Code and Terminal Tasks

Terminal-Bench 2.1 evaluates software engineering, system administration, data processing, model training, and security. This independently run 89-task refresh includes environment and instruction fixes, with pass@1 averaged over three repeats.

Previously: Claude Opus 5

Current leader
Claude Fable 5.1 Adaptive Reasoning Max Effort with Default Fallback
Anthropic

The highest verified Terminal-Bench 2.1 result.

Score
91.4%
  • 01Ranks first with 91.4% pass@1
  • 02Uses the Max Effort adaptive-reasoning configuration
  • 03Evaluated across five operational and technical task areas
Runners-up
2

Claude Fable 5.1 Adaptive Reasoning Xhigh with Default Fallback

Anthropic

Finishes only 0.4 percentage points behind the leading configuration.

  • Ranks second with 91.0% pass@1
  • Uses the Xhigh adaptive-reasoning configuration
  • Maintains a result above 90% on the 89-task refresh
91.0%
3

Claude Fable 5.1 Adaptive Reasoning High with Default Fallback

Anthropic

Completes an all-Claude Fable 5.1 top three.

  • Ranks third with 89.9% pass@1
  • Uses the High adaptive-reasoning configuration
  • Trails the leading Max Effort variant by 1.5 percentage points
89.9%
Change

Claude Fable 5.1 Adaptive Reasoning Max Effort with Default Fallback replaces the August issue’s Claude Opus 5 as the category leader. Because Terminal-Bench 2.1 is a refreshed evaluation, the September percentage should not be treated as a direct comparison with the prior issue’s score.

Market

The leaderboard’s top three results are effort variants of Claude Fable 5.1, spanning 89.9% to 91.4%. For enterprise buyers, this highlights the growing importance of selecting a reasoning-effort tier rather than evaluating only the base model name.

05
Sonic 3.6 takes the Provider Voice Arena lead

Voice and Text-to-Speech

Released in August 2026, Cartesia Sonic 3.6 leads the Provider Voice Arena at 1278 Elo. Inworld Realtime TTS-2 follows at 1243, ahead of SpeechifyAI Simba 3.2 at 1238 and Alibaba Qwen-Audio-3.0-TTS-Plus at 1235.

Previously: Realtime 2

Current leader
Cartesia Sonic 3.6
Cartesia

The August 2026 release is the new Provider Voice Arena leader.

Score
1278 Elo
  • 01Ranks first with 1278 Elo
  • 02API price listed at $49.0 per 1M characters
Runners-up
2

Inworld Realtime TTS-2

Inworld

Second overall in the current provider voice ranking.

  • Scores 1243 Elo
  • API price listed at $20.8 per 1M characters
1243 Elo
3

SpeechifyAI Simba 3.2

SpeechifyAI

Places third, five Elo points behind Realtime TTS-2.

  • Scores 1238 Elo
  • API price listed at $6.6 per 1M characters
1238 Elo
4

Alibaba Qwen-Audio-3.0-TTS-Plus

Alibaba

Alibaba's entry places fourth in a tightly grouped top tier.

  • Scores 1235 Elo
  • API price listed at $27.6 per 1M characters
1235 Elo
5

VUI Labs Luna TTS

VUI Labs

Luna TTS rounds out the arena's top five.

  • Scores 1230 Elo
  • API price listed at $80.0 per 1M characters
1230 Elo
6

Inworld Realtime TTS-2 Flash

Inworld

Inworld holds two of the top six positions with its standard and Flash variants.

  • Scores 1214 Elo
  • API price listed at $10.4 per 1M characters
1214 Elo
Change

Cartesia Sonic 3.6 replaces August leader Realtime 2, while Inworld Realtime TTS-2 enters the current table in second place.

Market

API pricing varies substantially across the top six: SpeechifyAI Simba 3.2 is listed at $6.6 per 1M characters, compared with $49.0 for Sonic 3.6 and $80.0 for VUI Labs Luna TTS. Breeze TTS 2 is the highest-ranked open-weights TTS model at 1203.

06
Blind-preference rankings for AI-generated instrumental music

Instrumental Music

Mureka V9 leads the Instrumental Music Arena with 1177 Elo, followed by Suno V5.5 at 1171 and Mureka V8 at 1143. The leaderboard is based on blind preference votes and evaluates instrumental and vocals modes separately.

Previously: Suno v5.5

Current leader
Mureka V9
Mureka

Instrumental Music Arena leader, supported by 2,237 samples.

Score
1177 Elo
  • 01Highest blind-preference Elo among the ranked instrumental models
  • 02Ranks ahead of Suno V5.5
Runners-up
2

Suno V5.5

Suno

Runner-up with 1171 Elo across 2,222 samples.

  • Second-highest blind-preference score
  • One of two Suno models in the top four
1171 Elo
3

Mureka V8

Mureka

Third place with 1143 Elo and 3,130 samples.

  • Second Mureka model in the top three
  • Largest sample count among the six listed leaders
1143 Elo
4

Suno V5

Suno

Fourth place with 1137 Elo across 2,713 samples.

  • Ranks within the arena's top four
  • Second Suno model among the four leaders
1137 Elo
5

StepAudio 3 Music

StepAudio

Fifth place with 1113 Elo and 2,212 samples.

  • Top-five position in blind instrumental preference voting
  • Ranks ahead of Google Lyria 3 Pro
1113 Elo
6

Google Lyria 3 Pro

Google

Sixth place with 1102 Elo across 2,959 samples.

  • Included among the six leading instrumental models
  • Evaluation supported by 2,959 samples
1102 Elo
Change

Mureka V9 replaces August leader Suno v5.5, while Suno V5.5 now ranks second.

Market

Mureka and Suno each hold two of the top four positions. Sample counts range from 2,212 for StepAudio 3 Music to 3,130 for Mureka V8; the source does not provide music-model pricing.

07
Vision Arena rankings based on more than 1.3 million votes

Vision

The Vision Arena table dated September 13, 2026 covers 152 models and 1,314,119 votes. Claude Fable 5 High leads at 1310±8, narrowly ahead of Qwen3.8 Max at 1302±8 and two Claude Opus 4.7 configurations.

Previously: Claude Fable 5

Current leader
Claude Fable 5 High
Anthropic

Leads the September Vision Arena.

Score
1310±8
  • 01Ranked No. 1 among 152 models
  • 02Backed by 1,314,119 arena votes across the full leaderboard
Runners-up
2

Qwen3.8 Max

Alibaba

Finishes eight Elo points behind the leader.

  • Ranked No. 2 overall
  • Listed at $2 per 1M input tokens and $6 per 1M output tokens
1302±8
3

Claude Opus 4.7 High

Anthropic

Places third, one point behind Qwen3.8 Max.

  • Ranked No. 3 overall
  • Score interval of 1301±7
1301±7
4

Claude Opus 4.7

Anthropic

The standard configuration follows its High variant by one point.

  • Ranked No. 4 overall
  • Only 10 points behind the leader
1300±7
5

Claude Opus 4.6 High

Anthropic

Extends Anthropic's presence to four of the top five positions.

  • Ranked No. 5 overall
  • Only one point behind Claude Opus 4.7
1299±7
Change

Claude Fable 5 High takes the top position, succeeding the August issue's Claude Fable 5 listing.

Market

The leader is priced at $10 per 1M input tokens and $50 per 1M output tokens. Runner-up Qwen3.8 Max is listed at $2 and $6 respectively, creating a substantial price gap between the top two Vision Arena models.

08
MiMo-V2.6-Pro takes the open-weights lead

Open-Weights Models

MiMo-V2.6-Pro ranks first among open-weights models with an Artificial Analysis Intelligence Index of 46, ahead of GLM-5.3 max at 45 and Kimi K3 max at 44.

Previously: Kimi K3

Current leader
MiMo-V2.6-Pro
Xiaomi

The highest-ranked open-weights model on the current Intelligence Index.

Score
46
  • 01Intelligence Index of 46
  • 02Ranks first among open-weights models
  • 03Listed with a 1M context window
Runners-up
2

GLM-5.3 max

Zhipu AI

Second among current open-weights models, one index point behind the leader.

  • Intelligence Index of 45
  • Ranks second among open-weights models
  • GLM-5.3 is listed with a 1M context window
45
3

Kimi K3 max

Moonshot AI

The August leader now ranks third in the open-weights table.

  • Intelligence Index of 44
  • Ranks third among open-weights models
  • Listed with a 1M context window
44
4

GLM-5.3-Flash

Zhipu AI

GLM-5.3-Flash holds fourth place in the current open-weights ranking.

  • Intelligence Index of 42
  • Ranks fourth among open-weights models
42
5

Qwen3.8 2.4T A95B

Alibaba

Qwen3.8 2.4T A95B records an Intelligence Index of 40.

  • Intelligence Index of 40
  • Listed among the six leading open-weights models
40
6

Qwen3.8-Flash-Next

Alibaba

Qwen3.8-Flash-Next matches Qwen3.8 2.4T A95B with an index score of 40.

  • Intelligence Index of 40
  • Listed among the six leading open-weights models
40
Change

MiMo-V2.6-Pro replaces August leader Kimi K3 at the top of the open-weights ranking.

Market

Open weights remain competitive but trail the absolute proprietary frontier: MiMo-V2.6-Pro scores 46, versus 58 for Claude Opus 5.5 max with fallback.

09
Lowest task cost with a stronger intelligence baseline

Cost-effectiveness

GPT-6 Luna (low) is the value leader at $0.0045 per task and an Intelligence Index of 21. This designation is an editorial inference from the Artificial Analysis comparison table: it combines the lowest stated positive task cost with higher intelligence than the table’s $0.00 utility models.

Previously: DeepSeek V4 Flash 0731

Current leader
GPT-6 Luna (low)
OpenAI

The lowest stated positive cost per task, paired with an Intelligence Index of 21.

Score
$0.0045/task
  • 01$0.0045 cost per task
  • 02Intelligence Index of 21
  • 03Best overall value by editorial comparison of cost and intelligence
Runners-up
2

Granite 4.2 3B

IBM

Listed among the current Cost per Task leaders at $0.01 per task.

  • $0.01 cost per task
  • 3B model designation
  • Second model in the published cost-leader ordering
$0.01/task
3

Ministral 3 3B

Mistral AI

Matches Granite 4.2 3B at a stated cost of $0.01 per task.

  • $0.01 cost per task
  • 3B model designation
  • Included among the three current cost leaders
$0.01/task
Change

GPT-6 Luna (low) replaces August leader DeepSeek V4 Flash 0731.

Market

Command A+ and North Mini Code are listed at $0.00 per task, but their Intelligence Index scores are only 13 and 10 respectively. They should not be treated as the best overall value without that performance qualification.

Editorial · 05 observations

What changed this month

What changed across the AI model landscape this month — distilled from the data above.

01

Frontier Rebase

The frontier text leader has shifted from Claude Opus 5 to Claude Opus 5.5, reflecting ongoing advancements in reasoning capabilities.

02

No verified September update for deepseek-flash-0731

The cited material contains no verified factual entry or URL for deepseek-flash-0731, so this issue makes no model, score, release-date, or ranking claim.

03

MiniMax H3 expands across open and post-trained video tiers

The H3 family holds two consecutive positions in the audio-video top four: Minimax H3 Max post-trained by fal ranks third at 1227 Elo and costs $2.40 per minute, while MiniMax H3 Open Weights ranks fourth at 1220 and leads the open-weights segment with audio.

04

Nano Banana 2 ranks sixth in a reshaped image leaderboard

Nano Banana 2 (Gemini 3.1 Flash Image) records an Elo score of 1122 and ranks sixth in the current Text to Image Arena. GPT Image 2.5 Sunburst leads at 1197, leaving a 75-point gap between the two models.

05

Media leadership is increasingly modality-specific

No single vendor leads every generative-media category. GPT Image 2.5 Sunburst tops text-to-image at 1197 Elo, Gemini Omni Flash leads text-to-video with audio at 1233, Cartesia Sonic 3.6 leads TTS at 1278, and Mureka V9 leads instrumental music at 1177.

Sources
  1. [01]
  2. [02]
  3. [03]
  4. [04]
  5. [05]
  6. [06]
  7. [07]
    Vision Arenacommunity leaderboard
Talk on WhatsApp