Claude Opus 5.5 (xhigh with fallback)
AnthropicStrong performance with fallback capabilities.
- Fallback logic
- Robust reasoning
Eight categories. Twenty-four leading models. Updated monthly. AI-friendly citations included.
One issue a month, all archived. Trace how the AI capability map shifts month over month.
The latest leaders in reasoning capabilities, showcasing the most advanced models in logical processing and decision-making.
Previously: Claude Opus 5
Leading in reasoning with a robust score.
Strong performance with fallback capabilities.
High reasoning capacity with fallback.
Maximized reasoning with fallback.
Xhigh reasoning with fallback.
Maximized reasoning capabilities.
Claude Opus 5.5 has taken the lead from Claude Opus 5, with an Intelligence Index score of 58.
Anthropic continues to dominate the reasoning category with its Claude Opus series.
GPT Image 2.5 Sunburst (max) leads the current Text to Image Arena at Elo 1197, followed by GPT Image 2.5 Flare (max) at 1191 and GPT Image 2 (high) at 1171. Sunburst is listed with 13,433 samples.
Previously: GPT Image 2 / Nano Banana 2
Current Text to Image Arena leader at Elo 1197.
Second in the current Text to Image Arena at Elo 1191.
Third in the current Text to Image Arena at Elo 1171.
Fourth in the current Text to Image Arena at Elo 1150.
Fifth in the current Text to Image Arena at Elo 1147.
Sixth in the current Text to Image Arena at Elo 1122.
The leader changed from the August issue's GPT Image 2 / Nano Banana 2 pairing.
Among open-weight image models, Ideogram 4.0 (Quality) leads at 1011, followed by Ideogram 4.0 at 1003 and FLUX.2 [dev] at 1000.
Gemini Omni Flash retains first place in the audio-enabled Text to Video arena with 1233 Elo, narrowly ahead of Wan 3.0 at 1229. The ordering differs without audio, where Wan 3.0 leads with 1336.
Retains the audio-video lead, four Elo points ahead of Wan 3.0.
Places second with audio and leads the separate no-audio ranking.
Ranks only two Elo points behind Wan 3.0 while posting the lowest listed price among the top five.
The highest-ranked open-weights video model with audio.
Completes a tightly grouped top five, 23 Elo points behind the leader.
The leader is unchanged from August, but Wan 3.0 has moved into second place while two MiniMax H3 variants occupy ranks three and four.
Competition is tight at the top: only 23 Elo points separate first-ranked Gemini Omni Flash from fifth-ranked Dreamina Seedance 2.0 720p. Pricing is more dispersed, ranging from $2.40 to $12.00 per minute among these five models.
Terminal-Bench 2.1 evaluates software engineering, system administration, data processing, model training, and security. This independently run 89-task refresh includes environment and instruction fixes, with pass@1 averaged over three repeats.
Previously: Claude Opus 5
The highest verified Terminal-Bench 2.1 result.
Finishes only 0.4 percentage points behind the leading configuration.
Completes an all-Claude Fable 5.1 top three.
Claude Fable 5.1 Adaptive Reasoning Max Effort with Default Fallback replaces the August issue’s Claude Opus 5 as the category leader. Because Terminal-Bench 2.1 is a refreshed evaluation, the September percentage should not be treated as a direct comparison with the prior issue’s score.
The leaderboard’s top three results are effort variants of Claude Fable 5.1, spanning 89.9% to 91.4%. For enterprise buyers, this highlights the growing importance of selecting a reasoning-effort tier rather than evaluating only the base model name.
Released in August 2026, Cartesia Sonic 3.6 leads the Provider Voice Arena at 1278 Elo. Inworld Realtime TTS-2 follows at 1243, ahead of SpeechifyAI Simba 3.2 at 1238 and Alibaba Qwen-Audio-3.0-TTS-Plus at 1235.
Previously: Realtime 2
The August 2026 release is the new Provider Voice Arena leader.
Second overall in the current provider voice ranking.
Places third, five Elo points behind Realtime TTS-2.
Alibaba's entry places fourth in a tightly grouped top tier.
Luna TTS rounds out the arena's top five.
Inworld holds two of the top six positions with its standard and Flash variants.
Cartesia Sonic 3.6 replaces August leader Realtime 2, while Inworld Realtime TTS-2 enters the current table in second place.
API pricing varies substantially across the top six: SpeechifyAI Simba 3.2 is listed at $6.6 per 1M characters, compared with $49.0 for Sonic 3.6 and $80.0 for VUI Labs Luna TTS. Breeze TTS 2 is the highest-ranked open-weights TTS model at 1203.
Mureka V9 leads the Instrumental Music Arena with 1177 Elo, followed by Suno V5.5 at 1171 and Mureka V8 at 1143. The leaderboard is based on blind preference votes and evaluates instrumental and vocals modes separately.
Previously: Suno v5.5
Instrumental Music Arena leader, supported by 2,237 samples.
Runner-up with 1171 Elo across 2,222 samples.
Third place with 1143 Elo and 3,130 samples.
Fourth place with 1137 Elo across 2,713 samples.
Fifth place with 1113 Elo and 2,212 samples.
Sixth place with 1102 Elo across 2,959 samples.
Mureka V9 replaces August leader Suno v5.5, while Suno V5.5 now ranks second.
Mureka and Suno each hold two of the top four positions. Sample counts range from 2,212 for StepAudio 3 Music to 3,130 for Mureka V8; the source does not provide music-model pricing.
The Vision Arena table dated September 13, 2026 covers 152 models and 1,314,119 votes. Claude Fable 5 High leads at 1310±8, narrowly ahead of Qwen3.8 Max at 1302±8 and two Claude Opus 4.7 configurations.
Previously: Claude Fable 5
Leads the September Vision Arena.
Finishes eight Elo points behind the leader.
Places third, one point behind Qwen3.8 Max.
The standard configuration follows its High variant by one point.
Extends Anthropic's presence to four of the top five positions.
Claude Fable 5 High takes the top position, succeeding the August issue's Claude Fable 5 listing.
The leader is priced at $10 per 1M input tokens and $50 per 1M output tokens. Runner-up Qwen3.8 Max is listed at $2 and $6 respectively, creating a substantial price gap between the top two Vision Arena models.
MiMo-V2.6-Pro ranks first among open-weights models with an Artificial Analysis Intelligence Index of 46, ahead of GLM-5.3 max at 45 and Kimi K3 max at 44.
Previously: Kimi K3
The highest-ranked open-weights model on the current Intelligence Index.
Second among current open-weights models, one index point behind the leader.
The August leader now ranks third in the open-weights table.
GLM-5.3-Flash holds fourth place in the current open-weights ranking.
Qwen3.8 2.4T A95B records an Intelligence Index of 40.
Qwen3.8-Flash-Next matches Qwen3.8 2.4T A95B with an index score of 40.
MiMo-V2.6-Pro replaces August leader Kimi K3 at the top of the open-weights ranking.
Open weights remain competitive but trail the absolute proprietary frontier: MiMo-V2.6-Pro scores 46, versus 58 for Claude Opus 5.5 max with fallback.
GPT-6 Luna (low) is the value leader at $0.0045 per task and an Intelligence Index of 21. This designation is an editorial inference from the Artificial Analysis comparison table: it combines the lowest stated positive task cost with higher intelligence than the table’s $0.00 utility models.
Previously: DeepSeek V4 Flash 0731
The lowest stated positive cost per task, paired with an Intelligence Index of 21.
Listed among the current Cost per Task leaders at $0.01 per task.
Matches Granite 4.2 3B at a stated cost of $0.01 per task.
GPT-6 Luna (low) replaces August leader DeepSeek V4 Flash 0731.
Command A+ and North Mini Code are listed at $0.00 per task, but their Intelligence Index scores are only 13 and 10 respectively. They should not be treated as the best overall value without that performance qualification.
What changed across the AI model landscape this month — distilled from the data above.
The frontier text leader has shifted from Claude Opus 5 to Claude Opus 5.5, reflecting ongoing advancements in reasoning capabilities.
The cited material contains no verified factual entry or URL for deepseek-flash-0731, so this issue makes no model, score, release-date, or ranking claim.
The H3 family holds two consecutive positions in the audio-video top four: Minimax H3 Max post-trained by fal ranks third at 1227 Elo and costs $2.40 per minute, while MiniMax H3 Open Weights ranks fourth at 1220 and leads the open-weights segment with audio.
Nano Banana 2 (Gemini 3.1 Flash Image) records an Elo score of 1122 and ranks sixth in the current Text to Image Arena. GPT Image 2.5 Sunburst leads at 1197, leaving a 75-point gap between the two models.
No single vendor leads every generative-media category. GPT Image 2.5 Sunburst tops text-to-image at 1197 Elo, Gemini Omni Flash leads text-to-video with audio at 1233, Cartesia Sonic 3.6 leads TTS at 1278, and Mureka V9 leads instrumental music at 1177.