VOL. 2026ISSUE 08Updated as of 2026-08-09

LLM Monthly Leaderboard

Eight categories. Twenty-four leading models. Updated monthly. AI-friendly citations included.

9
categories
39
models
9
sources
Share this issueXLinkedIn
01
Text Generation & Reasoning

Text Generation & Reasoning

August 9 scrape of Artificial Analysis shows Claude Opus 5 still #1, but the whole frontier band rebased upward: Opus 5 max 63 (July issue 61), Fable 5 with fallback 62, GPT-5.6 Sol max 61, open-weights Kimi K3 max 60. Muse Spark 1.2 (xhigh) enters at 57; Grok 4.5 high sits at 56.

Current leader
Claude Opus 5
Anthropic

Still #1 on AA Intelligence Index — now 63 max, up from 61 in the July issue.

Score
63
  • 01Intelligence Index: 63 (max)
  • 02SWE-bench Verified ~96%
  • 03$5 / $25 per 1M tokens
  • 041M context
  • 05Released 2026-07-24
Runners-up
2

Claude Fable 5

Anthropic

AA lists Fable 5 with fallback at 62 — #2 overall after the August rebase.

  • Intelligence Index: 62 (with fallback)
  • Mythos-class safeguards
  • 1M+ context · Adaptive Reasoning
  • Released 2026-06-09
62
3

GPT-5.6 Sol

OpenAI

Sol max climbs to 61 on AA; still the OpenAI flagship tier.

  • Intelligence Index: 61 (max)
  • Terminal-Bench 2.1: 88.8%
  • $5 / $30 per 1M tokens
  • 1M context
  • GA 2026-07-09
61
4

Kimi K3

Moonshot AI

Open-weights K3 max hits 60 — global #4, only 3 points off Opus 5.

  • Intelligence Index: 60 (open #1)
  • 2.8T MoE · 1M context
  • LMArena WebDev #1 (July)
  • Open weights · $3/$15 per 1M
  • Released 2026-07-17
60
5

Muse Spark 1.2

Muse

New mid-frontier entrant at AA Index 57 (xhigh) — displaces Grok from the July top-5 shape.

  • Intelligence Index: 57 (xhigh)
  • New vs July top tier
  • Closed-weights contender
57
6

Grok 4.5

xAI

AA high mode at 56; still the cheap always-on reasoning pick at $2/$6.

  • Intelligence Index: 56 (high)
  • Aggressive pricing: $2 / $6 per 1M
  • Always-on reasoning · 1M context
  • Released 2026-07-08
56
Change

AA Intelligence Index top-4 all +2–3 vs July issue. Muse Spark 1.2 is new in the mid-frontier; DeepSeek V4 Flash 0731 lands at 52.

Market

Opus 5 remains the default frontier workhorse at $5/$25; Fable 5 is the Mythos-class ceiling; Sol is OpenAI’s tiered answer; Kimi K3 is still the open model inside the frontier band.

02
Image Generation

Image Generation

GPT Image 2 (high) extends its Artificial Analysis Image Arena lead to ELO 1357 (July issue 1338). Reve 2.1 holds #2 at 1314. Google’s Nano Banana 2 (Gemini 3.1 Flash Image Preview) debuts in the top 3 at 1307. Seedream 5.0 Pro sits at 1266.

Current leader
GPT Image 2
OpenAI

AA Image Arena ELO 1357 (high) — quality and ecosystem crown extended.

Score
1357
  • 01AA T2I Arena ELO: 1357 (#1)
  • 02Token-based pricing
  • 03Batch API at 50% discount
  • 04High-fidelity inputs
Runners-up
2

Reve 2.1

Reve

Holds #2 at ELO 1314 with layered generation.

  • AA T2I Arena ELO: 1314 (#2)
  • Layered generation
  • $24 per 1k images
  • Day-one top-tier quality
1314
3

Nano Banana 2

Google

Gemini 3.1 Flash Image Preview branding — new AA top-3 entrant at 1307.

  • AA T2I Arena ELO: 1307 (#3)
  • Gemini 3.1 Flash Image Preview
  • Google stack integration
1307
4

MAI-Image-2.5

Microsoft

Enterprise image default at ELO 1298.

  • AA T2I Arena ELO: 1298
  • Enterprise integration
  • Microsoft stack
1298
5

Seedream 5.0 Pro

ByteDance

Top Chinese production pick at ELO 1266 with strong text rendering.

  • AA T2I Arena ELO: 1266
  • Layer separation
  • Strong text rendering
1266
Change

GPT Image 2 high 1338→1357; Nano Banana 2 enters top 3; Seedream 5.0 Pro still the leading Chinese production option in the top tier.

Market

Quality + ecosystem: GPT Image 2; layered generation value: Reve 2.1; Google preview stack: Nano Banana 2; Chinese text/layout: Seedream 5.0 Pro.

03
Video Generation

Video Generation

Gemini Omni Flash keeps the AA text-to-video (with audio) crown at ELO 1244 (July 1240). MiniMax H3 surges to #2 at 1240. Dreamina Seedance 2.0 720p is #3 at 1224; Wan2.7-260612 and HappyHorse-1.1 round out the top 5. Sora API sunset remains 2026-09-24.

Current leader
Gemini Omni Flash
Google

AA T2V with audio #1 at ELO 1244 — still the synced A/V crown.

Score
1244
  • 01AA T2V Arena (w/ audio) #1 · ELO 1244
  • 02Without-audio arena also #1 · ELO 1324
  • 03Native audio-visual sync
  • 04Gemini Omni stack
Runners-up
2

MiniMax H3

MiniMax

New #2 on AA T2V with audio at ELO 1240 — the biggest August video mover.

  • AA T2V Arena (w/ audio) #2 · ELO 1240
  • Without-audio #2 · ELO 1306
  • Fast climb vs July field
1240
3

Seedance 2.0

ByteDance

Dreamina Seedance 2.0 720p at ELO 1224 — still the ByteDance production workhorse.

  • AA T2V Arena (w/ audio) #3 · ELO 1224
  • Production pipeline maturity
  • Multimodal references
1224
4

Wan2.7-260612

Alibaba

Alibaba Wan2.7 build 260612 at ELO 1161 on the with-audio arena.

  • AA T2V Arena (w/ audio) ELO: 1161
  • Alibaba Cloud native
1161
5

Kling 3.0 Pro

Kuaishou 快手

Still the APAC short-form value line (July ELO band); not displaced by H3/Seedance fight at the very top.

  • Native 4K/60fps line
  • Turbo iteration speed
  • TikTok / Douyin native style
1110
Change

With-audio arena: Gemini Omni Flash 1240→1244; MiniMax H3 takes #2 from Seedance; Seedance 2.0 still #3.

Market

Synced A/V arena: Gemini Omni Flash; aggressive #2 challenger: MiniMax H3; ByteDance production: Seedance 2.0; APAC short-form still Kling-family adjacent.

04
Code Generation & Agentic Coding

Code Generation & Agentic Coding

Claude Opus 5 remains the agentic-coding leader on the July SWE-bench Verified high-water mark (~96%). Kimi K3 keeps the open-weights / LMArena WebDev crowd-vote crown (1679). GPT-5.6 Sol still leads Terminal-Bench 2.1 at 88.8%. No cleaner public number beating Opus 5 was confirmed on the August 9 pass.

Current leader
Claude Opus 5
Anthropic

SWE-bench Verified ~96% high-water mark still stands; AA Index now 63.

Score
96
  • 01SWE-bench Verified: ~96% (#1)
  • 02AA Intelligence Index: 63
  • 03$5 / $25 per 1M tokens
  • 04Long-horizon agentic edits
Runners-up
2

Kimi K3

Moonshot AI

Open-weights WebDev arena king at 1679; AA Index 60.

  • LMArena WebDev: 1679 (#1)
  • AA Intelligence Index: 60
  • Open weights · 2.8T MoE
  • $3 / $15 per 1M tokens
1679
3

Claude Fable 5

Anthropic

Mythos-class coding ceiling; AA Index 62 overall.

  • SWE-bench Verified: 95%
  • Mythos-class agentic coding
  • 1M+ context
95
4

GPT-5.6 Sol

OpenAI

Best sandboxed terminal agent — Terminal-Bench 2.1 88.8%.

  • Terminal-Bench 2.1: 88.8% (#1)
  • SWE-bench Pro: 64.6%
  • Sandboxed shell built-in
88.8
Change

No crown flip: Opus 5 still SWE-bench leader; Kimi K3 still WebDev #1; AA overall Index now 63/60 for Opus/Kimi.

Market

Hardest refactors: Opus 5; open frontend + deploy: Kimi K3; Mythos ceiling: Fable 5; sandboxed terminal agents: GPT-5.6 Sol.

05
Voice / Speech

Voice / Speech

No new speech-to-speech king — OpenAI Realtime 2 still leads agentic voice. On AA TTS, Qwen-Audio-3.0-TTS-Plus remains #1 at ELO 1229 (July 1234; same crown, slight Elo drift). Speechify Simba 3.2 is #2 at 1227; Gemini 3.1 Flash TTS #3 at 1210.

Current leader
Realtime 2
OpenAI

Still the agentic voice default — configurable-reasoning speech-to-speech.

  • 01Configurable reasoning
  • 02Speech-to-speech agents
  • 03Streaming translate + STT variants
  • 04Released 2026-05-07
Runners-up
2

Qwen-Audio-3.0-TTS-Plus

Alibaba

AA TTS #1 at ELO 1229 — China still holds the component TTS crown.

  • AA TTS ELO: 1229 (#1)
  • Strong multilingual + CJK
  • Alibaba Cloud native
1229
3

Simba 3.2

Speechify

AA TTS #2 at ELO 1227 — inches behind Qwen.

  • AA TTS ELO: 1227 (#2)
  • Consumer reading voice quality
1227
4

ElevenLabs v3

ElevenLabs

Still the character voice cloning default; Scribe v2 realtime STT remains in stack.

  • Character voice cloning SOTA
  • Scribe v2 realtime STT
  • Music v2 API
  • 100+ languages
Change

TTS crown unchanged (Qwen-Audio-3.0-TTS-Plus); Elo 1234→1229. Agentic S2S still Realtime 2.

Market

Agentic voice: Realtime 2; raw TTS quality: Qwen-Audio-3.0-TTS-Plus; character cloning: ElevenLabs v3.

06
Music Generation

Music Generation

Suno v6 still has not shipped as of the August 9 check. Suno v5.5 remains the full-song default. ElevenLabs Music v2 (API since June 15) and Udio v3 stems stay the secondary studio picks; Lyria remains on the Gemini Omni cross-modal path.

Current leader
Suno v5.5
Suno

Still the latest shipped Suno — best full-song coherence and lyric prosody.

  • 01Full-song coherence SOTA
  • 02Lyric prosody best
  • 03Multilingual vocal
  • 04Style transfer
Runners-up
2

Udio v3

Udio

Studio-grade mixing with stem-level export.

  • Stem-level export
  • Studio-grade mixing
  • Strong electronic genres
  • DAW-friendly workflow
3

ElevenLabs Music v2

ElevenLabs

API-first licensed-catalog music generation since June 15.

  • API-first music generation
  • Licensed catalog
  • Music Finetunes API
  • Voice stack integration
Change

No crown change. Suno v6 still unannounced/unshipped.

Market

Full songs: Suno v5.5; stems/studio: Udio v3; licensed API music: ElevenLabs Music v2; cross-modal: Lyria via Gemini Omni.

07
Vision / Multimodal Understanding

Vision / Multimodal Understanding

Claude Fable 5 remains the LMArena Vision leader at ELO 1327 from the July recalculation — no higher public Vision number was confirmed on August 9. Overall AA Intelligence still has Opus 5 at 63 and Fable 5 at 62, so multimodal understanding stays Anthropic-heavy. Gemini 3.6 Flash (Index 52) is the fast Google multimodal pick.

Current leader
Claude Fable 5
Anthropic

LMArena Vision #1 at 1327; AA overall Index 62.

Score
1327
  • 01LMArena Vision ELO: 1327 (#1)
  • 02AA Intelligence Index: 62
  • 03OCR + document SOTA
  • 041M+ context
Runners-up
2

Claude Opus 5

Anthropic

Multimodal flagship with AA overall #1 at 63.

  • AA Intelligence Index: 63 (#1 overall)
  • Document Q&A + charts
  • $5 / $25 per 1M tokens
63
3

GPT-5.6 Sol

OpenAI

Image-grounded reasoning at AA Index 61.

  • AA Intelligence Index: 61
  • Image-grounded reasoning
  • 1M context
61
4

Gemini 3.6 Flash

Google

New on AA overall at 52 — fast Google multimodal stack pick.

  • AA Intelligence Index: 52
  • High output speed on AA
  • Gemini multimodal stack
52
Change

Vision crown unchanged (Fable 5 @ 1327). Gemini 3.6 Flash appears on AA overall at 52.

Market

OCR/docs/charts: Anthropic; image-grounded reasoning: GPT-5.6 Sol; video understanding: Gemini line.

08
Open-Source / Open-Weights

Open-Source / Open-Weights

Kimi K3 max reaches AA Intelligence Index 60 — still open #1 and global #4, only 3 points behind Opus 5 (63). DeepSeek V4 Flash 0731 jumps to 52 without a price change ($0.14/$0.28). GLM-5.2 max is ~53; MiniMax-M3 is 45.

Current leader
Kimi K3
Moonshot AI

Open #1 at AA Index 60 — global #4, 3 points off Opus 5.

Score
60
  • 01AA Intelligence Index: 60 (open #1)
  • 022.8T MoE · 1M context
  • 03LMArena WebDev #1
  • 04$3 / $15 per 1M tokens
Runners-up
2

DeepSeek V4 Flash 0731

DeepSeek

0731 refresh to Index 52 at unchanged $0.14/$0.28 — open value king.

  • AA Intelligence Index: 52
  • $0.14 in / $0.28 out per 1M
  • MIT open weights · 1M context
52
3

GLM-5.2

Zhipu AI

AA max around 53 — strong open Chinese alternative.

  • AA Intelligence Index: ~53 (max)
  • Pay-as-you-go API available
53
4

Gemma 4 12B

Google

Apache 2.0 on-device multimodal default.

  • Apache 2.0 · open weights
  • 256K context · native multimodal
  • Runs on 16GB VRAM
Change

Kimi K3 57→60; DeepSeek V4 Flash 0731 47→52; open crown unchanged.

Market

Frontier open deploy: Kimi K3; pure $/IQ: DeepSeek V4 Flash 0731; on-device multimodal: Gemma 4 12B.

09
Intelligence per Dollar

Cost-Effectiveness / Value

DeepSeek V4 Flash 0731 is the new value headline: Intelligence Index 52 at the same $0.14/$0.28 pricing that made Flash the July king at 47. MiniMax-M3 remains the cheap high-coding tier at Index 45. Kimi K3 offers Index 60 but at $3/$15 — quality, not pure $/IQ.

Previously: DeepSeek V4 Flash

Current leader
DeepSeek V4 Flash 0731
DeepSeek

Index 52 at $0.14/$0.28 — still the intelligence-per-dollar king after the 0731 refresh.

Score
52
  • 01AA Intelligence Index: 52
  • 02$0.14 in / $0.28 out per 1M
  • 03Blended cost still ~1/10 of flagships
  • 04MIT open weights · 1M context
Runners-up
2

MiniMax-M3

MiniMax

AA Index 45 — cheap high-coding value tier.

  • AA Intelligence Index: 45
  • Strong coding-per-dollar
  • Standard API pricing tier
45
3

GLM-5.2

Zhipu AI

Higher absolute price than Flash, stronger index (~53 max).

  • AA Intelligence Index: ~53 (max)
  • Pay-as-you-go API
53
4

Kimi K3

Moonshot AI

Not the cheapest — but best open quality at $3/$15 with Index 60.

  • AA Intelligence Index: 60
  • $3 / $15 per 1M tokens
  • Open weights frontier-adjacent
60
Change

Leader name updates to V4 Flash 0731; Index 47→52 with unchanged price. MiniMax-M3 still the SWE-bench value pick from July narrative.

Market

Highest volume low-cost: V4 Flash 0731; stronger open reasoning on a budget: Kimi K2.6-class / GLM; quality-per-dollar frontier-adjacent: Kimi K3 if you can pay $15 out.

Editorial · 05 observations

What changed this month

What changed across the AI model landscape this month — distilled from the data above.

01

Frontier scores rebased upward

Artificial Analysis Intelligence Index top-4 all rose ~2–3 points vs the July issue (Opus 63, Fable 62, Sol 61, Kimi 60), without a crown flip at the very top.

02

DeepSeek Flash 0731 value jump

DeepSeek V4 Flash 0731 reaches Index 52 at unchanged $0.14/$0.28 pricing — the open cheap tier got smarter without getting pricier.

03

MiniMax H3 crashes the video podium

MiniMax H3 takes #2 on AA text-to-video with audio (ELO 1240), slotting between Gemini Omni Flash (1244) and Seedance 2.0 (1224).

04

Google Nano Banana 2 enters image top-3

Nano Banana 2 (Gemini 3.1 Flash Image Preview) debuts at AA Image Arena ELO 1307, behind only GPT Image 2 and Reve 2.1.

05

Sora sunset still on the clock

OpenAI’s Sora API shutdown date remains 2026-09-24; video demand continues consolidating around Gemini Omni, MiniMax H3, and Seedance.

Sources
  1. [01]
  2. [02]
  3. [03]
  4. [04]
  5. [05]
  6. [06]
    OpenAI Changelogofficial changelog
  7. [07]
    Anthropic Newsofficial changelog
  8. [08]
    DeepSeek API Pricingofficial changelog
  9. [09]
    Moonshot AI / Kimiofficial changelog
预约 demo