VOL. 2026ISSUE 06Updated as of 2026-06-10

LLM Monthly Leaderboard

Eight categories. Twenty-four leading models. Updated monthly. AI-friendly citations included.

9
categories
31
models
10
sources
Share this issueXLinkedIn
01
Text Generation & Reasoning

Text Generation & Reasoning

June 9 — Anthropic ships Claude Fable 5, the first publicly available Mythos-class model, and it debuts at #1 on the Artificial Analysis Intelligence Index v4.0 with a score of 65, clearing Opus 4.8 by 4 points. Anthropic now holds the entire reasoning frontier.

Previously: Claude Opus 4.8

Current leader
Claude Fable 5
Anthropic

Released June 9. First public Mythos-class model — debuts #1 on the AA Intelligence Index v4.0 at 65 and scores 53% on Humanity's Last Exam (7+ points clear).

Score
65
  • 01Intelligence Index v4.0: 65 (#1)
  • 02Humanity's Last Exam: 53%
  • 03Highest AA-Omniscience score to date
  • 041M+ context · Adaptive Reasoning
  • 05Released 2026-06-09
Runners-up
2

Claude Opus 4.8

Anthropic

Released May 28. Adaptive reasoning + max-effort mode; now #2 on the AA Intelligence Index v4.0 behind Fable 5.

  • Intelligence Index v4.0: 61.4
  • Adaptive reasoning mode
  • Coding + agentic upgrade over 4.7
  • Long-running work consistency
  • Released 2026-05-28
61.4
3

GPT-5.5

OpenAI

Apr 24 release. 1M-token context, native MCP + Skills + computer use + hosted shell.

  • Intelligence Index v4.0 (xhigh): 60.2
  • 1M token context
  • Native MCP + Skills
  • Computer use built-in
  • Tool search + web search
  • Released 2026-04-24
60.2
4

Gemini 3.1 Pro

Google

Strong agentic frontier with Antigravity 2.0 platform integration.

  • Intelligence Index v4.0: 57
  • Agentic platform (Antigravity 2.0) integration
  • Tied with GPT-5.5 medium and Qwen3.7 Max
57
5

Grok 4.3

xAI

Released Apr 30. AA Intelligence Index 53 with always-on reasoning, 1M context, and aggressive pricing — the cheapest of the frontier four.

  • Intelligence Index v4.0: 53
  • Always-on reasoning · 1M context
  • ~40% / 60% cheaper input/output vs Grok 4.20
  • Strong agentic + tool use
53
Change

Released June 9, 2026. AA Intelligence Index v4.0: 65 (Adaptive Reasoning, max effort) — 4 points ahead of Opus 4.8 (61.4) and GPT-5.5 xhigh (60.2); Gemini 3.1 Pro (57) and Grok 4.3 (53) follow.

Market

Fable 5 is the Mythos-class tier above Opus, gated with extra safeguards; Opus 4.8 stays the default workhorse for cost-sensitive high-volume work.

02
Image Generation

Image Generation

OpenAI ships GPT Image 2 with token-based pricing and 50% Batch API discount. Recraft V4.1 holds the Artificial Analysis quality leaderboard, while Adobe Firefly enterprise mode remains the rights-cleared default.

Previously: GPT Image-2

Current leader
GPT Image 2
OpenAI

Released April 21. State-of-the-art quality with token pricing and Batch API support.

Score
None
  • 01Token-based pricing
  • 02Batch API at 50% discount
  • 03Flexible image sizes
  • 04High-fidelity inputs
  • 05Released 2026-04-21
Runners-up
2

Recraft V4.1

Recraft

Leads Artificial Analysis text-to-image arena on raw output quality.

  • Top of AA text-to-image quality
  • Strong control / style transfer
  • Designer-grade output
None
3

Adobe Firefly Image 4

Adobe

IP-cleared training data; the enterprise-safe choice for commercial use.

  • Trained on licensed assets
  • Indemnification for enterprise
  • Native Adobe Creative Cloud integration
None
Change

Released April 21, 2026. Token-based pricing, Batch API at 50% off.

Market

Recraft V4.1 leads AA's text-to-image arena on raw quality; GPT Image 2 wins on ecosystem and pricing transparency.

03
Video Generation

Video Generation

Seedance 2.0 (ByteDance) tops the Artificial Analysis text-to-video Arena with audio at ELO 1215 — the first model to make synced audio-visual generation state of the art. Google's Veo 3.5 leads the silent-cinematic tier; Kling 4 holds the value end. OpenAI's Sora exited the field after its app was discontinued on April 26, 2026.

Previously: Veo 3.5

Current leader
Seedance 2.0
ByteDance

Tops the Artificial Analysis text-to-video Arena (with audio) at ELO 1215 — ahead of Veo. Native audio-visual generation: 15-second multi-shot clips with synced sound from text, image, audio and video inputs.

Score
1215
  • 01AA T2V Arena (w/ audio) #1 · ELO 1215
  • 0215s multi-shot · synced audio
  • 03Multimodal input (text/image/audio/video)
  • 04Dual-Branch Diffusion Transformer
Runners-up
2

Veo 3.5

Google

Production-grade film output with strongest temporal coherence.

  • 1080p output
  • Strong physics simulation
  • Long-shot temporal coherence
  • Native Gemini API integration
None
3

Kling 4

Kuaishou 快手

Dominant in APAC short-form ad creative; fastest iteration cycle in the space.

  • 9:16 vertical native
  • Fastest editorial iteration
  • TikTok / Douyin native style
  • Low-latency generation
None
Change

Seedance 2.0 holds #1 on the with-audio text-to-video Arena; Sora 2 drops out after OpenAI discontinued the Sora app (Apr 26, 2026).

Market

Seedance 2.0 for synced audio-visual and multi-shot narrative; Veo 3.5 for cinematic fidelity; Kling 4 for cost.

04
Code Generation & Agentic Coding

Code Generation & Agentic Coding

Claude Fable 5 launches June 9 and immediately tops SWE-bench Pro at 80.3% — about 11 points clear of Opus 4.8 (69.2%) on the same uncontaminated multi-language benchmark. Pro scores run far below the contaminated Verified set, so treat 80% as the honest new frontier.

Previously: Claude Opus 4.8

Current leader
Claude Fable 5
Anthropic

#1 on SWE-bench Pro at 80.3% (June 9 launch) — 11 points ahead of the field. Mythos-class agentic coding for the hardest multi-file, long-horizon work.

Score
80.3
  • 01SWE-bench Pro: 80.3% (#1)
  • 0211 points clear of Opus 4.8
  • 03Long-horizon agentic edits
  • 041M+ context · Mythos-class
  • 05Released 2026-06-09
Runners-up
2

Claude Opus 4.8

Anthropic

#2 on SWE-bench Pro at 69.2% (updated June 8), behind Fable 5. The cost-sensitive default for multi-file refactoring, code review, and long-horizon agentic edits.

  • SWE-bench Pro: 69.2% (#2)
  • Multi-file refactor SOTA
  • Code review top
  • Vision-aware coding
69.2
3

GPT-5.3 Codex (xhigh)

OpenAI

#3 on SWE-bench Pro (June 8). Specialized terminal-code agent; AA Intelligence Index 54.

  • AA Intelligence Index: 54
  • Sandboxed shell built-in
  • Strong agentic loops
  • Terminal-Bench best
54
4

Cursor Composer 2.5

Cursor

Ranks #3 on AA Coding Agent Index. IDE-native pair programming with multi-file context.

  • AA Coding Agent Index: #3
  • IDE-native context
  • Multi-file edits
  • Inline diff workflow
None
Change

Fable 5 debuts June 9 at #1 on SWE-bench Pro (80.3%); Opus 4.8 (69.2%) #2, GPT-5.3 Codex #3.

Market

Fable 5 for the hardest multi-file refactors and long-horizon agents; Opus 4.8 as the cost-sensitive default; GPT-5.3 Codex for sandboxed terminal agents; Cursor Composer 2.5 for IDE-native pair programming.

05
Voice / Speech

Voice / Speech

OpenAI's Realtime 2 (May 7) brings configurable-reasoning speech-to-speech to general availability; AA's text-to-speech crown goes to Fun-Realtime-TTS, and MAI-Transcribe-1.5 wins STT on accuracy-speed.

Previously: ElevenLabs v3

Current leader
Realtime 2
OpenAI

GA on May 7. Configurable-reasoning speech-to-speech with realtime translate + Whisper variants.

Score
None
  • 01Configurable reasoning
  • 02Speech-to-speech agents
  • 03Streaming translate variant
  • 04Streaming STT variant
  • 05Released 2026-05-07
Runners-up
2

ElevenLabs v3

ElevenLabs

Industry default for character voice cloning and audiobook production.

  • Character voice cloning SOTA
  • 100+ languages
  • Long-form audiobook quality
  • Emotion control
None
3

Fun-Realtime-TTS

Fun (Alibaba DAMO)

Tops AA text-to-speech leaderboard on quality metrics.

  • AA TTS leaderboard #1
  • Sub-200ms latency
  • Multi-speaker streaming
  • Strong CJK
None
Change

Realtime 2 family shipped May 7 (gpt-realtime-2 / -translate / -whisper).

Market

Realtime 2 for agentic voice; ElevenLabs v3 for character voice cloning; Fun-Realtime-TTS for raw TTS quality; MAI-Transcribe-1.5 for transcription.

06
Music Generation

Music Generation

Suno v6 widens the gap on full-song coherence and lyric prosody; Udio v3 keeps pushing studio-grade mixing; Lyria (Google) integrates into Gemini Omni for any-to-music workflows.

Previously: Suno v5.5

Current leader
Suno v6
Suno

Rolling release expected late June. Best full-song coherence and lyric prosody.

Score
None
  • 01Full-song coherence SOTA
  • 02Lyric prosody best
  • 03Multilingual vocal
  • 04Style transfer
Runners-up
2

Udio v3

Udio

Studio-grade mixing with stem-level output for producers.

  • Stem-level export
  • Studio-grade mixing
  • Strong electronic genres
  • DAW-friendly workflow
None
3

Lyria (via Gemini Omni)

Google

Folded into Gemini Omni for any-to-music + cross-modal generation.

  • Gemini Omni native
  • Cross-modal generation
  • Image / video → music workflows
None
Change

Suno v6 expected late June; v5.5 remains the deployed default.

Market

Suno v6 for full-song generation; Udio v3 for mixed studio-grade stems; Lyria via Gemini Omni for cross-modal generation.

07
Vision / Multimodal Understanding

Vision / Multimodal Understanding

Anthropic sweeps LMArena Vision top-3 with Opus 4.7-thinking, 4.6-thinking, and 4.7. Opus 4.8 is too freshly released to appear on Arena ELO but is expected to consolidate the lead by end of June.

Previously: GPT-4o Vision

Current leader
Claude Opus 4.7-thinking
Anthropic

Tops LMArena Vision. Best OCR + chart + document understanding.

Score
1309
  • 01LMArena Vision ELO: 1309
  • 02OCR SOTA
  • 03Chart understanding
  • 04Document Q&A
Runners-up
2

GPT-5.5

OpenAI

Strongest image-grounded reasoning chains; native computer-use vision pipeline.

  • Image-grounded reasoning best
  • Computer use vision
  • 1M token multimodal
  • Released 2026-04-24
None
3

Gemini 3.1 Pro

Google

Best video understanding and long-form temporal reasoning.

  • Video understanding SOTA
  • Long-form temporal
  • Robotics-ER 1.6 integration
  • Multimodal context 2M+
None
Change

LMArena Vision top-3 all Anthropic (1309 / 1303 / 1298 ELO).

Market

Anthropic for OCR + document Q&A + chart understanding; GPT-5.5 for image-grounded reasoning; Gemini 3.1 Pro for video understanding.

08
Open-Source / Open-Weights

Open-Source / Open-Weights

Kimi K2.6 (Moonshot) leads open weights at AA Intelligence Index 54 — within 7 points of frontier closed models. DeepSeek V4 Pro (MIT, 52) is the #2 open reasoning model, and Google's new Gemma 4 12B (Apache 2.0, 2026-06-03) packs native multimodal into a 16GB-laptop footprint. The closed-vs-open gap is the narrowest it has ever been.

Previously: Llama 4

Current leader
Kimi K2.6
Moonshot AI

Open-weights leader on AA Intelligence Index. Closes the closed-source gap to 7 points.

Score
54
  • 01AA Intelligence Index: 54
  • 02Open weights
  • 03Strong Chinese + English
  • 04Long-context retention
Runners-up
2

DeepSeek V4 Pro

DeepSeek

MIT-licensed open weights at AA Intelligence Index 52 — #3 of 89 overall and the #2 open reasoning model behind only Kimi K2.6.

  • AA Intelligence Index: 52 (#3/89)
  • MIT license · open weights
  • MoE 1.6T total / 49B active
  • 1M-token context
52
3

Gemma 4 12B

Google

Released 2026-06-03 under Apache 2.0. Encoder-free native multimodal (text/image/audio/video), 256K context, runs on a 16GB laptop — performance nearing last-gen 27B.

  • Apache 2.0 · open weights
  • 256K context · native multimodal
  • Runs on 16GB VRAM
  • MMLU-Pro 77.2 · GPQA-Diamond 78.8
4

Qwen3.7 Plus

Alibaba

Best open-source for Chinese-language self-host deployment.

  • AA Intelligence Index: 53
  • Best Chinese open-source
  • Strong tool use
  • Open weights
53
Change

Kimi K2.6 (54) tops open weights; DeepSeek V4 Pro (52) #2; Gemma 4 12B brings 16GB-laptop multimodal.

Market

Kimi K2.6 for general-purpose open deployment; DeepSeek V4 Pro for cheap frontier-grade reasoning; Gemma 4 12B for on-device multimodal; Qwen3.7 Plus for Chinese self-host.

09
Intelligence per Dollar

Cost-Effectiveness / Value

The 2026 value war is led by China's open-source camp. DeepSeek V4 Flash delivers near-flagship intelligence (Index 47) at roughly a tenth of the price — a blended cost near $0.06 per 1M tokens. The leaderboard makes one warning explicit: 'Flash' and 'mini' branding does not mean cheap. Gemini 3.5 Flash scores a strong 55 but costs over 20× more per full Intelligence-Index run.

Current leader
DeepSeek V4 Flash
DeepSeek

The intelligence-per-dollar king. AA Intelligence Index 47 at $0.14/$0.28 per 1M tokens — about a tenth of comparable Flash flagships, with the lowest cache-hit price of any 2026 frontier model.

Score
47
  • 01AA Intelligence Index: 47
  • 02$0.14 in / $0.28 out per 1M
  • 03Blended ≈ $0.06 / 1M
  • 04MIT open weights · 1M context
Runners-up
2

DeepSeek V4 Pro

DeepSeek

Best balance in the high-intelligence + low-price quadrant. Intelligence Index 52 (top-3 overall) at $0.435/$0.87 — a fraction of same-tier flagships like GPT-5.5 and Claude Opus.

  • AA Intelligence Index: 52 (#3/89)
  • $0.435 in / $0.87 out per 1M
  • Top-3 intelligence overall
  • MIT open weights
52
3

Qwen3.7 Plus

Alibaba

Keeps the value race from being a single-vendor story. Intelligence Index 53 at $0.40 input — a higher score than V4 Pro; the trade-off is pricier output and slower generation.

  • AA Intelligence Index: 53
  • $0.40 in / $1.16 out per 1M
  • Highest score in the value tier
  • Strong Chinese + tool use
53
Change

New category. DeepSeek V4 Flash leads intelligence-per-dollar; open-source models sweep the value tier.

Market

DeepSeek V4 Flash for highest-volume low-cost workloads; V4 Pro when you need stronger reasoning but still want to save; Qwen3.7 Plus to diversify away from a single vendor.

Editorial · 07 observations

What changed this month

What changed across the AI model landscape this month — distilled from the data above.

01

Anthropic's Mythos-Class Fable 5 Sweeps Reasoning + Code

Claude Fable 5 (Mythos-class, June 9) debuts #1 on the AA Intelligence Index (65) and #1 on SWE-bench Pro coding (80.3%, ~11 points clear); Opus 4.8 sits #2 on both, and Opus 4.7-thinking still holds the LMArena Vision and Document arenas. No single lab has held reasoning + vision + code this completely since GPT-4-era OpenAI.

02

GPT-5.5 Brings 1M Context + Native MCP/Skills

OpenAI's April 24 GPT-5.5 ships with 1M-token context, native MCP, Skills, hosted shell, computer use, tool search, and web search — turning the API itself into an agent runtime.

03

Google's Gemini Omni — Any-to-Any Generation

May launch of Gemini Omni unifies image / audio / video generation; Antigravity 2.0 platform turns Gemini 3.5 into an agentic substrate. Google's bet: not the smartest single model, but the most integrated stack.

04

Open-Source Closes to a 7-Point Gap

Kimi K2.6 (54) is within 7 points of Claude Opus 4.8 (61) on AA Intelligence Index. Meta's muse-spark cracks LMArena top-5 at 1489. Closed-source's moat is the smallest it has ever been.

05

Five Chinese Labs in AA Top 15

Qwen3.7 Max (Alibaba, 57), MiniMax-M3 (55), Kimi K2.6 (Moonshot, 54), MiMo-V2.5-Pro (Xiaomi, 54), and Qwen3.7 Plus (Alibaba, 53) all sit in the AA Intelligence Index top 15 — Chinese labs are no longer 'catching up', they're inside the frontier.

06

Sub-200ms Voice Agents Are Now Commodity

OpenAI Realtime 2 (May 7) + Gemini 3.1 Flash TTS (April) + Fun-Realtime-TTS push real-time voice agents from research to production. Speech-to-speech with reasoning is now a checkbox feature.

07

The value war is a Chinese open-source story

DeepSeek V4 Flash delivers Intelligence Index 47 at a blended ~$0.06 per 1M tokens — open-weights models from DeepSeek and Qwen now sweep the intelligence-per-dollar top tier, while 'Flash'/'mini'-branded closed models like Gemini 3.5 Flash cost over 20× more per Index run.

Sources
  1. [01]
  2. [02]
    LMArena Leaderboardcommunity leaderboard
  3. [03]
  4. [04]
    OpenAI Changelogofficial changelog
  5. [05]
    Anthropic Newsofficial changelog
  6. [06]
    Google DeepMind Blogofficial changelog
  7. [07]
    DeepSeek API Pricingofficial changelog
  8. [08]
    Google Gemma 4 Launchofficial changelog
  9. [09]
  10. [10]
预约 demo