LLM Leaderboard & AI Model Rankings

One page that aggregates the major AI benchmarks and leaderboards — Chatbot Arena Elo (arena.ai), the Artificial Analysis Intelligence Index and speech arenas, and OpenRouter usage — across chat, coding, TTS, speech-to-text, image, video and agents.

RankModelVendorElo ?Votes ?Input ?OutputCache ?
1OpenAIgpt-image-2.5-sunburstOpenAI1424 ±810,884———
2OpenAIgpt-image-2.5-flareOpenAI1401 ±810,058———
3OpenAIgpt-image-2 (medium)OpenAI1383 ±488,744———
4Microsoft AImai-image-2.6Microsoft AI1335 ±618,625———
5Revereve-2.1Reve1302 ±88,305———
6SpaceXAIgrok-imagine-image-2.0 (low)SpaceXAI1301 ±87,257———
7Metamuse-imageMeta1276 ±535,422———
8Revereve-2.0Reve1269 ±615,802———
9Googlegemini-3.1-flash-image (nano-banana-2) [web-search]Google1261 ±453,2280.5003.00—
10Bytedanceseedream-5.0-proBytedance1256 ±485,690———
11Alibabaqwen-image-3.0-proAlibaba1256 ±611,520———
12Microsoft AImai-image-2.5Microsoft AI1254 ±461,448———
13Googlegemini-3.1-flash-lite-image (nano-banana-2-lite)Google1250 ±617,8510.2501.50—
14Googlegemini-3-pro-image-2k (nano-banana-pro)Google1247 ±3168,447———
15OpenAIgpt-image-1.5-high-fidelityOpenAI1238 ±3165,437———
16Googlegemini-3-pro-image-preview (nano-banana-pro)Google1234 ±586,4512.0012.00.200
17Alibabaqwen-image-2.1Alibaba1228 ±94,445———
18Ideogramideogram-4.0-qualityIdeogram1205 ±453,371———
19Alibabaqwen-image-2.0-pro-2026-06-22Alibaba1192 ±613,123———
20Luma AIuni-1.1-maxLuma AI1189 ±614,083———
21Microsoft AImai-image-2Microsoft AI1184 ±551,477———
22Luma AIuni-1.1Luma AI1183 ±441,412———
23SpaceXAIgrok-imagine-imageSpaceXAI1170 ±3262,132———
24Recraftrecraft-v4.1-utility-proRecraft1169 ±112,618———
25NvidiaCosmos3-Super-Text2Image (Agentic)Nvidia1165 ±710,856———
26SpaceXAIgrok-imagine-image-proSpaceXAI1162 ±499,444———
27Black Forest Labsflux-2-maxBlack Forest Labs1162 ±3123,187———
28Black Forest Labsflux-2-flexBlack Forest Labs1157 ±3156,912———
29Revereve-v1.5Reve1155 ±439,755———
30Black Forest Labsflux-2-proBlack Forest Labs1154 ±3205,540———
31NvidiaCosmos3-Super-Text2ImageNvidia1152 ±618,581———
32Googlegemini-2.5-flash-image-preview (nano-banana)Google1150 ±2878,4510.3002.500.030
33Tencenthunyuan-image-3.0Tencent1150 ±3180,550———
34Googleimagen-ultra-4.0-generate-001Google1148 ±4388,772———
35Bytedanceseedream-4.5Bytedance1147 ±3287,128———
36Black Forest Labsflux-2-devBlack Forest Labs1146 ±478,080———
37Bytedanceseedream-4-2kBytedance1140 ±712,566———
38Bytedanceseedream-5.0-liteBytedance1138 ±3123,178———
39Alibabawan2.6-t2iAlibaba1137 ±3223,535———
40Recraftrecraft-v4.1-proRecraft1133 ±102,795———
41Googleimagen-4.0-generate-001Google1129 ±3539,788———
42Alibabaqwen-image-2512Alibaba1125 ±398,363———
43Kreakrea-2-mediumKrea1123 ±450,800———
44Bytedanceseedream-4-falBytedance1117 ±711,842———
45Alibabawan2.5-t2i-previewAlibaba1116 ±3271,922———
46HiDreamhidream-o1-imageHiDream1116 ±459,511———
47OpenAIgpt-image-1OpenAI1116 ±3268,466———
48Recraftrecraft-v4Recraft1115 ±3129,674———
49Bytedanceseedream-4-high-res-falBytedance1113 ±3182,669———
50Kreakrea-2-turboKrea1113 ±543,913———
51Kreakrea-2-largeKrea1110 ±449,906———
52OpenAIgpt-image-1-miniOpenAI1110 ±3171,058———
53Alibabawan2.7-image-proAlibaba1103 ±530,229———
54Alibabawan2.7-imageAlibaba1101 ±530,621———
55Recraftrecraft-v4.1-flashRecraft1100 ±93,533———
56Microsoft AImai-image-1Microsoft AI1093 ±498,414———
57Alibabaz-image-turboAlibaba1084 ±524,291———
58Bytedanceseedream-3Bytedance1082 ±536,890———
59Black Forest Labsflux-1-kontext-maxBlack Forest Labs1074 ±365,517———
60Black Forest Labsflux-2-klein-9bBlack Forest Labs1070 ±3150,780———
61Alibabaqwen-image-prompt-extendAlibaba1060 ±3716,829———
62Black Forest Labsflux-1-kontext-proBlack Forest Labs1059 ±3332,376———
63Googleimagen-3.0-generate-002Google1058 ±3359,786———
64Alibabaqwen-imageAlibaba1057 ±384,555———
65Ideogramideogram-v3-qualityIdeogram1048 ±4118,542———
66Luma AIphotonLuma AI1036 ±4130,158———
67?p-image—1033 ±3109,782———
68Black Forest Labsflux-2-klein-4bBlack Forest Labs1030 ±3152,358———
69Runwayrunway-gen4Runway1024 ±454,376———
70Recraftrecraft-v3Recraft1021 ±4195,623———
71Black Forest Labsflux-1.1-proBlack Forest Labs1016 ±370,460———
72Ideogramideogram-v2Ideogram1013 ±472,090———
73Leonardo AIlucid-originLeonardo AI1013 ±3287,345———
74Z.aiglm-imageZ.ai1012 ±94,852———
75Googlegemini-2.0-flash-preview-image-generationGoogle975 ±3257,241———
76Black Forest Labsflux-1-dev-fp8Black Forest Labs969 ±449,221———
77OpenAIdall-e-3OpenAI968 ±4238,749———
78Black Forest Labsflux-1-kontext-devBlack Forest Labs941 ±4216,238———
79?stable-diffusion-v35-large—938 ±523,396———
80BytedancebagelBytedance898 ±612,363———

Source: arena.ai · Updated Sep 24, 2026 · PandaNpc Updated Sep 28, 2026, 06:00 UTCView full leaderboard →

How to read these numbers

An LLM leaderboard tells you which model people prefer; an AI benchmark tells you what a model can objectively do; usage data tells you what developers actually ship. No single ranking answers "which is the best AI model", so this page aggregates all three — Chatbot Arena Elo from arena.ai, the Artificial Analysis Intelligence Index and speech arenas, and OpenRouter token share — and refreshes them daily.

arena.ai (formerly LMArena) shows users two anonymous answers to the same prompt and asks them to pick one, then derives an Elo rating from those battles. It measures which answer people prefer — not objective capability. A likeable style can win, while a rigorous but verbose answer can lose.

Votes are the battles a model has accumulated. A newly listed model has few votes and a wide confidence interval, so its rank swings easily — always read the rank together with the ± interval.

The Agent board works differently from the rest: it measures real agent sessions rather than blind-vote preference. Its headline metric, Net Improvement, is a percentage gain — not an Elo rating — so it cannot be compared across boards. arena.ai also reports five more dimensions (Confirmed Success, Steerability, Bash Recovery and others) on the full leaderboard.

The usage board comes from OpenRouter and counts real API token share — developers voting with their wallets, which complements the subjective preference boards.

The speech boards come from Artificial Analysis: TTS uses blind-listening Speech Arena Elo; ASR uses AA-WER v2, where lower is better; and native Speech-to-Speech uses an equal-weight composite of Big Bench Audio, Full Duplex Bench and τ-Voice, where higher is better. Compare scores only within the same board. Source: artificialanalysis.ai.

The Intelligence Index board comes from Artificial Analysis: it distills many public benchmarks into a single 0-100 Intelligence Index, measuring objective capability rather than subjective preference — a natural complement to arena.ai's human blind vote. Source: artificialanalysis.ai.