PandaNpc
PandaNpc

AI Model Leaderboards

Multi-source AI model rankings — arena.ai human blind-vote Elo. Chat, code, image & video boards.

RankModelVendorElo ?Votes ?Input ?OutputCache ?
1Anthropicclaude-opus-5-maxAnthropic1705 ±152,1925.0025.00.500
2Moonshotkimi-k3-maxMoonshot1676 ±124,3663.0015.00.300
3Anthropicclaude-opus-5-highAnthropic1669 ±113,8735.0025.00.500
4Alibabaqwen3.8-maxAlibaba1668 ±181,5632.006.000.250
5Anthropicclaude-fable-5Anthropic1630 ±96,31010.050.01.00
6OpenAIgpt-5.6-sol-xhigh (codex-harness)OpenAI1620 ±96,0005.0030.00.500
7Z.aiglm-5.2-maxZ.ai1586 ±96,3610.7602.420.140
8DeepSeekdeepseek-v4-flash-highDeepSeek1577 ±181,3190.0900.1800.018
9Anthropicclaude-opus-4-8-thinkingAnthropic1566 ±88,9345.0025.00.500
10Anthropicclaude-opus-4-7Anthropic1561 ±711,7555.0025.00.500
11Anthropicclaude-opus-4-7-thinkingAnthropic1557 ±712,3875.0025.00.500
12SpaceXAIgrok-4.5SpaceXAI1549 ±103,9922.006.000.300
13Anthropicclaude-opus-4-6-thinkingAnthropic1545 ±614,2735.0025.00.500
14Anthropicclaude-sonnet-5-highAnthropic1542 ±104,7222.0010.00.200
15Anthropicclaude-opus-4-8Anthropic1539 ±87,8955.0025.00.500

Source: arena.ai · Updated Aug 1, 2026 · PandaNpc Updated Aug 6, 2026, 14:00 GMT+8View full leaderboard →

How to read these numbers

arena.ai (formerly LMArena) shows users two anonymous answers to the same prompt and asks them to pick one, then derives an Elo rating from those battles. It measures which answer people prefer — not objective capability. A likeable style can win, while a rigorous but verbose answer can lose.

Votes are the battles a model has accumulated. A newly listed model has few votes and a wide confidence interval, so its rank swings easily — always read the rank together with the ± interval.

The Agent board works differently from the rest: it measures real agent sessions rather than blind-vote preference. Its headline metric, Net Improvement, is a percentage gain — not an Elo rating — so it cannot be compared across boards. arena.ai also reports five more dimensions (Confirmed Success, Steerability, Bash Recovery and others) on the full leaderboard.

The usage board comes from OpenRouter and counts real API token share — developers voting with their wallets, which complements the subjective preference boards.

The Intelligence Index board comes from Artificial Analysis: it distills many public benchmarks into a single 0-100 Intelligence Index, measuring objective capability rather than subjective preference — a natural complement to arena.ai's human blind vote. Source: artificialanalysis.ai.