Arena ranking
Browse Agent Task Outcome Explicit results, model ratings and supporting selection context.
Data as of 2026-09-08
This source leaderboard ranking orders models by their original Arena rank rather than an aggregated benchmark score.
Rank | Model | Source Agent score | LLMBoard Score | Observations | Official input / 1M | Output speed |
|---|
| Rank01 | ModelAN | Source Agent score22.45% | LLMBoard ScoreN/A | Observations2.5K | Official input / 1M$10 | Output speed7.67 tok/s |
| Rank02 | ModelMA | Source Agent score16.58% | LLMBoard ScoreN/A | Observations74.2K | Official input / 1M$3 | Output speedN/A |
| Rank03 | ModelAN | Source Agent score14.89% | LLMBoard ScoreN/A | Observations12.5K | Official input / 1M$5 | Output speed62.88 tok/s |
| Rank04 | ModelZA | Source Agent score14.14% | LLMBoard ScoreN/A | Observations13.5K | Official input / 1M$0.075 | Output speed74.28 tok/s |
| Rank05 | ModelTE | Source Agent score13.56% | LLMBoard ScoreN/A | Observations6.5K | Official input / 1MN/A | Output speedN/A |
| Rank06 | ModelAC | Source Agent score12.46% | LLMBoard ScoreN/A | Observations9.5K | Official input / 1MN/A | Output speedN/A |
| Rank07 | ModelZA | Source Agent score12.10% | LLMBoard ScoreN/A | Observations34.9K | Official input / 1M$1.4 | Output speed472.69 tok/s |
| Rank08 | ModelDE | Source Agent score11.75% | LLMBoard ScoreN/A | Observations20.4K | Official input / 1MN/A | Output speed45.48 tok/s |
| Rank09 | ModelGO | Source Agent score10.42% | LLMBoard ScoreN/A | Observations4.6K | Official input / 1M$0.75 | Output speed51.56 tok/s |
| Rank10 | ModelAC | Source Agent score10.37% | LLMBoard ScoreN/A | Observations14K | Official input / 1M$2 | Output speed60.87 tok/s |
| Rank11 | ModelGO | Source Agent score9.54% | LLMBoard ScoreN/A | Observations19.6K | Official input / 1M$0.75 | Output speed117.86 tok/s |
| Rank12 | ModelZA | Source Agent score8.63% | LLMBoard ScoreN/A | Observations42.5K | Official input / 1M$1.4 | Output speed3.96 tok/s |
| Rank13 | ModelDE | Source Agent score8.46% | LLMBoard ScoreN/A | Observations35.4K | Official input / 1M$0.14 | Output speed15.91 tok/s |
| Rank14 | ModelXA | Source Agent score7.83% | LLMBoard ScoreN/A | Observations12.9K | Official input / 1M$2 | Output speed132.69 tok/s |
| Rank15 | ModelAN | Source Agent score7.37% | LLMBoard ScoreN/A | Observations27.5K | Official input / 1M$10 | Output speed43.18 tok/s |
| Rank16 | ModelAC | Source Agent score6.98% | LLMBoard ScoreN/A | Observations13.9K | Official input / 1MN/A | Output speedN/A |
| Rank17 | ModelOP | Source Agent score6.68% | LLMBoard ScoreN/A | Observations22.8K | Official input / 1M$4 | Output speed2.36 tok/s |
| Rank18 | ModelME | Source Agent score6.60% | LLMBoard ScoreN/A | Observations12.9K | Official input / 1M$1.25 | Output speed28.44 tok/s |
| Rank19 | ModelXA | Source Agent score6.30% | LLMBoard ScoreN/A | Observations23.8K | Official input / 1M$2 | Output speed80.00 tok/s |
| Rank20 | ModelAN | Source Agent score5.93% | LLMBoard ScoreN/A | Observations21.1K | Official input / 1M$5 | Output speed42.00 tok/s |
| Rank21 | ModelME | Source Agent score5.01% | LLMBoard ScoreN/A | Observations56.5K | Official input / 1M$1.25 | Output speed14.00 tok/s |
| Rank22 | ModelAN | Source Agent score3.51% | LLMBoard ScoreN/A | Observations25.8K | Official input / 1M$5 | Output speed42.00 tok/s |
| Rank23 | ModelAN | Source Agent score2.80% | LLMBoard ScoreN/A | Observations19.4K | Official input / 1M$2 | Output speed42.00 tok/s |
| Rank24 | ModelAN | Source Agent score2.50% | LLMBoard ScoreN/A | Observations27.8K | Official input / 1M$5 | Output speed123.15 tok/s |
| Rank25 | ModelOP | Source Agent score2.07% | LLMBoard ScoreN/A | Observations39.1K | Official input / 1M$5 | Output speed134.94 tok/s |
| Rank26 | ModelMA | Source Agent score1.15% | LLMBoard ScoreN/A | Observations5.7K | Official input / 1M$0.95 | Output speedN/A |
| Rank27 | ModelOP | Source Agent score1.14% | LLMBoard ScoreN/A | Observations53.9K | Official input / 1M$2.5 | Output speed50.00 tok/s |
| Rank28 | ModelOP | Source Agent score0.83% | LLMBoard ScoreN/A | Observations15K | Official input / 1M$0.20 | Output speed41.87 tok/s |
| Rank29 | ModelZA | Source Agent score0.17% | LLMBoard ScoreN/A | Observations45.3K | Official input / 1M$1.4 | Output speedN/A |
| Rank30 | ModelGO | Source Agent score-1.62% | LLMBoard ScoreN/A | Observations68.3K | Official input / 1M$1.5 | Output speed210.94 tok/s |
This leaderboard overview compares the source ranking, score gaps and observed task volume.
Use confidence and observed task volume to judge how firmly each leaderboard position is supported; these source proportions remain separate from benchmark coverage.
Vendor representation among the top-ranked models in this leaderboard.
A practical summary of the first five entries in this agent task outcome explicit AI model leaderboard, with benchmark evidence and pricing kept in context.
Common questions about the Agent Task Outcome Explicit Leaderboard.
Claude Fable 5.1 is currently ranked first with 22.45 Agent score.
The current leaders are Claude Fable 5.1 (rank #1), Kimi K3 (rank #2), and Claude Opus 5 (rank #3).
GLM 5.3 Flash has the lowest matched official input price at $0.075 per 1M tokens.
The fastest matched records are GLM 5.3 (472.69 tok/s via FriendliAI), Gemini 3.5 Flash (210.94 tok/s via Google), and GPT-5.5 (134.94 tok/s via OpenAI).
No. Source Arena results describe a particular task or preference signal. Capability benchmarks, prices and runtime can produce different rankings.
This source-native ranking follows the original source metric. Benchmark evidence is shown separately and is not substituted for the source result.
This page currently compares 49 models.
A model may not yet have a matching Arena result, or multiple version results may resolve to the same model.
Ranking basisThis agent task outcome explicit AI model leaderboard uses the source agent score and rank. The leaderboard ranking keeps matched price and speed data separate from benchmark evidence.
Selection summary
Claude Fable 5.1 currently leads the agent task outcome explicit ranking at 22.45%. Compare evidence coverage, source activity and price separately before choosing a model for production.
Use this leaderboard with the supporting benchmark results and coverage details above. A leaderboard position summarizes the selected ranking signal; it does not replace workload-specific testing.