Arena ranking
Browse Agent Steerability results, model ratings and supporting selection context.
Data as of 2026-09-08
This source leaderboard ranking orders models by their original Arena rank rather than an aggregated benchmark score.
Rank | Model | Source Agent score | LLMBoard Score | Observations | Official input / 1M | Output speed |
|---|
| Rank01 | ModelAN | Source Agent score16.09% | LLMBoard ScoreN/A | Observations21.3K | Official input / 1M$5 | Output speed62.88 tok/s |
| Rank02 | ModelAN | Source Agent score12.59% | LLMBoard ScoreN/A | Observations27.2K | Official input / 1M$5 | Output speed42.00 tok/s |
| Rank03 | ModelAN | Source Agent score11.89% | LLMBoard ScoreN/A | Observations27.9K | Official input / 1M$10 | Output speed43.18 tok/s |
| Rank04 | ModelAN | Source Agent score11.52% | LLMBoard ScoreN/A | Observations19.8K | Official input / 1M$2 | Output speed42.00 tok/s |
| Rank05 | ModelOP | Source Agent score8.88% | LLMBoard ScoreN/A | Observations26.4K | Official input / 1M$4 | Output speed2.36 tok/s |
| Rank06 | ModelOP | Source Agent score8.54% | LLMBoard ScoreN/A | Observations43.5K | Official input / 1M$5 | Output speed134.94 tok/s |
| Rank07 | ModelOP | Source Agent score8.54% | LLMBoard ScoreN/A | Observations18.2K | Official input / 1M$2 | Output speed99.35 tok/s |
| Rank08 | ModelMA | Source Agent score7.73% | LLMBoard ScoreN/A | Observations5.8K | Official input / 1M$0.95 | Output speed1.24 tok/s |
| Rank09 | ModelAN | Source Agent score7.27% | LLMBoard ScoreN/A | Observations28.3K | Official input / 1M$5 | Output speed42.00 tok/s |
| Rank10 | ModelXA | Source Agent score6.84% | LLMBoard ScoreN/A | Observations26.9K | Official input / 1M$2 | Output speed80.00 tok/s |
| Rank11 | ModelMA | Source Agent score4.68% | LLMBoard ScoreN/A | Observations5.8K | Official input / 1M$0.95 | Output speedN/A |
| Rank12 | ModelAN | Source Agent score4.58% | LLMBoard ScoreN/A | Observations34.3K | Official input / 1M$5 | Output speed123.15 tok/s |
| Rank13 | ModelZA | Source Agent score4.48% | LLMBoard ScoreN/A | Observations52.5K | Official input / 1M$1.4 | Output speed3.96 tok/s |
| Rank14 | ModelOP | Source Agent score4.16% | LLMBoard ScoreN/A | Observations60.5K | Official input / 1M$2.5 | Output speed50.00 tok/s |
| Rank15 | ModelAC | Source Agent score4.15% | LLMBoard ScoreN/A | Observations17.1K | Official input / 1M$2 | Output speed60.87 tok/s |
| Rank16 | ModelXA | Source Agent score3.86% | LLMBoard ScoreN/A | Observations15K | Official input / 1M$2 | Output speed132.69 tok/s |
| Rank17 | ModelOP | Source Agent score1.83% | LLMBoard ScoreN/A | Observations18.5K | Official input / 1M$0.20 | Output speed41.87 tok/s |
| Rank18 | ModelMA | Source Agent score1.24% | LLMBoard ScoreN/A | Observations80.6K | Official input / 1M$3 | Output speedN/A |
| Rank19 | ModelDE | Source Agent score0.67% | LLMBoard ScoreN/A | Observations22K | Official input / 1MN/A | Output speed45.48 tok/s |
| Rank20 | ModelZA | Source Agent score0.63% | LLMBoard ScoreN/A | Observations43.7K | Official input / 1M$1.4 | Output speed472.69 tok/s |
| Rank21 | ModelAN | Source Agent score0.58% | LLMBoard ScoreN/A | Observations1.2K | Official input / 1M$10 | Output speed7.67 tok/s |
| Rank22 | ModelZA | Source Agent score0.23% | LLMBoard ScoreN/A | Observations15.1K | Official input / 1M$0.075 | Output speed74.28 tok/s |
| Rank23 | ModelDE | Source Agent score0.02% | LLMBoard ScoreN/A | Observations45.1K | Official input / 1M$0.14 | Output speed15.91 tok/s |
| Rank24 | ModelTE | Source Agent score-0.10% | LLMBoard ScoreN/A | Observations5.9K | Official input / 1MN/A | Output speedN/A |
| Rank25 | ModelTE | Source Agent score-0.15% | LLMBoard ScoreN/A | Observations11.8K | Official input / 1MN/A | Output speedN/A |
| Rank26 | ModelAC | Source Agent score-0.39% | LLMBoard ScoreN/A | Observations15.3K | Official input / 1MN/A | Output speedN/A |
| Rank27 | ModelMA | Source Agent score-0.50% | LLMBoard ScoreN/A | Observations5.5K | Official input / 1M$1.5 | Output speedN/A |
| Rank28 | ModelAN | Source Agent score-0.99% | LLMBoard ScoreN/A | Observations39K | Official input / 1M$3 | Output speed7.67 tok/s |
| Rank29 | ModelAC | Source Agent score-1.18% | LLMBoard ScoreN/A | Observations8.2K | Official input / 1MN/A | Output speedN/A |
| Rank30 | ModelZA | Source Agent score-1.27% | LLMBoard ScoreN/A | Observations51.6K | Official input / 1M$1.4 | Output speedN/A |
This leaderboard overview compares the source ranking, score gaps and observed task volume.
Use confidence and observed task volume to judge how firmly each leaderboard position is supported; these source proportions remain separate from benchmark coverage.
Vendor representation among the top-ranked models in this leaderboard.
A practical summary of the first five entries in this agent steerability AI model leaderboard, with benchmark evidence and pricing kept in context.
Common questions about the Agent Steerability Leaderboard.
Claude Opus 5 is currently ranked first with 16.09 Agent score.
The current leaders are Claude Opus 5 (rank #1), Claude Opus 4.8 (rank #2), and Claude Fable 5 (rank #3).
GLM 5.3 Flash has the lowest matched official input price at $0.075 per 1M tokens.
The fastest matched records are GLM 5.3 (472.69 tok/s via FriendliAI), Gemini 3.5 Flash (210.94 tok/s via Google), and GPT-5.5 (134.94 tok/s via OpenAI).
No. Source Arena results describe a particular task or preference signal. Capability benchmarks, prices and runtime can produce different rankings.
This source-native ranking follows the original source metric. Benchmark evidence is shown separately and is not substituted for the source result.
This page currently compares 49 models.
A model may not yet have a matching Arena result, or multiple version results may resolve to the same model.
Ranking basisThis agent steerability AI model leaderboard uses the source agent score and rank. The leaderboard ranking keeps matched price and speed data separate from benchmark evidence.
Selection summary
Claude Opus 5 currently leads the agent steerability ranking at 16.09%. Compare evidence coverage, source activity and price separately before choosing a model for production.
Use this leaderboard with the supporting benchmark results and coverage details above. A leaderboard position summarizes the selected ranking signal; it does not replace workload-specific testing.